IIT Madras’ Bodhan AI Partners with NVIDIA to Launch Open-Weight Indic Language AI Suite for Education
Bodhan AI IIT Madras: IIT Madras-incubated Bodhan AI has partnered with NVIDIA and AI4Bharat to release open-weight Indic language AI models for education. Covering speech recognition, text-to-speech, translation, and OCR across up to 27 languages, the initiative forms part of India's sovereign digital public infrastructure to personalize regional-language learning.
Bodhan AI IIT Madras: The Bodhan AI, Centre of Excellence in AI for Education, incubated at IIT Madras, has partnered with NVIDIA and AI4Bharat in the development of a set of open-weight Indic language AI models for Education & public digital infrastructure. Developed under the stewardship of Principal Investigator Prof. Mitesh Khapra, this project acts as the core technology layer of the Bharat EduAI Stack, India’s sovereign Digital Public Infrastructure (DPI) for Education. The set includes four basic open-weight models that cover core multimodal capabilities: Indic Transcribe for Automatic Speech Recognition (ASR) in 27 Indic languages, Indic Speak for Text-to-Speech (TTS) in 23 Indic languages, Indic Translate for Machine Translation (MT) in 22 Indic languages, and Indic OCR for document scanning in 23 Indic scripts.
Using the NVIDIA NeMo framework and post-trained on the NVIDIA Nemotron 3.5 ASR, the architecture is optimised to understand regional dialects, accents, and complex handwritten educational documents. Through NVIDIA TensorRT-LLM and vLLM inference microservices, this set of open-weight models & APIs enables edtech companies, researchers, and government organisations to develop NCERT/SCERT-aligned student bots and teacher workspaces for millions of students who study in their native language.
IIT Madras: Key Highlights of the Bodhan AI Suite
-
Four Pillar Capabilities: Offers open-source models in Automatic Speech Recognition (Indic-Transcribe), Text-to-Speech (Indic-Speak), Machine Translation (Indic-Translate), and Optical Character Recognition (Indic-OCR).
-
Linguistic Diversity: Covers up to 27 Indian languages, including accents and dialects, enabling native language communication between students and educators.
-
Based on NVIDIA Framework: Designed using NVIDIA NeMo framework, trained using post-NVIDIA Nemotron 3.5 ASR, and deployed using NVIDIA TensorRT-LLM/vLLM microservices for inference.
-
Bharat EduAI Stack: Acts as the foundational layer of open weights within India’s sovereign Digital Public Infrastructure (DPI), providing free-of-cost education applications for learners, educators, and state governments.
IIT Madras: Overview Of The Bodhan AI
Below mentioned are the overview of the IIT Madras’ Bodhan AI:
| Model Name | AI Capability | Language Support / Focus Area |
| Indic-Transcribe | Automatic Speech Recognition (ASR) | Covers 27 Indian languages; optimized for local dialects & accents |
| Indic-Speak | Text-to-Speech (TTS) | Covers 23 languages; generates natural voice responses for learning |
| Indic-Translate | Machine Translation (MT) | Covers 22 scheduled Indian languages for seamless content translation |
| Indic-OCR | Optical Character Recognition | Covers 23 scripts; digitizes printed textbooks and handwritten sheets |
Also Read:
NEET UG 2026 Round 2 Final Seat Allotment Released? Check Latest Update Here
Siddhi Sharma is an education journalist at Jagran Josh. A Journalism and Mass Communication graduate from IP University, she brings sharp newsroom instincts developed during her previous stint at Zee News. At Jagran Josh, Siddhi specializes in decoding the educational updates. Her coverage is highly exam-centric, ranging from curated news blogs for competitive exams to crucial school board and university news. Combining her strong media foundations with a research-driven approach, she creates reliable, high-utility content that helps students and aspirants stay ahead of the curve. Her writing is factual, engaging, and tailored to meet the fast-paced needs of modern learners and exam aspirants.
