|
Author :
Shilpa Nigam |
IIT Madras’ Bodhan AI and AI4Bharat have launched a suite of foundational, open-weight AI models tailored specifically for Indian languages.
Developed in partnership with NVIDIA under the Bharat EduAI Stack initiative, these models aim to bridge the digital divide by serving as Digital Public Goods for education and public infrastructure across India.
Research & Institutional Portal: Visit the official Bodhan.AI Hub, India's Center of Excellence in Artificial Intelligence for Education incubated at IIT Madras.
Language AI Research Platform: Explore specific model breakthroughs and dataset portfolios via the AI4Bharat Research Lab Platform.
Model Repositories: To download open weights and integrate the software, you can directly monitor the Bodhan AI Hugging Face Profile.
The newly released suite includes four foundational models trained using the NVIDIA NeMo framework:
Indic-Transcribe (ASR): Delivers automatic speech recognition for 27 languages, capturing diverse regional dialects and accents.
Indic-Speak (TTS): Provides high-quality text-to-speech synthesis across 23 languages.
Indic-Translate: Enables seamless machine translation across 22 Indian languages.
Indic-OCR: Optimised for document layout parsing and optical character recognition across 23 languages to digitise regional textbooks.
To ensure wide accessibility, the model weights have been made public on Hugging Face for local fine-tuning. For enterprise scaling, low-cost APIs are available on the Bodhan AI Console, which runs entirely on India’s sovereign cloud infrastructure.
While commercial EdTech entities can leverage the paid APIs, the technology remains completely free for students, teachers, and state governments.
Alongside the models, two application tools-a Student Tutor Bot for classes 6-12 and a Teacher Assistant Bot for lesson planning-have also been rolled out.
Chat with us