Bodhan AI Releases 4 New Open Models for Indian Languages

0
1
Bodhan AI Releases 4 New Open Models for Indian Languages


A Hindi lesson can mix English terms (loan words), scanned tables and handwritten equations. Making that content searchable, translating it and reading it aloud requires several kinds of AI. Bodhan AI and AI4Bharat’s four new models target those jobs across Indian languages. 

Released in September 2026, the models cover document parsing, translation, speech recognition and speech generation, with support for mixed languages and scripts. In this article, we break down each model, its benchmarks, limitations and access options. 

The 4 Model in a Nutshell 

Model Task Architecture
IndicOCR Page image → structured text 33M layout parser + 0.8B OCR
Indic-Translate Text → translated text 4B effective parameters; 32K context
Indic-Transcribe Speech → transcript Core and Flexible; 1.2B each
Indic-Speak Text → speech 3.36B stack; 45 voices

This is the biggest display of frontier development across Indic language that I’ve seen since the release of Indic-LM Arena back in November 2025. But what the platform offered initially as a blueprint, the following releases are making progress across different facets of that leaderboard.  

1. IndicOCR: Read the Text and Keep the Structure

IndicOCR

IndicOCR parses printed documents in English and all 22 scheduled Indian languages across 13 scripts. It also recognizes handwriting in English and 12 Indian languages, including Hindi, Bengali, Tamil, Telugu and Urdu. 

It uses two stages. IndicDocLayout, a 33M model based on PP-DocLayoutV3, detects page blocks and their reading order. IndicBlockOCR, built on Qwen3.5-0.8B with a Sarvam tokenizer, transcribes those blocks. Equations become LaTeX, while tables retain their structure. 

OmniDocBench

Bodhan reports 92.76 on OmniDocBench v1.6, evaluated on its 610-page English subset, and 82.20 on the English olmOCR-Bench subset. Its internal IndicOCR-Printed benchmark reports 86.2% word-level accuracy across 22 Indian languages and English. 

These measure different things. An English document-parsing score does not establish equal accuracy across Indian languages. The internal printed benchmark evaluates individual blocks, separating text recognition from page ordering. 

Where it fits: Digitizing textbooks, making regional archives searchable, or preparing scanned pages for RAG. AV’s guide to using Mistral OCR in a RAG system explains the broader document-to-retrieval workflow. 

What still needs work: Bodhan flags dense reading order, difficult handwriting and layouts outside education. Handwriting support for the remaining 10 Indian languages is planned. 

2. Indic-Translate: Translate Whole Documents

Indic-Translate

Indic-Translate is a translation-focused fine-tune of Gemma 4 E4B IT, described as having 4B effective parameters and a 32K-token context window. It supports English and all 22 scheduled Indian languages in both directions. 

Its main feature is document-level translation. It is trained to preserve Markdown, LaTeX, tables and code while translating the surrounding language. It also supports Romanized text, transliteration and code-mixed input. 

On the release’s in-house document test, Indic-Translate scores 58.97 dBLEU, compared with 47.44 for Sarvam Translate and 31.93 for IndicTrans2-1B. Its reported word error rate is 0.4326, versus 0.5553 and 0.8304, respectively. Higher dBLEU and lower WER indicate closer matches to reference translations. 

Bodhan reports leading both metrics across all 22 languages in that evaluation. Human evaluation is still in progress, so these results do not establish a universal winner across translation tasks. 

Where it fits: Localizing a lesson, technical manual or knowledge-base article while keeping headings, lists and tables usable. A 32K context window still limits document length; it does not mean an unlimited PDF can be translated in one request. 

What still needs work: Direct translation between two Indian languages is on the roadmap. The release describes the current path as translation through English. It also identifies sentence-level English-to-Indic fluency as an area for improvement. 

3. Indic-Transcribe: Choose Accuracy or Script Flexibility

Indic-Transcribe

Indic-Transcribe is a family of two 1.2B-parameter ASR models. Its coverage includes the 22 scheduled Indian languages, English, Bhili and Bhojpuri, with Flex also listing Haryanvi and Chhattisgarhi. 

Core prioritizes accurate native-script transcripts. Flex offers native, Romanized and mixed-script output. Mixed mode keeps native words in their script while allowing English terms and numerals in Latin characters. 

The release chart reports 8.7 OIWER for Core and 11.1 for Flex on Voice of India, covering 15 languages. OIWER accepts documented spelling and transliteration variants, reducing penalties for valid alternative spellings. 

The Hugging Face card lists a slightly different Flex average, 11.3. The figure above reproduces the release blog’s evaluation; its values should not be mixed with the model-card comparison. 

Underneath, both use a Canary-derived FastConformer encoder and a newly trained 24-layer Transformer decoder. Bodhan reports training on 1.3 million hours of audio, combining weak supervision, synthetic speech and human-labelled data. 

Where it fits: Transcribing recorded lessons, interviews or regional-language voice notes. Choose Core when native-script accuracy matters most, and Flex when transcript format is part of the product requirement. 

What still needs work: Audio is processed in windows of up to 30 seconds. Longer recordings need chunking. Real-time streaming, speaker diarization and overlapping-speaker separation are listed as future work in the release. 

4. Indic-Speak: Read Mixed-Language Text Aloud

Indic-Speak

Indic-Speak generates speech across 22 Indian languages and 12 scripts, with 45 voices. It accepts native and Latin scripts within the same sentence without requiring a language tag for every span. 

The roughly 3.36B-parameter stack uses a Llama-3.2-3B backbone extended with audio tokens, followed by a vocoder. A normalizer converts notation, numbers and dates into spoken forms before generation. 

Bodhan evaluated 30,000 readings from 15,000 code-mixed sentences across 10 languages. An ASR system transcribed the audio, then an LLM judge assessed content fidelity. About 93% reached the highest scoring band; 0.7% scored two or below out of five. 

This measures whether the generated audio preserves the content. It is not a human preference score for naturalness. Human listening comparisons were still in progress, and the other 12 supported languages did not yet have equivalent scored evidence. 

Where it fits: Regional-language narration, accessible learning material and support responses containing English terms. Each voice can read different languages, but its original accent carries over. Start with a recommended native voice when that matters. 

What still needs work: Quality varies by voice, and some generations repeat or omit content. The 5:36 audiobook example on the release page joins six separately generated paragraphs; it is not a single uninterrupted generation. 

For an original test, try: “Kal ka science test 9:30 AM par hai. Chapter 4 revise kar lena.” Then compare a Romanized and native-script version for pronunciation, numbers and pauses. This is a suggested test input, not a measured result. 

How to Access the Four Models

Use the Bodhan API console for hosted access, or the Hugging Face weights linked below for local deployment. The hosted APIs use OpenAI-compatible request shapes with the base URL https://api.bodhan.ai/v1. Keys are issued per model. 

Model / weights Hosted price
IndicOCR ₹0.20 per image
Indic-Translate ₹0.20 per 10,000 output tokens
Indic-Transcribe ₹0.10 per input audio minute
Indic-Speak ₹6 per 10,000 input characters

Weights: IndicOCR · Translate · Transcribe Core / Flex · Speak

New accounts are listed with ₹10 credit. Not much but considering the cost, it would be sufficient to do some tests.  .

The hosted API documentation has narrower operating guidance than some model demonstrations: transcription requests accept up to 30 seconds, and speech generation recommends short inputs. The speech API also requires a language setting, even though the model does not need per-span language tags.

What Can You Build With Them?

One possible classroom workflow is to extract a scanned lesson with IndicOCR, translate the verified text with Indic-Translate, and narrate it with Indic-Speak. Indic-Transcribe can turn a teacher’s recorded explanation into searchable notes. These are proposed integrations, not a prebuilt four-model application.

For the document side, OCR tutorial with Tesseract, OpenCV and Python is a useful starting point. 

Conclusion

Bodhan’s releases give developers four focused tools for Indian-language documents and audio. Their value will depend on the languages, scripts and input quality a project encounters. Start with one representative page or recording, inspect the output, and expand once the results hold up. 

Frequently Asked Questions

Q1. Are all four models one system? 

A. No. They are separate models for OCR, translation, transcription and speech generation. Developers can connect them in an application. 

Q2. Does IndicOCR support handwriting in all 22 languages? 

A. No. Handwriting currently covers 12 Indian languages plus English. Printed-text coverage spans all 22 Indian languages plus English. 

Q3. Which Indic-Transcribe model should I use? 

A. Start with Core for native-script accuracy. Choose Flex when you need Romanized or mixed-script output. 

Studying, evaluating, and explaining AI systems for over 6 years.

“𝘖𝘯𝘤𝘦 𝘮𝘦𝘯 𝘵𝘶𝘳𝘯𝘦𝘥 𝘵𝘩𝘦𝘪𝘳 𝘵𝘩𝘪𝘯𝘬𝘪𝘯𝘨 𝘰𝘷𝘦𝘳 𝘵𝘰 𝘮𝘢𝘤𝘩𝘪𝘯𝘦𝘴 𝘪𝘯 𝘵𝘩𝘦 𝘩𝘰𝘱𝘦 𝘵𝘩𝘢𝘵 𝘵𝘩𝘪𝘴 𝘸𝘰𝘶𝘭𝘥 𝘴𝘦𝘵 𝘵𝘩𝘦𝘮 𝘧𝘳𝘦𝘦. 𝘉𝘶𝘵 𝘵𝘩𝘢𝘵 𝘰𝘯𝘭𝘺 𝘱𝘦𝘳𝘮𝘪𝘵𝘵𝘦𝘥 𝘰𝘵𝘩𝘦𝘳 𝘮𝘦𝘯 𝘸𝘪𝘵𝘩 𝘮𝘢𝘤𝘩𝘪𝘯𝘦𝘴 𝘵𝘰 𝘦𝘯𝘴𝘭𝘢𝘷𝘦 𝘵𝘩𝘦𝘮.” — 𝖥𝗋𝖺𝗇𝗄 𝖧𝖾𝗋𝖻𝖾𝗋𝗍, 𝖣𝗎𝗇𝖾

Login to continue reading and enjoy expert-curated content.