Indic speech recognition accuracy: what is actually published
Word error rates for Hindi and other Indian languages, compiled from the papers and vendor pages that published them. Each figure says which dataset, which metric, which audio condition and which model version it came from — because almost none of them are comparable to each other.
- 15.0%
- Lowest published Hindi WER on the IndicVoices benchmark — AI4Bharat's own IndicASR, 2024
- 3.5×
- Same model, same language: Hindi error rate on telephone audio versus studio-read audio (Vistaar, 2023)
- 0
- Indian languages in the Artificial Analysis word-error-rate index, the most-cited independent ASR leaderboard
- 0
- Accuracy figures in Bhashini's public model catalogue, which lists models, languages and service IDs only
If you are choosing a speech stack for a citizen-facing line in an Indian language, the published record is thinner than it looks. There is a lot of it, and very little of it answers the question a department actually has, which is: how often will this system mis-hear the people who call my helpline.
This page is a compilation, not a measurement. We ran no evaluation for it and we publish no accuracy figure of our own. Every number below comes from a paper or a vendor page that we fetched, and every row states what that source was measuring. Where a vendor publishes nothing for a language, the cell says so — that absence is one of the more useful findings here.
It is the accuracy companion to a latency compilation written on the same principle: gather what exists, label what each figure measures, refuse to add up numbers that describe different things, and hand over a method for producing your own.
Read the metric before you read the number
Six things separate two error rates that look like they belong in the same column.
WER versus CER
Word error rate counts substituted, deleted and inserted words. Character error rate counts characters. For Indian languages the gap between them is wide, because Devanagari, Bengali and Dravidian orthographies pack more meaning per word and a single wrong vowel sign costs a whole word in WER and one character in CER. Meta's MMS reports Hindi at 10.6 WER and 4.7 CER on the same benchmark family. Both are correct. Only one of them is comparable to a WER quoted elsewhere.
Read speech versus spontaneous speech
IndicTTS is studio-recorded read sentences. Kathbath is read speech over a phone app. Gramvaani is telephone-quality spontaneous speech from a rural helpline. IndicVoices is 74% extempore and 17% conversation. The same model scores 7.6 and 26.8 on two of these, in Hindi.
Clean versus noisy splits
IndicSUPERB ships explicit clean and noisy test splits for the same speakers and sentences, which is the cleanest published isolation of the noise variable in an Indian-language benchmark. IndicWav2Vec with a language model goes from 8.3 to 9.2 WER in Hindi between the clean and noisy unknown-speaker splits, and from 12.5 to 15.0 in Tamil.
Text normalisation
Whisper's published numbers use Whisper's own text normaliser; Meta re-scored Whisper with the same normaliser so its comparison would hold. Two teams scoring the same transcript with different normalisation rules for numerals, punctuation and transliteration will report different WERs for identical audio. When a paper does not say which normaliser it used, its number is not comparable to one that does.
Model version and date
"Whisper" is six models. Large-v2 scores 21.5 on FLEURS Hindi; medium scores 26.8; small scores 38.4; base scores 101.1. A benchmark row that says only "Whisper" has told you very little, and several published comparisons do exactly that.
Why some figures exceed 100%
WER can pass 100 when a model inserts more words than the reference contains — typically when it has effectively no coverage of the language and starts hallucinating. Whisper large-v2 scores 102.7 on FLEURS Gujarati and 156.5 on FLEURS Sindhi. Those are not bad scores on a hard language. They mean the output is unusable.
Where the gaps are, and who left them
Four categories of missing number, each with a different consequence for a procurement decision.
The commercial STT vendors publish no per-language Indic figure
Deepgram, AssemblyAI, Google Cloud and OpenAI all list Hindi in their supported-language tables. None of the four publishes a word error rate for it. AssemblyAI goes furthest, grouping languages into accuracy tiers measured by WER, but publishes the tier rather than the number. A supported-language list is a statement about routing, not about accuracy.
The open multilingual models publish averages, not languages
Meta's SeamlessM4T reports a 45% average WER reduction against Whisper across 77 FLEURS languages and does not break that out per language. Google's USM reports averages across FLEURS-102. An average over 77 languages tells a department in Odisha nothing about Odia, and the per-language tables that do exist show exactly why: the spread inside those averages runs from single digits to over 100.
Bhashini publishes a catalogue, not a scorecard
The Bhashini API documentation lists ASR service IDs, the languages each one covers and the institution that built it — IIT Madras, AI4Bharat and others. It carries no accuracy figure of any kind. The same is true of the AI4Bharat IndicConformer model card on Hugging Face, which is the model behind several of those service IDs: architecture, parameter count, 22 languages, no error rate. The numbers for those models exist, but they live in the research papers, not in the catalogue a department procures from.
The independent leaderboard does not cover India
Artificial Analysis publishes the most widely cited independent word-error-rate index for speech-to-text. Its v2 index is built from three datasets — AA-AgentTalk, VoxPopuli-Cleaned and Earnings22-Cleaned — all English. There is no Indian language in it. So for Indic ASR there is no neutral scoreboard at all: every figure is either an academic benchmark run by the team that built one of the models, or a vendor scoring itself.
Why a phone call is harder than any of these benchmarks
Four effects that move the number, each with something published behind it. None of the figures in this section is ours.
Every figure in the table above was measured on a file. A citizen helpline measures on a call, and four things change between the two.
Narrowband audio. Telephony carries roughly 300–3,400 Hz at 8 kHz sampling, while every model in the table expects 16 kHz — the IndicConformer model card resamples to 16 kHz as a required step. The published evidence for what this costs in Indian languages is the Gramvaani row: telephone-quality spontaneous Hindi collected through a farmer-facing helpline platform. On Vistaar, IndicWhisper scores 7.6 on studio-read Hindi and 26.8 on Gramvaani; Google STT scores 18.3 and 59.9 on the same pair. Same language, same models, different channel.
Code-switching. The 2021 MUCS challenge built test sets specifically for Hindi-English and Bengali-English code-switched speech. Its three baseline systems land between 24.7 and 31.6 WER on Hindi-English and between 30.3 and 37.0 on Bengali-English, on conversational audio — well above where the same era's monolingual Hindi systems sit on read benchmarks, though those are different corpora and the pair is not a controlled comparison. The organisers also report transliteration-WER separately, because a system can hear an English word correctly and write it in the wrong script: on the Hindi-English blind set the best baseline scores 24.7 WER and 22.7 transliteration-WER.
Numbers and named entities. A helpline call is mostly numbers and names: a scheme name, an account reference, a village, a date. The IndicContextEval benchmark measures this as a separate metric — named entity error rate — across eight Indic languages, and reports NEER running well above overall WER for every model tested: 25.9 NEER against 16.9 WER for the best system, and 35.6 against 28.6 for GPT-4o Transcribe. A system with an acceptable WER can still be unusable on the part of the sentence that carries the transaction.
Dialect and register. IndicVoices was collected across 145 districts precisely because district-level variation is the thing a national benchmark averages away, and Gramvaani was built around regional and dialectal variation in Hindi. Neither dataset lets you predict how a model behaves on your callers' dialect. It only establishes that the variation is large enough that two teams thought it worth building a corpus around.
Meta's SeamlessM4T paper also publishes WER-against-signal-to-noise curves showing error rates rising as input noise rises, over four languages. It is a chart rather than a table, so there is no figure to quote, but the direction is published and it is the direction you would expect.
How a department should evaluate this for itself
Nothing in the published record substitutes for an evaluation on your own calls. This is the shape one should take.
- 1
Sample from your own helpline, not from a benchmark
Take held-out recordings from the line you intend to automate, across the hours it actually runs, on the telephony path it actually uses. A benchmark file tells you how a model handles that benchmark. Your recordings carry your codec, your line noise, your callers' dialects and your scheme vocabulary, and none of those four are in any published number.
- 2
Transcribe to a written protocol, with two passes
Decide before you start how numerals are written, how English words inside an Indian-language sentence are scripted, how hesitations and false starts are handled, and how unintelligible audio is marked. Have a second transcriber re-do a share of the set blind and measure how far the two disagree. That disagreement is the floor of your evaluation: no model can be measured as more accurate than your reference is consistent. The IndicVoices collection ran a published transcription guideline and a quality-control pass for exactly this reason.
- 3
Break the result down by language and by dialect region
Report every language separately, and within the main language report by district or region where the sample allows. A single pooled number hides the case you are about to fail: the model that scores acceptably overall and badly in the two districts that generate a third of the calls. The published per-language tables show spreads of 20 to 80 points inside a single model, so assume your dialect spread is not zero until you have measured it.
- 4
Score numbers and entities as their own metric
Extract the scheme names, amounts, dates, phone numbers and place names from each reference, and score those separately from the running text. IndicContextEval's named entity error rate is the published precedent for doing this, and the reason to copy it is that entity errors and word errors move independently — a system can improve on one while getting worse on the other.
- 5
Report CER alongside WER
For Indian-language scripts, WER alone overstates the damage from orthographic near-misses and understates nothing. Reporting both tells you whether a model is mis-hearing the word or mis-spelling it, and those have different fixes: the first needs a different model, the second often needs normalisation or a lexicon.
- 6
Size the sample and publish the interval
Enough calls per language that the number stops moving when you add more, and enough per dialect region that the per-region breakdown is not noise. Report the sample size in calls and in minutes of audio alongside every figure, and put a confidence interval on it — a bootstrap over utterances is the standard way and takes minutes. A point estimate with no interval and no sample size is not a result anyone can act on, including you in six months.
- 7
Re-run it when anything changes
A model version, a telephony provider, a new district, a new scheme vocabulary — each one invalidates the measurement. Keep the evaluation set and the scoring script in version control so a re-run costs an afternoon rather than a new project.
Every per-language figure we could find, with what it measures
Sorted by source. Figures in the same column are not comparable across rows unless the dataset, metric and condition all match. A cell reading "no published figure" means we fetched the source and it publishes none.
| Model · language · dataset | Metric and audio condition | Figure | Source · date |
|---|---|---|---|
| IndicASR / IndicConformer · Hindi · IndicVoices | Unlabelled in the paper; read as WER. Read, extempore and conversational speech, phone-recorded across 145 districts | 15.0 | AI4Bharat · Mar 2024 |
| IndicASR · Bengali · IndicVoices | As above | 15.9 | AI4Bharat · Mar 2024 |
| IndicASR · Punjabi · IndicVoices | As above | 12.9 | AI4Bharat · Mar 2024 |
| IndicASR · Urdu · IndicVoices | As above | 14.4 | AI4Bharat · Mar 2024 |
| IndicASR · Marathi · IndicVoices | As above | 18.2 | AI4Bharat · Mar 2024 |
| IndicASR · Tamil · IndicVoices | As above | 31.2 | AI4Bharat · Mar 2024 |
| IndicASR · Telugu · IndicVoices | As above | 26.8 | AI4Bharat · Mar 2024 |
| IndicASR · Kannada · IndicVoices | As above | 30.3 | AI4Bharat · Mar 2024 |
| IndicASR · Malayalam · IndicVoices | As above | 40.5 | AI4Bharat · Mar 2024 |
| IndicASR · Kashmiri · IndicVoices | As above | 39.3 | AI4Bharat · Mar 2024 |
| IndicASR · Santali · IndicVoices | As above | 35.4 | AI4Bharat · Mar 2024 |
| Google USM · Hindi · IndicVoices | As above. Model version not stated by the evaluating paper | 20.5 | AI4Bharat · Mar 2024 |
| Microsoft Azure STT · Hindi · IndicVoices | As above. Version not stated | 27.2 | AI4Bharat · Mar 2024 |
| OpenAI Whisper · Hindi · IndicVoices | As above. Version not stated | 34.0 | AI4Bharat · Mar 2024 |
| Meta MMS · Hindi · IndicVoices | As above. Version not stated | 38.9 | AI4Bharat · Mar 2024 |
| Google USM · Tamil · IndicVoices | As above | 58.9 | AI4Bharat · Mar 2024 |
| OpenAI Whisper · Telugu · IndicVoices | As above — output unusable | 151.9 | AI4Bharat · Mar 2024 |
| IndicWhisper · Hindi · IndicTTS | WER. Studio-recorded read speech | 7.6 | AI4Bharat · May 2023 |
| IndicWhisper · Hindi · Kathbath | WER. Read speech captured through a phone app | 10.3 | AI4Bharat · May 2023 |
| IndicWhisper · Hindi · FLEURS | WER. Clean read speech | 11.4 | AI4Bharat · May 2023 |
| IndicWhisper · Hindi · MUCS | WER. Read speech | 12.0 | AI4Bharat · May 2023 |
| IndicWhisper · Hindi · Common Voice | WER. Crowd-recorded read speech | 15.0 | AI4Bharat · May 2023 |
| IndicWhisper · Hindi · Gramvaani | WER. Telephone-quality spontaneous speech | 26.8 | AI4Bharat · May 2023 |
| Google STT · Hindi · Gramvaani | WER. Telephone-quality spontaneous speech | 59.9 | AI4Bharat · May 2023 |
| Azure STT · Hindi · Gramvaani | WER. Telephone-quality spontaneous speech | 42.3 | AI4Bharat · May 2023 |
| IndicWav2Vec · Hindi · Gramvaani | WER. Telephone-quality spontaneous speech | 42.1 | AI4Bharat · May 2023 |
| Google STT · Hindi · IndicTTS | WER. Studio-recorded read speech, same model as the Gramvaani row above | 18.3 | AI4Bharat · May 2023 |
| IndicWav2Vec + LM · Hindi · Kathbath clean | WER. Unknown-speaker split, clean audio | 8.3 | AI4Bharat · Aug 2022 |
| IndicWav2Vec + LM · Hindi · Kathbath noisy | WER. Same speakers and sentences, noisy split | 9.2 | AI4Bharat · Aug 2022 |
| IndicWav2Vec + LM · Tamil · Kathbath clean | WER. Unknown-speaker split, clean audio | 12.5 | AI4Bharat · Aug 2022 |
| IndicWav2Vec + LM · Tamil · Kathbath noisy | WER. Same speakers and sentences, noisy split | 15.0 | AI4Bharat · Aug 2022 |
| IndicWav2Vec + LM · Malayalam · Kathbath clean | WER. Unknown-speaker split, clean audio | 24.5 | AI4Bharat · Aug 2022 |
| IndicWav2Vec + LM · Malayalam · Kathbath noisy | WER. Same speakers and sentences, noisy split | 25.9 | AI4Bharat · Aug 2022 |
| Whisper large-v2 · Hindi · FLEURS | WER. Clean read speech, Whisper text normaliser | 21.5 | OpenAI · Dec 2022 |
| Whisper large-v2 · Tamil · FLEURS | WER. Clean read speech | 17.5 | OpenAI · Dec 2022 |
| Whisper large-v2 · Urdu · FLEURS | WER. Clean read speech | 22.6 | OpenAI · Dec 2022 |
| Whisper large-v2 · Kannada · FLEURS | WER. Clean read speech | 37.0 | OpenAI · Dec 2022 |
| Whisper large-v2 · Marathi · FLEURS | WER. Clean read speech | 38.3 | OpenAI · Dec 2022 |
| Whisper large-v2 · Nepali · FLEURS | WER. Clean read speech | 47.1 | OpenAI · Dec 2022 |
| Whisper large-v2 · Telugu · FLEURS | WER. Clean read speech — output unusable | 99.0 | OpenAI · Dec 2022 |
| Whisper large-v2 · Malayalam · FLEURS | WER. Clean read speech — output unusable | 100.7 | OpenAI · Dec 2022 |
| Whisper large-v2 · Punjabi · FLEURS | WER. Clean read speech — output unusable | 102.4 | OpenAI · Dec 2022 |
| Whisper large-v2 · Gujarati · FLEURS | WER. Clean read speech — output unusable | 102.7 | OpenAI · Dec 2022 |
| Whisper large-v2 · Sindhi · FLEURS | WER. Clean read speech — output unusable | 156.5 | OpenAI · Dec 2022 |
| Whisper medium · Hindi · FLEURS | WER. Clean read speech | 26.8 | OpenAI · Dec 2022 |
| Whisper small · Hindi · FLEURS | WER. Clean read speech | 38.4 | OpenAI · Dec 2022 |
| Whisper base · Hindi · FLEURS | WER. Clean read speech — output unusable | 101.1 | OpenAI · Dec 2022 |
| Whisper large-v2 · Hindi · Common Voice 9 | WER. Crowd-recorded read speech | 21.9 | OpenAI · Dec 2022 |
| Whisper large-v2 · Tamil · Common Voice 9 | WER. Crowd-recorded read speech | 16.1 | OpenAI · Dec 2022 |
| Whisper large-v2 · Urdu · Common Voice 9 | WER. Crowd-recorded read speech | 24.2 | OpenAI · Dec 2022 |
| Whisper large-v2 · Malayalam · Common Voice 9 | WER. Crowd-recorded read speech — output unusable | 103.2 | OpenAI · Dec 2022 |
| MMS-1107 LSAH + LM · Hindi · FLEURS | WER. Clean read speech, Whisper text normaliser | 10.6 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Bengali · FLEURS | WER. Clean read speech | 12.1 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Gujarati · FLEURS | WER. Clean read speech | 12.8 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Kannada · FLEURS | WER. Clean read speech | 13.3 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Marathi · FLEURS | WER. Clean read speech | 13.4 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Telugu · FLEURS | WER. Clean read speech | 13.6 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Malayalam · FLEURS | WER. Clean read speech | 15.3 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Tamil · FLEURS | WER. Clean read speech | 16.3 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Assamese · FLEURS | WER. Clean read speech | 19.2 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Punjabi · FLEURS | WER. Clean read speech | 19.8 | Meta AI · May 2023 |
| MMS-1107 LSAH + LM · Urdu · FLEURS | WER. Clean read speech | 20.5 | Meta AI · May 2023 |
| MMS FL-102 · Hindi · FLEURS | CER, not WER — not comparable to any row above | 4.7 | Meta AI · May 2023 |
| MMS FL-102 · Gujarati · FLEURS | CER, not WER | 5.1 | Meta AI · May 2023 |
| MMS FL-102 · Bengali · FLEURS | CER, not WER | 5.3 | Meta AI · May 2023 |
| Organisers' baselines, best of three · Hindi-English · MUCS code-switching test | WER. Code-switched conversational speech, realigned segments. End-to-end Transformer baseline; the GMM-HMM and TDNN baselines score 31.6 and 28.4 | 25.9 | MUCS challenge · Apr 2021 |
| Organisers' baselines, best of three · Hindi-English · MUCS blind test | WER. Same audio type, blind set. GMM-HMM baseline; the TDNN and Transformer baselines score 29.0 and 31.2 | 24.7 | MUCS challenge · Apr 2021 |
| Organisers' baselines, best of three · Bengali-English · MUCS blind test | WER. Code-switched conversational speech, blind set. GMM-HMM baseline | 32.4 | MUCS challenge · Apr 2021 |
| Organisers' baselines, best of three · Bengali-English · MUCS code-switching test | WER. Code-switched conversational speech. TDNN baseline | 30.3 | MUCS challenge · Apr 2021 |
| GMM-HMM baseline · Hindi-English · MUCS blind test, transliteration-scored | Transliteration-WER after script normalisation, against 24.7 WER on the same audio | 22.7 | MUCS challenge · Apr 2021 |
| Organisers' monolingual baseline · Hindi · MUCS blind test | WER. Monolingual read speech. A different corpus from the code-switched rows above, so not a controlled comparison with them | 27.5 | MUCS challenge · Apr 2021 |
| Sarvam Audio · Hindi · IndicContextEval | WER. Natural audio across 23 domains, entity list supplied in native script | 12.4 | AI4Bharat · Interspeech 2026 |
| Gemini 3 Flash · Hindi · IndicContextEval | WER. Same setting | 14.3 | AI4Bharat · Interspeech 2026 |
| GPT-4o Transcribe · Hindi · IndicContextEval | WER. Same setting | 17.5 | AI4Bharat · Interspeech 2026 |
| Sarvam Audio · Malayalam · IndicContextEval | WER. Same setting | 30.8 | AI4Bharat · Interspeech 2026 |
| GPT-4o Transcribe · Malayalam · IndicContextEval | WER. Same setting | 42.6 | AI4Bharat · Interspeech 2026 |
| IndicConformer · 8 Indic languages pooled · IndicContextEval | Named entity error rate, a different metric from WER | 29.6 | AI4Bharat · Interspeech 2026 |
| Sarvam Audio · 8 Indic languages pooled · IndicContextEval | Named entity error rate, against 16.9 WER on the same audio | 25.9 | AI4Bharat · Interspeech 2026 |
| GPT-4o Transcribe · 8 Indic languages pooled · IndicContextEval | Named entity error rate, against 28.6 WER on the same audio | 35.6 | AI4Bharat · Interspeech 2026 |
| Sarvam Saaras v2.5 · 11 languages pooled · IndicVoices | WER. Vendor-run; no per-language breakdown published | ~22 | Sarvam AI · stated Feb 2026 for a Jun 2025 model |
| Sarvam Saaras v3 · 10 most-spoken languages pooled · IndicVoices | WER. Vendor-run; no per-language breakdown published | 19.31 | Sarvam AI · Feb 2026 |
| Sarvam Saaras v3 · 22 languages pooled · IndicVoices | WER. Vendor-run | ~19 | Sarvam AI · Feb 2026 |
| Sarvam Saaras v4 · 8 languages pooled · IndicContextEval | WER. Vendor-run on a third-party benchmark, entity list in native script | 16.03 | Sarvam AI · Aug 2026 |
| Sarvam Saaras v4 · per language · Vistaar | Published as an interactive chart; no tabulated value in the page source | no published figure | Sarvam AI · Aug 2026 |
| Deepgram Nova-3 · Hindi | Hindi is listed as supported; accuracy is claimed only as a relative reduction in English word error rate | no published figure | Deepgram docs · fetched Sep 2026 |
| AssemblyAI Universal-2 · Hindi | Languages are grouped into accuracy tiers defined by WER; the tier is published, the rate is not | no published figure | AssemblyAI docs · fetched Sep 2026 |
| Google Chirp 3 · Hindi | hi-IN is documented as supported; the model page carries no error rate for any language | no published figure | Google Cloud docs · fetched Sep 2026 |
| OpenAI gpt-4o-transcribe · Hindi | Hindi appears in the supported-language list; the API docs carry no accuracy figure | no published figure | OpenAI docs · fetched Sep 2026 |
| ElevenLabs Scribe · Hindi | The launch post claims the lowest WER across 99 languages on FLEURS and Common Voice and names figures for Italian only | no published figure | ElevenLabs · fetched Sep 2026 |
| Meta SeamlessM4T-Large · any Indic language | Reports a 45% average WER reduction over Whisper across 77 FLEURS languages; no per-language table | no published figure | Meta AI · Aug 2023 |
| Google USM · any Indic language | Reports averages across FLEURS-102; no per-language Indic table in the paper | no published figure | Google · Mar 2023 |
| Bhashini ASR services · all 22 languages | The catalogue lists service IDs, languages and the institution that built each model | no published figure | Bhashini API docs · fetched Sep 2026 |
| AI4Bharat IndicConformer 600M · all 22 languages | The model card gives architecture, parameter count and language list; the figures live in the papers above | no published figure | Hugging Face model card · fetched Sep 2026 |
| Any model · any Indian language · Artificial Analysis WER index | The index is built from three English datasets; no Indian language appears in it | no published figure | Artificial Analysis · fetched Sep 2026 |
One caveat that applies to the whole first block: the IndicVoices paper labels its comparison table "Performance" and never states the metric, the normalisation or the versions of USM, Azure, Whisper and MMS it tested. Everyone who cites it, Sarvam included, reads it as WER, and the neighbouring figures are consistent with that. We have reproduced it as published and flagged it rather than quietly calling it WER. Treat the four non-AI4Bharat rows in that block as the weakest-sourced numbers on this page.
The sources, and what class of evidence each one is
Every figure above traces to one of these. They fall into three classes that carry different weight: peer-reviewed papers, vendor-published claims, and independent leaderboards. For Indic ASR the third class is effectively empty, which is why the second class deserves more scrutiny here than it would in English.
Read this first. None of these figures is an AiSewak result. We have published no accuracy number of our own and nothing on this page should be read as one. Each entry below is another organisation's measurement, linked to the page we read it from. This page is a technical reference, not procurement or legal advice — a department should treat it as a reading list, not as a basis for an award.
- 15.0Measured result
IndicVoices: 7,348 hours across 22 languages and 145 districts, 74% extempore and 17% conversational, with a public train/test split and a per-language comparison of IndicASR against USM, Azure, Whisper and MMS.
Peer-reviewed. The broadest Indian-language ASR benchmark that exists, and the only one covering all 22 scheduled languages. Its comparison table does not state the metric or the competitor model versions — read the note under the table above before quoting it.
AI4Bharat, IIT Madras · IndicVoices (arXiv 2403.01926) — Hindi, IndicVoices test set · 2024
- 7.6 → 26.8Measured result
Vistaar: 59 benchmarks over 12 Indian languages, spanning studio-read, crowd-read, phone-app and telephone-quality spontaneous speech, with the same models scored across all of them.
Peer-reviewed. The single most useful source on this page for a helpline decision, because it holds the model and the language fixed and varies the recording channel. Gramvaani is described in the paper as telephone-quality speech with a focus on regional and dialectal variation in Hindi.
AI4Bharat · Vistaar (arXiv 2305.15386) — IndicWhisper, Hindi, IndicTTS versus Gramvaani · 2023
- 8.3 → 9.2Measured result
IndicSUPERB and the Kathbath corpus: matched clean and noisy test splits across 12 languages, for the same speakers and sentences.
Peer-reviewed. The cleanest published isolation of added noise as a variable in an Indian-language benchmark, because everything except the noise is held constant.
AI4Bharat · IndicSUPERB (arXiv 2208.11761) — IndicWav2Vec with LM, Hindi, clean versus noisy unknown-speaker split · 2022
- 10.5Measured result
IndicWav2Vec: self-supervised pretraining across 40 Indian languages, with published WER on MUCS, MSR and OpenSLR benchmarks and a comparison against the 2021 MUCS leaderboard.
Peer-reviewed. Useful mainly as the earlier baseline the later Indic models are measured against; its Hindi MUCS figure is read speech, not conversational.
AI4Bharat · IndicWav2Vec (arXiv 2111.03945) — Hindi, MUCS benchmark, best configuration · 2021
- 21.5Measured result
Whisper: per-language WER appendix tables for FLEURS, Common Voice 9, MLS and VoxPopuli, for all six model sizes.
Peer-reviewed, and the most complete per-language disclosure any of the commercial labs has made. It is also the source that shows the failure mode most clearly: the same model scores 21.5 on Hindi and above 100 on Gujarati, Telugu, Malayalam, Punjabi and Sindhi.
OpenAI · Robust Speech Recognition via Large-Scale Weak Supervision (arXiv 2212.04356) — large-v2, Hindi, FLEURS · 2022
- 10.6Measured result
MMS: per-language FLEURS WER for 54 languages against Whisper medium and large-v2, plus per-language CER for 102 languages.
Peer-reviewed, but note that Meta is reporting its own model against a competitor it re-scored itself. It states the normaliser it used, which is what makes the comparison readable at all.
Meta AI · Scaling Speech Technology to 1,000+ Languages (arXiv 2305.13516) — MMS-1107 LSAH with LM, Hindi, FLEURS test · 2023
- no per-language Indic figureOfficial figure
Google USM: a 2-billion-parameter multilingual speech model reporting averaged WER and CER across FLEURS-102 and YouTube long-form sets.
Peer-reviewed, but everything Indic in it is inside a 102-language average. The only per-language Hindi number for USM anywhere in this compilation came from AI4Bharat evaluating it, not from Google.
Google · Google USM (arXiv 2303.01037) · 2023
- no per-language Indic figureOfficial figure
SeamlessM4T: a 45% average WER reduction against Whisper across the 77 FLEURS languages both models support, and WER-versus-signal-to-noise curves over four languages.
Peer-reviewed. Reports ASR quality only as averages over language groups. Its section on input noise is the clearest published statement that error rates rise as signal-to-noise falls, but it is a chart, so there is no figure to cite.
Meta AI · SeamlessM4T (arXiv 2308.11596) · 2023
- 25.9Measured result
MUCS 2021: separate multilingual and code-switching ASR tracks, with Hindi-English and Bengali-English test and blind sets, and transliteration-WER reported alongside WER.
Peer-reviewed. The only published Indian-language benchmark built specifically around code-switching, which is how most urban helpline callers actually speak. Its data table also records the channel compression used per language — 3GP, M4A or PCM — which is rare and worth copying.
Interspeech 2021 special session · Multilingual and code-switching ASR challenges for low resource Indian languages (arXiv 2104.00235) — best of three organiser baselines, Hindi-English code-switched test set · 2021
- 25.9 NEER vs 16.9 WERMeasured result
IndicContextEval: 56 hours of natural audio across 23 domains and 8 Indic languages, scoring named entity error rate separately from word error rate at seven levels of supplied context.
Peer-reviewed. The published precedent for scoring entities as their own metric. For every model tested, entity error runs above overall word error — the gap is 9 points for the best system and 7 for GPT-4o Transcribe.
AI4Bharat · IndicContextEval (arXiv 2606.19157), Interspeech 2026 — best model, 8 languages pooled · 2026
- 19.31Official figure
Saaras v3: ~19% WER on the IndicVoices benchmark across 22 languages, and 19.31% on the ten most-spoken subset, compared against GPT-4o Transcribe, Gemini 3 Pro, Deepgram Nova-3 and Scribe v2.
Vendor-published: Sarvam running a third-party benchmark on its own model against named competitors. The benchmark is independent and the number is specific, which puts it well above a marketing claim; the scoring was still done by the party being scored, and no per-language breakdown is given.
Sarvam AI · Introducing Saaras V3 — vendor-published, ten-language subset of IndicVoices · 2026
- 16.03Official figure
Saaras v4: 16.03% WER on IndicContextEval with a native-script entity list supplied, 5.22% language-identification error across 22 languages, and per-dataset Vistaar results across ten languages.
Vendor-published. The English benchmark table is given as numbers; the Indic per-language results are rendered as an interactive chart with no tabulated values in the page, so they cannot be quoted or checked from the published page.
Sarvam AI · Introducing Saaras V4 — vendor-published, on the AI4Bharat IndicContextEval benchmark · 2026
- no published figureOfficial figure
The Bhashini API catalogue lists every ASR service ID, the languages it covers and the institution behind it — IIT Madras, AI4Bharat and others — across the 22 scheduled languages plus regional variants.
The national platform's own catalogue carries no accuracy figure of any kind. For a department procuring through Bhashini, the numbers behind those service IDs are in the AI4Bharat papers above, not in the catalogue itself.
Bhashini (MeitY) · Bhashini APIs — available models for usage · fetched Sep 2026
- no published figureOfficial figure
The IndicConformer 600M model card documents a hybrid CTC and RNN-T conformer across 22 languages, with a required resample to 16 kHz before inference.
The most widely deployed open Indic ASR model publishes no error rate on its own card. The 16 kHz requirement is the line that matters for telephony: 8 kHz call audio is upsampled to meet it, which adds no information back.
AI4Bharat · IndicConformer 600M multilingual — Hugging Face model card · fetched Sep 2026
- no published figureOfficial figure
Deepgram lists Hindi for Nova-3 and Flux and Gujarati and Kannada for Nova-2, and quantifies accuracy only as a relative reduction in English word error rate.
Vendor documentation. Language support is documented precisely; per-language accuracy is not documented at all.
Deepgram · Models and languages overview · fetched Sep 2026
- no published figureOfficial figure
AssemblyAI groups its Universal-2 languages into accuracy tiers explicitly defined by word error rate, and publishes the tier rather than the rate.
Vendor documentation. A tier is more than nothing — it is an ordering — but it cannot be compared to any figure in the table above.
AssemblyAI · Supported languages · fetched Sep 2026
- no published figureOfficial figure
Google's Chirp 3 model page documents hi-IN support and streaming recognition, with no error rate for any language.
Vendor documentation. The only Google per-language Indic figures in this compilation were produced by AI4Bharat evaluating Google's own services.
Google Cloud · Chirp 3 model documentation · fetched Sep 2026
- no published figureOfficial figure
OpenAI's speech-to-text guide lists Hindi, Tamil, Kannada, Marathi, Nepali and Urdu among supported languages for its transcription models, with no accuracy figure.
Vendor documentation. OpenAI's per-language Hindi numbers exist, but in the 2022 Whisper paper rather than in the docs for the models currently on sale.
OpenAI · Speech-to-text API guide · fetched Sep 2026
- no published figureOfficial figure
The Scribe launch post claims the lowest transcription word error rate across FLEURS and Common Voice benchmark tests in 99 languages, naming figures for Italian but none for any Indian language.
Vendor claim. It names the benchmarks it ran, which is more than most, and then publishes no per-language table for them.
ElevenLabs · Meet Scribe · fetched Sep 2026
- 0 Indian languagesMeasured result
The Artificial Analysis speech-to-text word-error-rate index is built from AA-AgentTalk, VoxPopuli-Cleaned and Earnings22-Cleaned, weighted by audio duration across roughly eight hours.
Independent leaderboard, and the one class of evidence that would settle vendor disagreements. All three of its datasets are English, so for Indic ASR it contributes nothing and no equivalent exists.
Artificial Analysis · Speech-to-text leaderboard, AA-WER v2 index · fetched Sep 2026
Call a live agent before you decide
These are running agents, not recordings. Open one, press call and speak to it in Hindi or English — the same stack that runs the deployments described above.
Yojana Didi — scheme helpline
A live Hindi agent over a real phone path. Worth calling while reading this page: it is the audio condition the benchmarks above do not measure.
Open the demo →Santhali language line
Santhali appears in exactly one benchmark on this page, at 35.4. A useful listen for what a low-resource Indian language sounds like in production.
Open the demo →Go deeper
The latency companion to this page
The same compilation method applied to response time: every published latency figure, what each one measures, and why they cannot be added together. Written by AiSewak's founder on his own site.
Bhashini and the government language stack
What Bhashini is, how a department reaches it, and where it sits relative to the commercial models in the table above.
Indian-language voice agents
The delivery side of the same question: which languages a deployed agent covers and what it takes to add one.
Procuring government voice AI
NICSI, C-DAC and GeM routes — and where an accuracy requirement belongs in a tender if it is to mean anything.
Related pages: AI voice agent for government · Indian-language coverage · Helpline call automation · Grievance redressal · TRAI and ECI rules for outbound calls · Official helpline numbers · AiSewak's AI ethics guidelines
Frequently asked questions
What is a good word error rate for an Indian-language voice agent?
There is no published standard, and anyone quoting one is quoting a benchmark rather than a service level. The best published Hindi figures sit around 10 to 15 on read and semi-spontaneous benchmarks, and the same models score in the mid-twenties to high-fifties on telephone-quality Hindi. Set your own threshold by measuring your own calls and deciding what proportion of mis-heard turns your escalation path can absorb.
Does Bhashini publish accuracy figures for its ASR models?
No. The Bhashini API catalogue lists service IDs, the languages each model covers and the institution that built it, with no error rate of any kind. The accuracy figures for the AI4Bharat models behind several of those service IDs are published in AI4Bharat's research papers instead — IndicVoices, Vistaar and IndicSUPERB, all linked on this page.
Can I compare Whisper's Hindi number to Sarvam's Hindi number?
Not directly. Whisper's 21.5 is FLEURS, studio-read, scored with Whisper's own text normaliser, from 2022. Sarvam's 19.31 is a ten-language pooled average on IndicVoices, which is mostly extempore and conversational, scored by Sarvam, in 2026. They differ in dataset, in speaking style, in normalisation, in scope and by four years. The only way to compare two systems is to run both on the same audio with the same scoring.
Why do some published word error rates exceed 100%?
Word error rate counts insertions as well as substitutions and deletions, so a model that emits more words than the reference contains can exceed 100. In practice it means the model has almost no coverage of that language and is generating text rather than transcribing. Whisper large-v2 does this on FLEURS Gujarati at 102.7 and Sindhi at 156.5, while scoring 21.5 on Hindi — the same model, the same benchmark.
Why is accuracy on a phone call worse than the benchmark number?
Four reasons, all with published support. Telephony is narrowband at 8 kHz while every model in the table expects 16 kHz. Callers code-switch between an Indian language and English mid-sentence. Helpline calls are dense with numbers and names, which are scored separately as named entity error rate and run consistently worse than overall word error. And dialect varies by district in a way that no national benchmark captures. The size of the channel effect is visible in one published pair: the same model on the same language scores 7.6 on studio-read Hindi and 26.8 on telephone-quality Hindi.
Is there an independent leaderboard for Indic speech recognition?
No. Artificial Analysis publishes the most cited independent word-error-rate index for speech-to-text, and all three datasets behind it are English. Every Indic figure in circulation is either an academic benchmark, usually run by the team that also built one of the models being compared, or a vendor scoring itself. That is the main reason a department should run its own evaluation rather than choose from a table.
How many calls does a department need to evaluate this properly?
Enough per language that the figure stops moving when you add more, and enough per dialect region that the regional breakdown is not noise. Report the sample in both calls and minutes of audio, put a confidence interval on every figure, and report character error rate alongside word error rate. A single pooled number with no interval and no sample size cannot be acted on, and will not survive the first argument about it.
Want this run on your own helpline recordings?
We will not quote you an accuracy number before measuring one. If you have held-out recordings from a line you are considering automating, we will run the evaluation described above on them and hand you the figures, the sample sizes and the intervals — including the languages where the answer is that it is not ready.
Every AiSewak agent identifies itself as an AI at the start of the call, never asks for an OTP or a payment, and honours DND. Election deployments require ECI / state CEO registration as a political advertiser.