Executive Summary
Bhashini, the national language platform run by the Digital India Bhashini Division (DIBD) under MeitY, exposes three services through an API: speech recognition (ASR), machine translation (NMT) and text-to-speech (TTS). A helpline can chain them. The caller speaks, the system transcribes, answers from the department's own records and speaks back in the caller's language.
This guide is for the state IT department, NIC unit or helpline cell that already runs an IVR and a team of agents. Today the caller picks a language from a menu and reaches an agent who may or may not speak it, and the agent types the complaint into a portal. After integration, the caller speaks in their own language, the portal stays the record, and officers keep every decision.
Executive Callout Settle three things before any code. First, Bhashini's public API documentation says usage is "for the purposes of PoC only" and sends production users to the Bhashini team for a paid version. Second, the model catalogue lists languages but publishes no accuracy or latency figure, so every language must be tested on your callers' audio. Third, the pages we read do not state retention or hosting terms for call audio, so get them in writing. AiSewak is an implementer that builds on services like this one. It is not the Bhashini portal and claims no endorsement or partnership from it.
Introduction: What the National Chatbot Shows
On 30 May 2026 the Department of Administrative Reforms and Public Grievances (DARPG) launched Samadhan Didi, a CPGRAMS voice chatbot built with Bhashini. According to the PIB release, a citizen can describe a complaint by speaking, the chatbot asks a few clarifying questions, identifies the Ministry, Department, category and sub-category, and files the grievance. It supports the 22 languages of the Eighth Schedule, and Bhojpuri, Garo, Khasi and others are being added in phases. The minister urged states to bring similar voice tools into their own grievance portals.
Two details matter for a state department. The release says the system combines Bhashini's language capabilities with grievance-classification models trained on CPGRAMS data, so language and case logic are separate parts. And it says the chatbot was developed within secure Government infrastructure. The release does not describe the API arrangement behind it. A state department calling a hosted API has to decide for itself where call audio goes.
What the Public Documentation Gives You
The Bhashini API documentation lists model "service IDs" per task, each tied to a model provider. This is what its models page showed on 5 October 2026.
| Task | What the page lists | What it does not say |
|---|---|---|
| ASR | 12 service IDs from AI4Bharat, IIT Madras, IISc and Bodhan.AI. One lists the 22 scheduled languages; Hindi also has a dedicated model | Accuracy, latency or readiness per language |
| NMT | 10 service IDs. The broadest covers more than 30 languages and dialects, including Awadhi, Bhojpuri, Braj and Magahi | Quality per language pair |
| TTS | 7 service IDs. One lists all 22 scheduled languages plus English and Rajasthani; Santali, Sindhi and Kashmiri are marked "Devanagari Script" | Voice quality, number of voices |
| Also listed | Transliteration, audio and text language detection |
A language in this table is catalogue metadata. Our earlier guide to the Bhashini language stack argues that catalogue is not deployment, and the page above supports the point: it carries no accuracy figure. For the numbers that have been published elsewhere, see the reference of Indic speech recognition accuracy.
The Santali, Sindhi and Kashmiri note is an integration trap. If your text layer produces a script different from the one the voice model expects, the call returns audio that sounds wrong or says nothing. Check the script per language, not just the language code.
How a Call Moves Through the API
The documentation describes a pipeline model. A Pipeline Search call is optional and finds available pipelines. A Pipeline Config call is mandatory: you name the task or task sequence, and the response returns the service ID for each task, a callback URL and an inference API key. A Pipeline Compute call is the one that produces output, using those values.
For a phone line, four details in the compute request decide the design:
- Audio goes in as a base64 string inside a JSON body, with
audioFormat(wav, flac, or mp3 accepted) andsamplingRate. The documented minimum sampling rate is 8000 Hz, which matches narrowband telephone audio. The values you send must match the audio actually recorded. - Pre- and post-processors are options. The documentation's example uses voice activity detection ("vad") before recognition and inverse text normalisation ("itn") after it. ITN matters on a helpline, because callers say reference numbers, amounts and dates.
- A streaming path exists. The WebSocket ASR page specifies 8000 Hz, 16-bit, mono PCM in wav format, and notes that voice activity detection depends on the speaker being loud and clear enough.
- No latency figure is published. A turn is ASR, then your logic, then TTS, and each hop crosses the network. Measure the full turn on your own telecom path before promising citizens anything about wait times.
The citizen journey
- The caller reaches the helpline and says what the problem is. The IVR asks for a language only as a fallback.
- The audio chunk goes to ASR with the language hint, or a language-detection call picks one.
- Your layer decides what the caller wants: status of a reference number, a new complaint, a transfer to a person.
- The layer reads or writes the department's portal through its own API. Read-only comes first; see the voice-layer guide.
- The reply text passes through translation if your logic runs in another language, then TTS.
- Numbers and the reference are read back and confirmed. A low-confidence turn or a repeated failure moves the call to a human.
Who decides what
| Decision | Owner |
|---|---|
| Which words the caller said | Bhashini ASR, checked by your confidence rules |
| What the caller wants | Your integration layer |
| Case state, reference number, rung | The department's portal |
| Closing, escalating, assigning | A named officer |
| When to hand to a person | A written rule the department approves |
Seven Questions to Put in Writing
The documentation is a developer guide. A procurement officer needs answers it does not give.
| # | Question | Why it matters |
|---|---|---|
| 1 | Is this use covered? | The documentation limits use to proof of concept and points commercial use to a paid version |
| 2 | What are the rate limits and call quotas? | The pages we read specify none, and a 10 am surge is the test |
| 3 | What availability is committed? | A helpline cannot go silent when a hosted API is down |
| 4 | Is call audio stored, and where? | The pages we read are silent; this decides your DPDP position |
| 5 | Is audio used to train models? | Recordings hold names, addresses and complaint details |
| 6 | Who is the vendor of record? | Registration is by integrator, with credentials issued after DIBD approves the account |
| 7 | What happens if a language model changes? | Service IDs point to named models, so a change can alter accuracy without a code change |
The DPDP Act's consent and penalty sections commence on 13 May 2027, so a contract signed now will run into them. The DPDP guide for government voice AI covers retention and consent in more detail.
A Pilot That Tests Languages, Not Demos
Pick one helpline, one district and no more than three languages. Include Hindi, because a working Hindi line proves the plumbing, and add the languages your callers actually use.
Record real calls, with consent and the department's data rules in place, and build a test set per language from them. Include what breaks Indian phone speech: numerals and reference numbers, a caller switching to English mid-sentence, scheme names, and local dialect. The multilingual voice agent page describes these failure modes and how acceptance testing handles them.
| Measure | How to read it |
|---|---|
| Word error rate per language on your own 8 kHz recordings | Compare languages, never average them |
| Reference-number capture, exact match after read-back | The error a citizen feels most |
| Calls handed to a human, per language | A language with a high handoff rate is not ready |
| Full-turn latency at the 90th percentile | Measured from caller silence to first audio |
| Completed task rate per language | Status obtained or complaint filed, versus abandoned |
Set targets from a baseline measured on the department's current line. No honest figure exists until that baseline does. Treat a successful API response as nothing more than a successful response: listen to the output in each language, because a call can return without producing intelligible speech.
Risks and Mitigation
- Production use on a trial footing. Resolve question 1 before the pilot expands past a proof of concept.
- One language failing quietly. Report per-language completion and handoff every week. An average hides the language where citizens give up.
- Single point of failure. Keep a fallback path per language, a second recogniser or a human queue, and test the switch.
- Overstated language claims. "22 languages" describes a catalogue. Tell citizens and officials which languages the line has passed acceptance in.
- Personal data in audio. Fix purpose, retention and access roles in the pilot order before the first recorded call.
Key Takeaways
- Bhashini offers ASR, translation and TTS through a config-then-compute API; the department's portal, not the language layer, stays the system of record.
- The public documentation limits use to proof of concept and publishes no accuracy, latency or rate-limit figures. Treat those as contract items.
- Telephone audio is narrowband. Test every language on 8 kHz recordings of your own callers.
- Keep decisions with named officers and write the handoff rule down.
- Test language by language, and report the results the same way.
Conclusion
The Samadhan Didi launch shows that a spoken grievance line in many Indian languages is something a department can build. For a state helpline the harder work is not the API call. It is the terms of use, the per-language testing and the rules for who decides what. A district cell can start with one practical step: write down which of the seven questions above it can already answer, and which it would need Bhashini to answer.
Government leaders exploring AI-powered citizen engagement can begin with a focused pilot in one department or constituency to validate impact before scaling statewide. Aisewak helps public institutions deploy multilingual Voice AI solutions designed specifically for Indian governance.
FAQ
What is the Bhashini API? It is a set of APIs from the Digital India Bhashini Division for speech recognition, machine translation and text-to-speech in Indian languages. Integrators register on Bhashini's dashboard and receive a user ID and API keys once the account is approved.
Is the Bhashini API free for a government helpline? The public documentation says usage of the APIs is for proof-of-concept purposes only. For production systems, or where integrators charge end-users, it directs you to the Bhashini team for the paid version and pricing plans. Confirm the terms that apply to your department in writing.
How many languages does Bhashini support for voice? Its models page lists one ASR model covering the 22 scheduled languages and one TTS model covering them as well, with three in Devanagari script. A listing is not a test result. The page publishes no accuracy figure, so measure each language on your own calls.
Can Bhashini handle a live phone call? The documentation accepts audio at a minimum of 8000 Hz and describes a WebSocket ASR path at 8 kHz, 16-bit mono. It publishes no latency figure, so a department must measure the full turn on its own telephony path.
Does a helpline need Bhashini to replace its portal? No. The portal remains the record of cases and references. The language layer reads and writes it through the department's own API, starting read-only.
Is AiSewak part of Bhashini? No. AiSewak is an implementer that builds voice agents using language services, and claims no endorsement or partnership from Bhashini or MeitY.
What audio format should a helpline send? The documentation accepts wav, flac and mp3, and says the sampling rate and format you declare must match the recorded audio. Telephone audio is 8 kHz, so test at that rate.
Who should answer when the voice agent fails? A human on a queue the department staffs. Write the trigger, such as a repeated low-confidence turn or a caller request, into the pilot order.
Schema Markup Suggestions
- TechArticle —
headline,description,datePublished,dateModified,author(Organization: AiSewak Editorial),publisher - FAQPage — each Q/A above as
Question+acceptedAnswer - HowTo — for the citizen journey:
name: Add Indian-Language Voice to a Government Helpline with Bhashini APIs, sixHowToStepitems - GovernmentService — on the linked
/multilingual-ai-voice-agent-indialanding:serviceType: Multilingual Citizen Helpline,areaServed: India
Suggested Internal Links
/multilingual-ai-voice-agent-india— multilingual voice agents for India/blog/multilingual-voice-ai-government-bhashini— the Bhashini language stack and readiness tiers/blog/voice-layer-existing-grievance-portal— adding a phone channel to a portal/blog/government-voice-ai-dpdp-privacy-security— DPDP, retention and consent/indic-speech-recognition-accuracy-benchmark— published Indic speech recognition accuracy
Suggested External References
- Bhashini APIs documentation, Digital India Bhashini Division — Overall Understanding of the API Calls
- Bhashini APIs: Available Models for usage, Pipeline Compute request payload, WebSocket ASR API, Pre-requisites and Onboarding — read 5 October 2026
- PIB, Ministry of Personnel, Public Grievances and Pensions, 30 May 2026 — launch of the CPGRAMS voice chatbot Samadhan Didi
- Digital Personal Data Protection Act, 2023 — commencement of consent and penalty provisions on 13 May 2027
Social Media Summary
Bhashini's own API documentation limits use to proof of concept and publishes no accuracy or latency figures. A guide for state IT departments adding Indian-language voice to a helpline: the call flow, the audio specs, seven contract questions and a language-by-language pilot.
LinkedIn Executive Summary
Samadhan Didi, launched by DARPG with Bhashini on 30 May 2026, lets a citizen file a grievance by speaking in any of the 22 scheduled languages. The minister urged states to bring similar tools into their own portals.
For a state IT department that runs a helpline, the integration is a config call and a compute call per task. The harder part is what the public documentation leaves open. It limits use to proof of concept and points production users to a paid version. It lists ASR and TTS models for the scheduled languages but publishes no accuracy, latency or rate-limit figure, and the pages we read are silent on audio retention.
The practical route is a pilot in one district and no more than three languages, tested on recorded 8 kHz calls. Measure error rate, reference-number capture, handoff to a human and full-turn latency per language, and report them the same way. The portal stays the record and named officers keep every decision.
AI Search Optimization Summary
Entities: Bhashini, Digital India Bhashini Division, MeitY, DARPG, CPGRAMS, Samadhan Didi, PIB, AI4Bharat, IIT Madras, IISc, Bodhan.AI, Eighth Schedule, DPDP Act, Aisewak
Topics: Bhashini API integration, multilingual government helpline, Indian language speech recognition for phone calls, state helpline voice AI, ASR NMT TTS pipeline, proof-of-concept versus production API terms
Semantic keywords: Bhashini pipeline config and compute call, 8 kHz telephone audio ASR, service ID per language, per-language acceptance testing, helpline handoff to human officer, read-only portal integration, Devanagari script TTS for Santali, call audio retention under DPDP