AiSewak
Integration Guide · Bhashini and Indic Delivery

Using Bhashini APIs in a Government Helpline

How a state IT department can wire Bhashini speech recognition, translation and voice into a helpline: call flow, audio, usage terms and pilot checks.

13 min readUpdated 5 Oct 20262,676 words

Executive Summary

Bhashini, the national language platform run by the Digital India Bhashini Division (DIBD) under MeitY, exposes three services through an API: speech recognition (ASR), machine translation (NMT) and text-to-speech (TTS). A helpline can chain them. The caller speaks, the system transcribes, answers from the department's own records and speaks back in the caller's language.

This guide is for the state IT department, NIC unit or helpline cell that already runs an IVR and a team of agents. Today the caller picks a language from a menu and reaches an agent who may or may not speak it, and the agent types the complaint into a portal. After integration, the caller speaks in their own language, the portal stays the record, and officers keep every decision.

Executive Callout Settle three things before any code. First, Bhashini's public API documentation says usage is "for the purposes of PoC only" and sends production users to the Bhashini team for a paid version. Second, the model catalogue lists languages but publishes no accuracy or latency figure, so every language must be tested on your callers' audio. Third, the pages we read do not state retention or hosting terms for call audio, so get them in writing. AiSewak is an implementer that builds on services like this one. It is not the Bhashini portal and claims no endorsement or partnership from it.

Introduction: What the National Chatbot Shows

On 30 May 2026 the Department of Administrative Reforms and Public Grievances (DARPG) launched Samadhan Didi, a CPGRAMS voice chatbot built with Bhashini. According to the PIB release, a citizen can describe a complaint by speaking, the chatbot asks a few clarifying questions, identifies the Ministry, Department, category and sub-category, and files the grievance. It supports the 22 languages of the Eighth Schedule, and Bhojpuri, Garo, Khasi and others are being added in phases. The minister urged states to bring similar voice tools into their own grievance portals.

Two details matter for a state department. The release says the system combines Bhashini's language capabilities with grievance-classification models trained on CPGRAMS data, so language and case logic are separate parts. And it says the chatbot was developed within secure Government infrastructure. The release does not describe the API arrangement behind it. A state department calling a hosted API has to decide for itself where call audio goes.

What the Public Documentation Gives You

The Bhashini API documentation lists model "service IDs" per task, each tied to a model provider. This is what its models page showed on 5 October 2026.

TaskWhat the page listsWhat it does not say
ASR12 service IDs from AI4Bharat, IIT Madras, IISc and Bodhan.AI. One lists the 22 scheduled languages; Hindi also has a dedicated modelAccuracy, latency or readiness per language
NMT10 service IDs. The broadest covers more than 30 languages and dialects, including Awadhi, Bhojpuri, Braj and MagahiQuality per language pair
TTS7 service IDs. One lists all 22 scheduled languages plus English and Rajasthani; Santali, Sindhi and Kashmiri are marked "Devanagari Script"Voice quality, number of voices
Also listedTransliteration, audio and text language detection

A language in this table is catalogue metadata. Our earlier guide to the Bhashini language stack argues that catalogue is not deployment, and the page above supports the point: it carries no accuracy figure. For the numbers that have been published elsewhere, see the reference of Indic speech recognition accuracy.

The Santali, Sindhi and Kashmiri note is an integration trap. If your text layer produces a script different from the one the voice model expects, the call returns audio that sounds wrong or says nothing. Check the script per language, not just the language code.

How a Call Moves Through the API

The documentation describes a pipeline model. A Pipeline Search call is optional and finds available pipelines. A Pipeline Config call is mandatory: you name the task or task sequence, and the response returns the service ID for each task, a callback URL and an inference API key. A Pipeline Compute call is the one that produces output, using those values.

For a phone line, four details in the compute request decide the design:

  • Audio goes in as a base64 string inside a JSON body, with audioFormat (wav, flac, or mp3 accepted) and samplingRate. The documented minimum sampling rate is 8000 Hz, which matches narrowband telephone audio. The values you send must match the audio actually recorded.
  • Pre- and post-processors are options. The documentation's example uses voice activity detection ("vad") before recognition and inverse text normalisation ("itn") after it. ITN matters on a helpline, because callers say reference numbers, amounts and dates.
  • A streaming path exists. The WebSocket ASR page specifies 8000 Hz, 16-bit, mono PCM in wav format, and notes that voice activity detection depends on the speaker being loud and clear enough.
  • No latency figure is published. A turn is ASR, then your logic, then TTS, and each hop crosses the network. Measure the full turn on your own telecom path before promising citizens anything about wait times.

The citizen journey

  1. The caller reaches the helpline and says what the problem is. The IVR asks for a language only as a fallback.
  2. The audio chunk goes to ASR with the language hint, or a language-detection call picks one.
  3. Your layer decides what the caller wants: status of a reference number, a new complaint, a transfer to a person.
  4. The layer reads or writes the department's portal through its own API. Read-only comes first; see the voice-layer guide.
  5. The reply text passes through translation if your logic runs in another language, then TTS.
  6. Numbers and the reference are read back and confirmed. A low-confidence turn or a repeated failure moves the call to a human.

Who decides what

DecisionOwner
Which words the caller saidBhashini ASR, checked by your confidence rules
What the caller wantsYour integration layer
Case state, reference number, rungThe department's portal
Closing, escalating, assigningA named officer
When to hand to a personA written rule the department approves

Seven Questions to Put in Writing

The documentation is a developer guide. A procurement officer needs answers it does not give.

#QuestionWhy it matters
1Is this use covered?The documentation limits use to proof of concept and points commercial use to a paid version
2What are the rate limits and call quotas?The pages we read specify none, and a 10 am surge is the test
3What availability is committed?A helpline cannot go silent when a hosted API is down
4Is call audio stored, and where?The pages we read are silent; this decides your DPDP position
5Is audio used to train models?Recordings hold names, addresses and complaint details
6Who is the vendor of record?Registration is by integrator, with credentials issued after DIBD approves the account
7What happens if a language model changes?Service IDs point to named models, so a change can alter accuracy without a code change

The DPDP Act's consent and penalty sections commence on 13 May 2027, so a contract signed now will run into them. The DPDP guide for government voice AI covers retention and consent in more detail.

A Pilot That Tests Languages, Not Demos

Pick one helpline, one district and no more than three languages. Include Hindi, because a working Hindi line proves the plumbing, and add the languages your callers actually use.

Record real calls, with consent and the department's data rules in place, and build a test set per language from them. Include what breaks Indian phone speech: numerals and reference numbers, a caller switching to English mid-sentence, scheme names, and local dialect. The multilingual voice agent page describes these failure modes and how acceptance testing handles them.

MeasureHow to read it
Word error rate per language on your own 8 kHz recordingsCompare languages, never average them
Reference-number capture, exact match after read-backThe error a citizen feels most
Calls handed to a human, per languageA language with a high handoff rate is not ready
Full-turn latency at the 90th percentileMeasured from caller silence to first audio
Completed task rate per languageStatus obtained or complaint filed, versus abandoned

Set targets from a baseline measured on the department's current line. No honest figure exists until that baseline does. Treat a successful API response as nothing more than a successful response: listen to the output in each language, because a call can return without producing intelligible speech.

Risks and Mitigation

  • Production use on a trial footing. Resolve question 1 before the pilot expands past a proof of concept.
  • One language failing quietly. Report per-language completion and handoff every week. An average hides the language where citizens give up.
  • Single point of failure. Keep a fallback path per language, a second recogniser or a human queue, and test the switch.
  • Overstated language claims. "22 languages" describes a catalogue. Tell citizens and officials which languages the line has passed acceptance in.
  • Personal data in audio. Fix purpose, retention and access roles in the pilot order before the first recorded call.

Key Takeaways

  • Bhashini offers ASR, translation and TTS through a config-then-compute API; the department's portal, not the language layer, stays the system of record.
  • The public documentation limits use to proof of concept and publishes no accuracy, latency or rate-limit figures. Treat those as contract items.
  • Telephone audio is narrowband. Test every language on 8 kHz recordings of your own callers.
  • Keep decisions with named officers and write the handoff rule down.
  • Test language by language, and report the results the same way.

Conclusion

The Samadhan Didi launch shows that a spoken grievance line in many Indian languages is something a department can build. For a state helpline the harder work is not the API call. It is the terms of use, the per-language testing and the rules for who decides what. A district cell can start with one practical step: write down which of the seven questions above it can already answer, and which it would need Bhashini to answer.

Government leaders exploring AI-powered citizen engagement can begin with a focused pilot in one department or constituency to validate impact before scaling statewide. Aisewak helps public institutions deploy multilingual Voice AI solutions designed specifically for Indian governance.


FAQ

What is the Bhashini API? It is a set of APIs from the Digital India Bhashini Division for speech recognition, machine translation and text-to-speech in Indian languages. Integrators register on Bhashini's dashboard and receive a user ID and API keys once the account is approved.

Is the Bhashini API free for a government helpline? The public documentation says usage of the APIs is for proof-of-concept purposes only. For production systems, or where integrators charge end-users, it directs you to the Bhashini team for the paid version and pricing plans. Confirm the terms that apply to your department in writing.

How many languages does Bhashini support for voice? Its models page lists one ASR model covering the 22 scheduled languages and one TTS model covering them as well, with three in Devanagari script. A listing is not a test result. The page publishes no accuracy figure, so measure each language on your own calls.

Can Bhashini handle a live phone call? The documentation accepts audio at a minimum of 8000 Hz and describes a WebSocket ASR path at 8 kHz, 16-bit mono. It publishes no latency figure, so a department must measure the full turn on its own telephony path.

Does a helpline need Bhashini to replace its portal? No. The portal remains the record of cases and references. The language layer reads and writes it through the department's own API, starting read-only.

Is AiSewak part of Bhashini? No. AiSewak is an implementer that builds voice agents using language services, and claims no endorsement or partnership from Bhashini or MeitY.

What audio format should a helpline send? The documentation accepts wav, flac and mp3, and says the sampling rate and format you declare must match the recorded audio. Telephone audio is 8 kHz, so test at that rate.

Who should answer when the voice agent fails? A human on a queue the department staffs. Write the trigger, such as a repeated low-confidence turn or a caller request, into the pilot order.


Schema Markup Suggestions

  • TechArticle — headline, description, datePublished, dateModified, author (Organization: AiSewak Editorial), publisher
  • FAQPage — each Q/A above as Question + acceptedAnswer
  • HowTo — for the citizen journey: name: Add Indian-Language Voice to a Government Helpline with Bhashini APIs, six HowToStep items
  • GovernmentService — on the linked /multilingual-ai-voice-agent-india landing: serviceType: Multilingual Citizen Helpline, areaServed: India

  • /multilingual-ai-voice-agent-india — multilingual voice agents for India
  • /blog/multilingual-voice-ai-government-bhashini — the Bhashini language stack and readiness tiers
  • /blog/voice-layer-existing-grievance-portal — adding a phone channel to a portal
  • /blog/government-voice-ai-dpdp-privacy-security — DPDP, retention and consent
  • /indic-speech-recognition-accuracy-benchmark — published Indic speech recognition accuracy

Suggested External References


Social Media Summary

Bhashini's own API documentation limits use to proof of concept and publishes no accuracy or latency figures. A guide for state IT departments adding Indian-language voice to a helpline: the call flow, the audio specs, seven contract questions and a language-by-language pilot.


LinkedIn Executive Summary

Samadhan Didi, launched by DARPG with Bhashini on 30 May 2026, lets a citizen file a grievance by speaking in any of the 22 scheduled languages. The minister urged states to bring similar tools into their own portals.

For a state IT department that runs a helpline, the integration is a config call and a compute call per task. The harder part is what the public documentation leaves open. It limits use to proof of concept and points production users to a paid version. It lists ASR and TTS models for the scheduled languages but publishes no accuracy, latency or rate-limit figure, and the pages we read are silent on audio retention.

The practical route is a pilot in one district and no more than three languages, tested on recorded 8 kHz calls. Measure error rate, reference-number capture, handoff to a human and full-turn latency per language, and report them the same way. The portal stays the record and named officers keep every decision.


AI Search Optimization Summary

Entities: Bhashini, Digital India Bhashini Division, MeitY, DARPG, CPGRAMS, Samadhan Didi, PIB, AI4Bharat, IIT Madras, IISc, Bodhan.AI, Eighth Schedule, DPDP Act, Aisewak

Topics: Bhashini API integration, multilingual government helpline, Indian language speech recognition for phone calls, state helpline voice AI, ASR NMT TTS pipeline, proof-of-concept versus production API terms

Semantic keywords: Bhashini pipeline config and compute call, 8 kHz telephone audio ASR, service ID per language, per-language acceptance testing, helpline handoff to human officer, read-only portal integration, Devanagari script TTS for Santali, call audio retention under DPDP

AiSewak (AI Sewak) is a Voxdonna company, made in India.

© 2026 Donna AI Labs Private Limited · CIN U62013DL2026PTC464877. All rights reserved.