Executive Summary
India's government helplines measure the wrong things. Rajasthan Sampark claims 99.36% disposal while running 1 lakh pending cases. CPGRAMS reports 95% disposal while BSNL's own feedback surveys return 44–51% citizen satisfaction. The gap between what departments count and what citizens experience is not a data-quality problem — it is a measurement-design problem.
Executive Callout Every major government voice AI deployment risks failing not because the technology underdelivers, but because the measurement framework was designed to show activity rather than resolution. Haryana's AI-powered 112 dispatch is the single published Indian exception: it tracked real outcomes — emergency response time from 12 to 7 minutes, 92.60% citizen satisfaction — and earned MHA national recognition. The framework in this article replicates that approach across any government helpline. (Aisewak Government Helpline Report, 2026, citing Haryana 112 operational data and MHA recognition records.)
This article gives government leaders the measurement vocabulary, the benchmark numbers, and an implementation roadmap to prove impact — to the CAG, to parliamentary committees, and to the citizens the system was built to serve.
Introduction: The Disposal Rate Delusion
When a government department "disposes" a complaint, it marks the case closed in the system. The citizen may have received a callback, a template SMS, or nothing at all — the disposal rate does not distinguish. This is the foundational measurement failure across India's government helpline ecosystem.
The data is unambiguous. DARPG claims a 95% disposal rate on CPGRAMS grievances, with an average resolution time of 15 days. The BSNL Feedback Call Centre — which phones citizens after purported resolution — found only 44% satisfaction in March 2024 and 51% in December 2024 (Aisewak Government Helpline Report, 2026, citing BSNL feedback data and Lok Sabha Question 5217, April 2025). Nearly half of every "resolved" complaint left the citizen worse off.
The same pattern repeats. Rajasthan Sampark's 99.36% disposal rate coexists with 1 lakh pending cases — a contradiction the Chief Secretary himself acknowledged when he began personally visiting the call centre daily. The 108 Ambulance receives 250,000 calls daily but answers only 86,000 — a 66% abandonment rate that never appears in any official performance dashboard (Aisewak Government Helpline Report, 2026, citing CAG Karnataka and Odisha audit data).
Voice AI generates granular, auditable interaction data on every call — every second, every transfer decision, every resolution pathway. Departments that deploy voice AI without upgrading their measurement framework will replicate the disposal-rate delusion at scale. Those that get the metrics right will have, for the first time, evidence that withstands CAG audit and parliamentary scrutiny.
Why Traditional Metrics Fail
Three structural flaws explain the systematic misrepresentation.
Closure does not equal resolution. A grievance officer who reads a complaint, marks it "forwarded to field office," and closes the ticket has disposed it in 30 seconds. The citizen has received nothing. DARPG's own data shows only 42.4% of Grievance Redressal Officers were active as of June 2024 — well below the 100% mandate — yet disposal rates remain officially above 90% (Aisewak Government Helpline Report, 2026, citing DARPG CPGRAMS 28th and 29th Reports).
Volume hides failure modes. Call centres report total calls received and disposed. They do not report calls abandoned before answer, calls requiring repeat contact for the same issue, or cases where citizens give up. The 181 Women Helpline's 88% no-response rate never appears in departmental performance reports because calls that are never answered do not enter the system (Aisewak Government Helpline Report, 2026, citing AALI survey across Uttar Pradesh, Uttarakhand, and Jharkhand).
Satisfaction surveys are not routine. The BSNL mechanism on CPGRAMS is one of very few systems where citizens are proactively asked whether their issue was actually resolved. Most helplines have no feedback loop at all — meaning there is no mechanism to detect the 49–56% of "resolved" cases where citizens remain dissatisfied.
The Five KPI Categories That Actually Matter
A measurement framework for government voice AI must track five distinct dimensions. Each maps to a different accountability audience: citizens, department heads, budget officers, and the CAG.
1. Accessibility Metrics
These measure whether citizens can actually reach the service.
| KPI | Definition | Benchmark Target |
|---|---|---|
| Call answer rate | % of calls answered within 30 seconds | ≥95% |
| Abandonment rate | % of callers who hang up before answer | ≤5% |
| 24/7 uptime | % of hours the system is operational | ≥99.5% |
| Language coverage | % of callers served in their preferred language | ≥90% |
The Kisan Call Centre's 45.7% answer rate and the 181 Women Helpline's 88% no-response rate make these the first line of accountability — and the most politically visible (Aisewak Government Helpline Report, 2026, citing IIM Ahmedabad study and NITI Aayog 2021 survey).
2. Efficiency Metrics
These measure what happens once a citizen is connected.
| KPI | Definition | Benchmark Target |
|---|---|---|
| Average Handle Time (AHT) | Mean call duration for AI-handled contacts | ≤3 minutes for enquiries |
| First-Call Resolution (FCR) | % of issues resolved without callback or repeat contact | ≥65% for AI-handled queries |
| Containment Rate | % of calls AI resolves without human transfer | ≥60% |
| Triage accuracy | Correct routing of emergency vs. non-emergency calls | ≥92% |
The Rajasthan Sampark pilot design targets containment ≥60% and AHT under 3 minutes, versus 7-day human resolution for grievance registration. The 108 Ambulance pilot targets triage accuracy ≥92% — emergency calls must never be wrongly held in the non-emergency queue (Aisewak Government Helpline Report, 2026).
3. Quality Metrics
These measure resolution depth, not closure volume.
| KPI | Definition | Benchmark Target |
|---|---|---|
| Citizen Satisfaction Score (CSAT) | Post-call survey: "Was your issue resolved?" | ≥75% for AI-handled calls |
| Escalation quality | % of AI-to-human transfers confirmed as correctly routed | ≥90% |
| Resolution recurrence | % of citizens who call again for the same issue within 30 days | ≤15% |
| Complaint re-opening rate | % of closed cases citizens re-open | ≤10% |
CSAT is the single most important number. Haryana's 92.60% post-AI satisfaction is the only published Indian benchmark for a government voice AI deployment at scale — and it was earned by tracking resolution quality, not disposal counts (Aisewak Government Helpline Report, 2026, citing Haryana 112 data).
4. Cost Metrics
These make the business case for budget officers and finance departments.
| KPI | Definition | Benchmark |
|---|---|---|
| Cost per call (AI) | Total AI operating cost ÷ AI-handled calls | Rs 2–5 per call |
| Cost per call (human) | Total agent cost ÷ human-handled calls | Rs 20–30 per call |
| Cost per resolved case | Total cost ÷ cases with confirmed resolution | Track improvement from baseline |
| Agent time freed | Hours per week reclaimed for complex cases | Track from baseline |
AI voice pricing of Rs 2–5 per call is materially below fully-loaded human agent cost across every helpline profiled (Aisewak Government Helpline Report, 2026). Cost-per-resolved-case — not cost-per-call — is the correct government efficiency metric: it reveals that a high-disposal, low-satisfaction system costs more per actual resolution than an AI system with genuine FCR.
5. Equity Metrics
These are unique to government and absent from most private-sector AI measurement frameworks.
| KPI | Definition | Target |
|---|---|---|
| Rural answer rate | Answer rate for calls from non-metro exchanges | Within 5% of urban rate |
| Dialect recognition accuracy | Speech recognition accuracy per regional dialect | ≥80% per dialect |
| Literacy-independent access | % of calls resolved without requiring the caller to read or type | ≥95% |
| Gender-differentiated CSAT | CSAT scores broken down by gender | No group below 70% |
A government voice AI that achieves 90% satisfaction in urban Hindi-speaking districts but fails in Marwari-speaking rural Rajasthan has not served its constitutional mandate (Aisewak Government Helpline Report, 2026). Aggregate national averages conceal these gaps; district-level breakdowns reveal them.
Indian Benchmarks: Before and After
| Helpline | Baseline (Pre-AI) | Target / Actual (Post-AI) | Source |
|---|---|---|---|
| Haryana 112 response time | 12 minutes average | 7 minutes actual | Haryana CS data; MHA recognition |
| Haryana 112 CSAT | Not formally tracked | 92.60% actual | MHA national recognition records |
| CPGRAMS citizen satisfaction | 44–51% (BSNL feedback) | Target: ≥75% with AI follow-up | Lok Sabha Q.5217, April 2025 |
| Kisan Call Centre answer rate | 45.7% (peak season) | Target: ≥90% with AI first-response | IIM Ahmedabad study; DAC&FW |
| Rajasthan Sampark cost per call | ~Rs 25 (human agent) | Target: Rs 5 (AI-handled) | Aisewak Report pilot design |
| 108 Ambulance abandonment | 66% (250K calls, 86K answered) | Target: ≤5% with AI overflow | CAG Karnataka; NHSRC data |
Haryana 112 is the only case with documented post-AI data because it is the only deployment designed to measure outcomes from the start. Every other helpline profiled in the research data is working from a baseline of unmeasured failure.
International Reference Points
India's measurement challenge is not unique. The UK's HMRC telephone service has published First-Call Resolution data since 2019, enabling consistent year-on-year improvement tracking and providing comparable benchmarks across service categories. Singapore's Singpass helpline reports language-differentiated satisfaction quarterly, allowing the government to identify and remediate minority-language service gaps within weeks of their emergence.
The EU Digital Decade framework targets 80% of government services measuring citizen satisfaction with standardised indicators by 2030, comparable across member states. India's DARPG has acknowledged the CPGRAMS satisfaction paradox — the gap between 95% disposal and 44–51% CSAT — but has not yet published a standardised measurement framework that states can replicate.
Implementation Roadmap
A government department can operationalise voice AI KPIs in four phases over twelve months.
Phase 1 — Baseline (Weeks 1–4). Audit current data. Identify what is tracked (calls received, cases disposed) and what is not (abandonment, satisfaction, recurrence). Establish pre-AI baselines for every metric in the five categories. This phase typically produces data that justifies the AI investment to budget officers without further persuasion — the failure numbers speak for themselves. See the Governance AI Maturity Model for assessing current measurement capability.
Phase 2 — Instrument (Weeks 5–12). Configure the voice AI system to capture interaction data at the call level. Every AI-handled call should produce: timestamp, duration, language detected, query category, resolution pathway (self-serve / transfer / callback), and a post-call CSAT prompt. A single-question IVR after the call ends — "Press 1 if your issue was resolved, Press 2 if not" — costs nothing additional if built into the system from deployment day.
Phase 3 — Pilot Dashboard (Months 4–6). Build a real-time dashboard visible to department heads, district officers, and the CM's office. Track weekly, not monthly — monthly aggregates mask the volatility that reveals system failures. The AI for the Chief Minister's Command Centre framework outlines how this rolls up to executive visibility.
Phase 4 — External Audit (Months 7–12). Commission an independent satisfaction survey — parallel to the BSNL CPGRAMS feedback model — to cross-validate internal CSAT data, and publish the results. Departments that publish honest KPI data, even showing imperfect results, build more durable political credibility than those that publish only disposal rates. For procurement guidance on selecting vendors who can deliver fully-instrumented deployments, see Procuring AI Voice Agents in Government: NICSI, C-DAC, GeM.
Risks and Mitigation
Gaming the metric. Once departments know they are measured on CSAT, some will restrict the survey population to satisfied callers, or route difficult calls away from the AI to avoid negative scores. Mitigation: randomise survey delivery to 100% of AI-handled calls. Design post-call IVR prompts that cannot be suppressed by operators or supervisors.
Dialect bias in speech recognition. A system with 95% overall recognition accuracy may perform at 70% for a specific dialect, systematically under-serving that community while the aggregate number looks acceptable. Mitigation: report recognition accuracy broken down by language and dialect, not as a single national average. Monthly district-level dialect reports, visible to DMs, surface bias before it compounds.
Narrow baselines invalidating before/after comparisons. Departments that cannot measure current performance cannot prove AI improvement. Mitigation: run a parallel measurement sprint on the existing human system for four weeks before deployment. Even an incomplete baseline is more defensible than none — and it reframes the deployment from "technology procurement" to "service improvement initiative," which is a more durable political framing.
Key Takeaways
- Disposal rate is a bureaucratic process metric, not a citizen outcome. Any framework that relies on it will systematically overstate performance.
- Five categories matter: Accessibility, Efficiency, Quality, Cost, and Equity. All five must be tracked from deployment day one.
- CSAT is the single most important number. Haryana's 92.60% is the only published Indian benchmark; every other government helpline is starting from an unmeasured baseline.
- Equity metrics are not optional. A government AI that serves urban Hindi speakers well but fails rural dialect speakers has not served its mandate.
- Design measurement into the system before the first call is handled. Retrofitting analytics onto a live deployment costs more and produces less defensible data.
Conclusion
The measurement framework determines whether a government voice AI deployment becomes a reference success or a quiet failure. Departments that track disposal rates will continue to demonstrate high performance in procurement documents while citizens remain unserved. Departments that track first-call resolution, containment rate, CSAT, and equity indicators will generate evidence that survives CAG audit, parliamentary questioning, and — most importantly — the experience of a citizen calling at 2 AM with nowhere else to turn.
The technology is available. The data infrastructure for measurement adds weeks, not months. The political return on a genuinely measurable AI success — Haryana's 92.60% satisfaction and MHA recognition — is demonstrably higher than a deployment that produces no auditable impact evidence. The choice between these outcomes is a measurement design decision, made before the first AI call is handled.
Government leaders exploring AI-powered citizen engagement can begin with a focused pilot in one department or constituency to validate impact before scaling statewide. Aisewak helps public institutions deploy multilingual Voice AI solutions designed specifically for Indian governance.
FAQ
Q: What is the most important KPI for a government voice AI deployment? First-Call Resolution (FCR) is the most operationally meaningful single metric — it directly measures whether a citizen's issue was resolved in one contact. Citizen Satisfaction Score (CSAT), measured through a post-call survey, is the most politically significant because it reflects the citizen's own judgement about the service, not a bureaucratic closure count.
Q: Why does CPGRAMS show 95% disposal but only 44–51% citizen satisfaction? Disposal means the case was marked closed by a Grievance Redressal Officer — it does not mean the citizen's problem was solved. BSNL's Feedback Call Centre, which contacts citizens after closure, consistently finds that roughly half of "disposed" grievances left citizens dissatisfied because the GRO closed the case without meaningful action. This is the satisfaction paradox documented in DARPG's own feedback data (Lok Sabha Question 5217, April 2025).
Q: What answer rate should a government helpline target after deploying voice AI? A target of ≥95% call answer rate within 30 seconds is achievable with 24/7 AI first-response regardless of volume spikes. The Kisan Call Centre's 45.7% answer rate and the 181 Women Helpline's 88% no-response rate both represent failure modes that AI can structurally eliminate by ensuring every inbound call is answered within seconds, at any hour.
Q: How do you measure voice AI quality without invading caller privacy? Post-call CSAT surveys can be delivered via a brief IVR prompt after the call ends — "Press 1 if your issue was resolved" — without recording or storing call content. Aggregate data is anonymised before reporting. The DPDP Act 2023 framework for government voice AI data handling is covered in the article on DPDP Act, Data Privacy and Security for Government Voice AI.
Q: What containment rate is realistic for government voice AI? 60–75% containment is achievable for high-volume enquiry helplines like Railway 139 (PNR status, schedules) and DISCOM customer care (bill status, outage reporting). For grievance registration helplines like CPGRAMS or Rajasthan Sampark, 50–60% is realistic: AI handles intake, verification, and status queries while complex cases transfer to trained officers.
Q: How should departments track equity in voice AI service delivery? Report recognition accuracy and CSAT broken down by language and district, not as national averages. A national CSAT of 80% that masks 55% satisfaction in tribal or dialect-majority districts is not equity-compliant. Monthly district-level dashboards, accessible to DMs and Municipal Commissioners, provide the granularity needed to detect and fix equity gaps before they compound.
Q: What does a 30-day pilot need to measure to justify statewide rollout? The minimum viable pilot dataset: answer rate, containment rate, triage accuracy (for emergency services), CSAT from post-call surveys, cost per call, and at least one equity metric such as dialect recognition accuracy or rural/urban answer rate comparison. The 30-Day Pilot to Statewide Scale Roadmap outlines how these measurements feed a go/no-go decision framework.
Q: How do government voice AI KPIs differ from private-sector contact centre KPIs? Two differences are significant. First, equity metrics — dialect accessibility, rural reach, literacy-independent access — are legally and constitutionally relevant for government services but absent from private-sector frameworks. Second, cost-per-resolved-case (not cost-per-call) is the correct government efficiency metric, because government services must resolve issues regardless of difficulty, while private contact centres optimise primarily for call deflection. The ROI and Cost-Benefit of Voice AI in Government article covers the full cost metric framework.
Schema Markup Suggestions
- Article — headline, author (Aisewak Editorial), datePublished (2026-09-18), dateModified, publisher (Aisewak)
- FAQPage — all 8 FAQ pairs with Question and Answer entities
- GovernmentService — serviceType: "Citizen Helpline Performance Measurement", areaServed: India
- HowTo — for the four-phase Implementation Roadmap (Baseline, Instrument, Pilot Dashboard, External Audit)
- Dataset — for the five KPI tables referencing Aisewak Government Helpline Report 2026 as primary source
Suggested Internal Links
- AI for Governance in India: The 2026 Executive Guide
- Voice AI for Government: How It Works and Why Now
- Why Traditional Government Helplines Fail
- The 10-Crore-Call Crisis in Indian Citizen Services
- AI vs Traditional Government Call Centres
- A Governance AI Maturity Model
- ROI and Cost-Benefit of Voice AI in Government
- The 30-Day Pilot to Statewide Scale Roadmap
- Procuring AI Voice Agents in Government: NICSI, C-DAC, GeM
- DPDP Act, Data Privacy and Security for Government Voice AI
- AI for the Chief Minister's Command Centre
- Human-in-the-Loop: Augmenting Government Call-Centre Agents
- Aisewak Voice AI for Government
Suggested External References
- Lok Sabha Question 5217, April 2025 — BSNL CPGRAMS feedback data (44% March 2024, 51% December 2024)
- DARPG CPGRAMS 28th and 29th Annual Reports — disposal rate and GRO activity data
- CAG Report on 108 Ambulance Services, Odisha and Karnataka — response-time and non-emergency call data
- NITI Aayog: Women Helpline Awareness Survey, 2021 — 23.5% awareness; 88% no-response finding
- AALI Survey — 181 Women Helpline no-response rate across Uttar Pradesh, Uttarakhand, Jharkhand
- Haryana 112 AI Dispatch Operational Data — MHA recognition; response time 12→7 minutes; 92.60% satisfaction
- IIM Ahmedabad Study on Kisan Call Centre 1551 — 45.7% answer rate; seasonal peak drop data
- Aisewak Government Helpline Report, 2026 — primary research document, 300+ source searches
- EU Digital Decade Policy Programme 2030 — government digital service measurement standards
- DPDP Act, 2023 — Section 4 (lawful processing), Section 7 (legitimate uses for state)
Social Media Summary
X/LinkedIn caption: India's helplines claim 95–99% disposal rates. BSNL surveys show 44–51% citizen satisfaction. Haryana's 112 AI deployment tracked real outcomes — response time from 12 to 7 minutes, 92.60% CSAT, national MHA recognition. The difference is measurement design. Here is the five-category KPI framework every government leader needs before deploying voice AI. aisewak.com/blog/government-voice-ai-kpis-citizen-satisfaction
LinkedIn Executive Summary
India's government helplines operate on a measurement paradox. CPGRAMS disposes 95% of grievances; BSNL's feedback surveys show only 44–51% citizen satisfaction. Rajasthan Sampark claims 99.36% disposal and simultaneously maintains 1 lakh pending cases. These numbers coexist because disposal means "case closed" — not "issue resolved."
Voice AI changes this, but only if measurement is designed into the deployment from day one. Haryana's AI-powered 112 dispatch is the proof point: it tracked emergency response time (12 → 7 minutes) and post-call satisfaction (92.60%), earned MHA national recognition, and created the template that other states are now studying.
The framework is five KPI categories — Accessibility, Efficiency, Quality, Cost, and Equity — with specific metrics, benchmark targets, and a four-phase implementation roadmap. Departments that measure correctly will have, for the first time, evidence that withstands CAG audit and parliamentary scrutiny. Those that do not will continue producing disposal reports that satisfy no one. This choice is a measurement design decision, made before the first AI call is handled.
AI Search Optimization Summary
Primary entities: Government Voice AI India, CPGRAMS KPIs, citizen satisfaction score government, first-call resolution India, Haryana 112 AI, BSNL CPGRAMS feedback, DARPG disposal rate paradox, Aisewak Government Helpline Report 2026, CAG audit government helplines India
Key topics: government helpline performance measurement India, voice AI ROI government, citizen satisfaction vs disposal rate India, KPI framework government contact centre, AI measurement framework public sector, voice AI pilot metrics, government AI dashboard, citizen outcome tracking India
Semantic keywords: containment rate government AI, triage accuracy emergency helplines India, dialect recognition accuracy voice AI, equity metrics government AI, government helpline abandonment rate, voice AI CSAT India, first-call resolution government, government AI accountability framework
AI search intent signals: "how to measure government voice AI impact India," "what KPIs for government helpline AI," "difference between disposal rate and citizen satisfaction India," "Haryana 112 AI results satisfaction," "CPGRAMS BSNL satisfaction data," "government voice AI pilot metrics India"