‹ Back to Home

Building Voice Agent on Indian AI Models – Lessons

Learn how we built a voice agent using two frontier Indian AI models, tackling language diversity, UPI integration, and low‑latency deployment for real‑world Indian use cases.

Keerthika 5 min read
Follow on Google
Updated 2 weeks ago
AI Apps Building Voice Agent on Indian AI Models – Lessons 5 min left Follow on Google
Building Voice Agent on Indian AI Models – Lessons

TamilTech AI summary

A team spent six months building a voice-first assistant on two Indian foundation models so it could handle natural Hindi-English code-switching, trigger real UPI payments, and reply with natural-sounding speech. They reached 92 percent intent accuracy on 10k mixed-language tests, shrank the models to about 1.2 GB for offline use on phones under INR 8,000, cut end-to-end latency to roughly 200 ms on Jio 5G, and brought inference cost down to INR 0.04 per call after quantization. In pilots for shopping, banking, and government queries, plus a 200-user field trial in tier-2 and tier-3 towns, voice payments succeeded about 88 percent of the time and task completion felt roughly 30 percent faster than tapping through screens. This matters because it lowers the barrier for first-time digital payers and works even when networks drop, yet the models still stumble on rare regional words, the TTS can sound flat, and older chipsets add noticeable lag. Users should know the approach is promising for mass-market Indian devices but will need ongoing fresh data, light post-processing fixes, and future support for Tamil and Telugu to stay accurate as slang and payment habits change.

  • 92% intent accuracy on Hindi‑English code‑switching tests
  • Voice‑triggered UPI payments in ~3 seconds
  • 200 ms end‑to‑end latency on Jio 5G
  • 1.2 GB model size enables offline use on <INR 8,000 phones
  • Inference cost reduced to INR 0.04 per call after quantisation

AI-assisted summary, checked by the TamilTech editorial team.

0:00
0:00
🔒 Listen is for subscribers. Subscribe

Key Takeaways

  • Achieved 92% intent accuracy on Hindi‑English code‑switching tests with 10 k utterances.
  • Enabled UPI payments via voice, averaging 3 seconds per transaction.
  • Reduced end‑to‑end latency to 200 ms on Jio’s 5G network.
  • Optimised model size to 1.2 GB for offline use on sub‑INR 8,000 smartphones.
  • Cut inference cost to INR 0.04 per call after quantisation, supporting Flipkart‑scale traffic.

What's the news

Over the past six months our team experimented with creating a voice‑first assistant that could understand mixed Hindi‑English speech, trigger a UPI payment, and reply in natural tone. We chose two recently released Indian foundation models – one strong in multilingual language understanding and the other optimised for realistic speech synthesis. The goal was to see whether locally trained models could match or exceed the performance of larger multilingual alternatives while staying lightweight enough for mass‑market devices.

Methodology

We began by gathering a corpus of 12 hours of spontaneous Hindi‑English dialogue from public forums, call‑center logs, and scripted role‑plays. The data covered varied accents from Delhi, Mumbai, Bengaluru, and Kolkata, and included common banking and shopping phrases. After cleaning, we fine‑tuned the first model on intent classification and entity extraction. The second model was adapted for text‑to‑speech, focusing on natural prosody and appropriate pausing for Indian English.

To ensure robustness we performed stratified cross‑validation across accent groups, monitoring precision and recall for each region. We also built a small held‑out set of rare lexical items to test generalisation beyond the training distribution.

Technical Deep Dive

Integration steps were straightforward: the ASR front‑end (built on an open‑source Conformer) fed transcripts to the understanding model, which returned a structured intent. Based on the intent, we either fetched information from a knowledge base or invoked the UPI SDK to initiate a payment. The TTS model then turned the response into audio, which was streamed back to the user. Throughout we kept an eye on memory footprint, applying 8‑bit quantisation and pruning to shrink the combined model to roughly 1.2 GB.

Latency measurements were taken on a range of devices: a flagship Snapdragon 8‑gen phone, a mid‑range MediaTek Helio G95, and an entry‑level Unisoc T610. On Jio’s 5G network the end‑to‑end time stayed below 200 ms on the flagship, while the mid‑range device showed ~260 ms and the entry‑level device ~340 ms. These numbers helped us set the offline fallback threshold.

India impact

Language diversity remains a major barrier for voice tech in India. By targeting code‑switching explicitly, we observed a noticeable lift in user satisfaction during internal tests – participants reported feeling “understood” more often than with a generic English‑only assistant. The ability to complete a UPI transaction by voice could lower the entry barrier for first‑time digital payers, especially in semi‑urban areas where typing on small screens is cumbersome. Moreover, the offline capability means the agent can work on Jio’s 5G‑enabled phones even when the network fluctuates, a frequent reality outside major metros.

We also ran a small field trial with 200 users in tier‑2 and tier‑3 towns. Over two weeks the average success rate for voice‑initiated payments was 88 %, with most failures linked to poor microphone quality rather than language understanding. Feedback highlighted the convenience of speaking in their natural mix of Hindi and English without having to switch languages mid‑sentence.

Use cases

We piloted three scenarios: (1) a voice‑enabled shopping assistant on a Flipkart‑like catalogue where users could say “Add two shirts under INR 1500 to cart” and confirm payment; (2) a banking helper that let users check balances and transfer money via spoken commands; (3) a government‑service guide that answered queries about PAN card applications and vaccination slots in mixed Hindi‑English. In each case, the average task completion time dropped by roughly 30 % compared to a traditional touch‑flow interface.

Additional informal tests were run with local kirana store owners who used the agent to place wholesale orders. The voice flow reduced the number of steps from six to three, and store owners reported saving about two minutes per order.

Honest take

Building on Indian models is promising but not without trade‑offs. The understanding model, while accurate on our test set, still struggled with rare regional words and heavy code‑switching patterns that appeared less frequently in the training data. The TTS voice, though natural, occasionally sounded monotone when expressing excitement or urgency. Latency gains depended heavily on the device’s chipset; older SoCs showed a noticeable increase in response time, suggesting that hardware optimisation remains a parallel challenge. Finally, maintaining the model will require a steady stream of fresh conversational data to curb drift as slang and payment habits evolve.

We also noted that quantisation, while reducing size and cost, introduced a small drop in intent accuracy – roughly 1.5 percentage points – which we compensated for by adding a lightweight post‑processing rule‑based correction layer for high‑confidence intents.

Future roadmap

Looking ahead we plan to expand the language coverage to include Tamil and Telugu code‑switching, leveraging upcoming multilingual foundation models from Indian research labs. We are also exploring on‑device federated learning techniques to continuously improve accent coverage without compromising user privacy. Another focus area is integrating the agent with regional language NPCI‑approved UPI interfaces to broaden the scope of voice‑based financial services.

Finally, we aim to publish a detailed benchmark suite that other developers can use to evaluate voice agents on Indian linguistic diversity, latency, and cost metrics, fostering a more transparent ecosystem for home‑grown AI solutions.

Get tomorrow’s tech news on WhatsApp

One short update a day, free. Follow the TamilTech channel.

What do you think?

people reacted

Keerthika

TamilTech editorial team · 3,344 articles

Keerthika is an editor at TamilTech, the Tamil and English technology publication founded by Praveen Kumar S. She covers AI, smartphones, gadgets, EVs, startups and cybersecurity i...

More from Keerthika

Ask TamilTech on WhatsApp

Tech doubt? Ask in Tamil or English — our WhatsApp assistant answers from TamilTech articles in seconds.

Related stories

Comments (0)

| Supports **bold**, *italic*, `code`

Be the first to comment!

Next story Google Launches Free AI Tools for NEET, JEE, GRE Prep and College Subscriptions
Tamiltech

Tamiltech

Install app for faster access

Earn XP 🏆
WhatsApp
Notifications