Key Takeaways
- Achieved 92% intent accuracy on Hindi‑English code‑switching tests with 10 k utterances.
- Enabled UPI payments via voice, averaging 3 seconds per transaction.
- Reduced end‑to‑end latency to 200 ms on Jio’s 5G network.
- Optimised model size to 1.2 GB for offline use on sub‑INR 8,000 smartphones.
- Cut inference cost to INR 0.04 per call after quantisation, supporting Flipkart‑scale traffic.
What's the news
Over the past six months our team experimented with creating a voice‑first assistant that could understand mixed Hindi‑English speech, trigger a UPI payment, and reply in natural tone. We chose two recently released Indian foundation models – one strong in multilingual language understanding and the other optimised for realistic speech synthesis. The goal was to see whether locally trained models could match or exceed the performance of larger multilingual alternatives while staying lightweight enough for mass‑market devices.
Methodology
We began by gathering a corpus of 12 hours of spontaneous Hindi‑English dialogue from public forums, call‑center logs, and scripted role‑plays. The data covered varied accents from Delhi, Mumbai, Bengaluru, and Kolkata, and included common banking and shopping phrases. After cleaning, we fine‑tuned the first model on intent classification and entity extraction. The second model was adapted for text‑to‑speech, focusing on natural prosody and appropriate pausing for Indian English.
To ensure robustness we performed stratified cross‑validation across accent groups, monitoring precision and recall for each region. We also built a small held‑out set of rare lexical items to test generalisation beyond the training distribution.
Technical Deep Dive
Integration steps were straightforward: the ASR front‑end (built on an open‑source Conformer) fed transcripts to the understanding model, which returned a structured intent. Based on the intent, we either fetched information from a knowledge base or invoked the UPI SDK to initiate a payment. The TTS model then turned the response into audio, which was streamed back to the user. Throughout we kept an eye on memory footprint, applying 8‑bit quantisation and pruning to shrink the combined model to roughly 1.2 GB.
Latency measurements were taken on a range of devices: a flagship Snapdragon 8‑gen phone, a mid‑range MediaTek Helio G95, and an entry‑level Unisoc T610. On Jio’s 5G network the end‑to‑end time stayed below 200 ms on the flagship, while the mid‑range device showed ~260 ms and the entry‑level device ~340 ms. These numbers helped us set the offline fallback threshold.
India impact
Language diversity remains a major barrier for voice tech in India. By targeting code‑switching explicitly, we observed a noticeable lift in user satisfaction during internal tests – participants reported feeling “understood” more often than with a generic English‑only assistant. The ability to complete a UPI transaction by voice could lower the entry barrier for first‑time digital payers, especially in semi‑urban areas where typing on small screens is cumbersome. Moreover, the offline capability means the agent can work on Jio’s 5G‑enabled phones even when the network fluctuates, a frequent reality outside major metros.
We also ran a small field trial with 200 users in tier‑2 and tier‑3 towns. Over two weeks the average success rate for voice‑initiated payments was 88 %, with most failures linked to poor microphone quality rather than language understanding. Feedback highlighted the convenience of speaking in their natural mix of Hindi and English without having to switch languages mid‑sentence.
Use cases
We piloted three scenarios: (1) a voice‑enabled shopping assistant on a Flipkart‑like catalogue where users could say “Add two shirts under INR 1500 to cart” and confirm payment; (2) a banking helper that let users check balances and transfer money via spoken commands; (3) a government‑service guide that answered queries about PAN card applications and vaccination slots in mixed Hindi‑English. In each case, the average task completion time dropped by roughly 30 % compared to a traditional touch‑flow interface.
Additional informal tests were run with local kirana store owners who used the agent to place wholesale orders. The voice flow reduced the number of steps from six to three, and store owners reported saving about two minutes per order.
Honest take
Building on Indian models is promising but not without trade‑offs. The understanding model, while accurate on our test set, still struggled with rare regional words and heavy code‑switching patterns that appeared less frequently in the training data. The TTS voice, though natural, occasionally sounded monotone when expressing excitement or urgency. Latency gains depended heavily on the device’s chipset; older SoCs showed a noticeable increase in response time, suggesting that hardware optimisation remains a parallel challenge. Finally, maintaining the model will require a steady stream of fresh conversational data to curb drift as slang and payment habits evolve.
We also noted that quantisation, while reducing size and cost, introduced a small drop in intent accuracy – roughly 1.5 percentage points – which we compensated for by adding a lightweight post‑processing rule‑based correction layer for high‑confidence intents.
Future roadmap
Looking ahead we plan to expand the language coverage to include Tamil and Telugu code‑switching, leveraging upcoming multilingual foundation models from Indian research labs. We are also exploring on‑device federated learning techniques to continuously improve accent coverage without compromising user privacy. Another focus area is integrating the agent with regional language NPCI‑approved UPI interfaces to broaden the scope of voice‑based financial services.
Finally, we aim to publish a detailed benchmark suite that other developers can use to evaluate voice agents on Indian linguistic diversity, latency, and cost metrics, fostering a more transparent ecosystem for home‑grown AI solutions.




Comments (0)
Be the first to comment!