Edited by Editor-in-Chief, The Indus Pulse 25 Sept 2026, 01:39 PM 3 min readai

Sarvam AI Launches Saaras V4 Speech Model Supporting 22 Indian Languages

Bangalore-based artificial intelligence firm Sarvam AI has released Saaras V4, an automatic speech recognition model engineered to support 22 Indian languages alongside English. The system combines an audio encoder with a 3-billion-parameter hybrid state-space language model trained in-house from scratch to process complex acoustic environments, including background noise and regional dialect variations.
The enterprise is backed by forty-one million dollars in Series A funding led by Lightspeed, alongside contributions from Peak XV Partners and Khosla Ventures. Meanwhile, the Indian banking, financial services, and insurance chatbot and voice artificial intelligence market anticipates expansion from approximately one hundred ten million dollars in 2022 to approximately five hundred fifty million dollars by 2030.
Designed for native voice applications across the subcontinent, the model introduces five integrated output modes capable of converting a single audio input into verbatim transcription, normalized text, code-mixed script, transliteration, and English translation without separate post-processing pipelines.

Unified Five-Mode Output Architecture

Sarvam AI reported that Saaras V4 manages all five transcription formats directly within the core model framework. According to the company, processing requirements are handled internally to avoid external preprocessing stages that typically trigger cascading errors across multi-step pipelines.
The system features verbatim mode to capture speech exactly as spoken, a transcription mode that formats native scripts while standardizing dates and numbers, and a code-mixed mode that preserves English terminology within vernacular speech. Additional formats include transliteration into Roman script and direct English translation.

Error Rates and Benchmark Performance

To evaluate accuracy, Sarvam AI utilized verified IndicVoices data, recording a language-identification error rate of 5.22 per cent across all 22 supported Indian languages. Across the 10 most widely spoken regional languages, the error rate dropped to 2.9 per cent.
For regional evaluations, the company tested Saaras V4 on the Vistaar benchmark across 10 languages using both standard Word Error Rate and LLM-WER, a metric that incorporates semantic assessments to verify whether transcription variances alter spoken meaning. In English evaluations spanning seven benchmarks, including AMI for meetings and GigaSpeech for multimedia, the model achieved the lowest average word error rate.

Low-Latency Streaming and Developer Access

Saaras V4 supports real-time WebSocket streaming and long-form audio processing capable of handling multi-minute recordings in under a second. Sarvam AI reported a time-to-first-token metric below 150 milliseconds for live streaming deployments.
The model is accessible through Sarvam AI's API utilizing Python and Node.js software development kits. Integration support includes frameworks such as the Vercel AI SDK, LiveKit Agents, and Pipecat Agents, accommodating synchronous transcription, batch processing for files up to two hours, and live streaming.
The Indus Pulse is committed to accuracy and transparency.