Interestana
Home/News/Sarvam AI Releases Saaras V4 Speech Model
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Sarvam AI Releases Saaras V4 Speech Model

Sarvam AI has released Saaras V4, the latest iteration of its speech recognition model, which now supports all 22 scheduled Indian languages and global English accents. The company reports that Saaras V4 achieves state-of-the-art accuracy across these languages. The model is available for deployment via Sarvam's API using the model identifier "saaras:v4". While the model weights are not publicly available, Sarvam's documentation for self-hosting currently covers the previous version, Saaras v3.

Saaras V4 is built on an encoder-decoder architecture. The audio encoder processes the waveform into embeddings that capture phonetic and acoustic details. A temporal-downsampling adapter then reduces the sequence length and projects it into the language model's embedding space, enabling the decoder to handle longer audio recordings within its context budget. The decoder component is Sarvam-3B, a 3 billion parameter hybrid state-space language model developed in-house. This language model processes audio features in conjunction with a text prompt and generates the transcript autoregressively, feeding each generated token back as input for the subsequent token.

In benchmark evaluations for English, Saaras V4 was tested on seven datasets, including six from Hugging Face's Open ASR Leaderboard: AMI, GigaSpeech, LibriSpeech clean, LibriSpeech other, SPGISpeech, and VoxPopuli. The seventh dataset, AI4Bharat's Indian-accented Svarah, was also included. The scoring methodology followed the normalization code used on the leaderboard. Sarvam reported that Saaras V4 achieved the lowest average Word Error Rate (WER) among the models it benchmarked on these English datasets.

For Indic languages, Sarvam reported results on the Vistaar dataset across 10 Indian languages, utilizing both WER and LLM-WER metrics. The LLM-WER metric incorporates a semantic check to differentiate between genuine meaning errors and minor variations in spelling or formatting that are common in Indic scripts. On the Kathbath Noisy dataset, which features compressed, clipped, and background-heavy audio recordings, Saaras V4 demonstrated an error rate less than half that of Deepgram Nova-3 and GPT-4o Transcribe, as measured by LLM-WER. Furthermore, in language identification tests on verified IndicVoices utterances, Saaras V4 exhibited a language identification error rate of 2.9% across the top 10 Indian languages, and 5.22% for a broader set of languages.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next