Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English
Sarvam AI's Saaras V4 covers all 22 Indian languages in one model. First-token latency under 150ms. ₹30/hour.

Why it matters
A multilingual speech-to-text model targeting underserved language communities with practical streaming performance and keyterm prompting. Relevant to practitioners building voice interfaces for Indian markets and enterprises evaluating non-English ASR coverage.
The key facts
15 to knowCovers all 22 Indian languages plus global English
3B hybrid state-space decoder architecture
Keyterm prompting: up to 50 custom terms per request
5 output modes from single model
First-token latency: <150ms for streaming
Pricing: ₹30 per hour via Sarvam API
GA today via API
Covers all 22 Indian languages plus global English in single model
Architecture: audio encoder + 3B hybrid state-space decoder
Keyterm prompting: up to 50 terms per request
5 output modes from 1 model (streaming, batch, transcription variants implied)
First-token latency: <150ms
Availability: API via Sarvam, ₹30 per hour pricing
Status: Available today (as of Sep 26, 2026)
No independent benchmark data vs. competitors provided
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Sarvam AI's Saaras V4 is a speech-to-text model covering all 22 Indian languages plus global English. It pairs an audio encoder with a 3B hybrid state-space decoder. It adds keyterm prompting for up to 50 terms, 5 output modes from 1 model, and streaming with first-token latency under 150 ms. It is…