FrontierThe story, in brief

Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English

Sarvam AI's Saaras V4 covers all 22 Indian languages in one model. First-token latency under 150ms. ₹30/hour.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A multilingual speech-to-text model targeting underserved language communities with practical streaming performance and keyterm prompting. Relevant to practitioners building voice interfaces for Indian markets and enterprises evaluating non-English ASR coverage.

The key facts

15 to know
  1. Covers all 22 Indian languages plus global English

  2. 3B hybrid state-space decoder architecture

  3. Keyterm prompting: up to 50 custom terms per request

  4. 5 output modes from single model

  5. First-token latency: <150ms for streaming

  6. Pricing: ₹30 per hour via Sarvam API

  7. GA today via API

  8. Covers all 22 Indian languages plus global English in single model

  9. Architecture: audio encoder + 3B hybrid state-space decoder

  10. Keyterm prompting: up to 50 terms per request

  11. 5 output modes from 1 model (streaming, batch, transcription variants implied)

  12. First-token latency: <150ms

  13. Availability: API via Sarvam, ₹30 per hour pricing

  14. Status: Available today (as of Sep 26, 2026)

  15. No independent benchmark data vs. competitors provided

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Sarvam AI's Saaras V4 is a speech-to-text model covering all 22 Indian languages plus global English. It pairs an audio encoder with a 3B hybrid state-space decoder. It adds keyterm prompting for up to 50 terms, 5 output modes from 1 model, and streaming with first-token latency under 150 ms. It is…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier