FrontierThe story, in brief

Interfaze Ships diffusion-gemma-asr-small, an Open-Source Diffusion ASR Model Transcribing Six Languages via DiffusionGemma’s Parallel Denoising Decoder

Interfaze just shipped a diffusion-based ASR model that does what autoregressive can't: transcribe six languages with a single 42M adapter.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Open-source diffusion ASR challenges the autoregressive paradigm for speech recognition, offering a novel inference approach (denoising steps vs. transcript length) that could reshape how companies think about transcription cost structures and multilingual deployment.

The key facts

10 to know
  1. Interfaze open-sourced diffusion-gemma-asr-small

  2. Uses diffusion-based decoding instead of autoregressive generation

  3. ~42M-parameter adapter on frozen Google DiffusionGemma

  4. Single adapter covers six languages

  5. Transcription cost determined by denoising steps, not transcript length

  6. Multilingual capability with unified model architecture

  7. 42M-parameter adapter architecture

  8. Covers 6 languages with single adapter

  9. Built on Google's frozen DiffusionGemma base

  10. Cost scales with denoising steps, not transcript length

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Interfaze open-sourced diffusion-gemma-asr-small, a multilingual ASR model that transcribes via diffusion, not autoregression. It adds audio to Google's frozen DiffusionGemma using a ~42M-parameter adapter. One adapter covers six languages, with transcription cost set by denoising steps, not…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier