FrontierThe story, in brief

NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text Intelligence of Its Backbone

NVIDIA just unified audio, speech, and text in one 30B model. Here's why the MoE architecture matters for inference costs.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

NVIDIA's Audex represents a significant capability expansion in multimodal LLMs—consolidating five audio/speech tasks into a single backbone while maintaining text performance. For enterprises, this signals a path toward unified inference endpoints that reduce deployment complexity and inference overhead.

The key facts

6 to know
  1. Model: Nemotron-Labs-Audex-30B-A3B

  2. Architecture: MoE (Mixture of Experts)

  3. Capabilities: audio understanding, speech recognition, translation, TTS, audio generation

  4. Backbone: Nemotron-Cascade-2

  5. Text performance: marginal regression (claimed minimal loss)

  6. Publication date: July 8, 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: NVIDIA's Nemotron-Labs-Audex-30B-A3B unifies audio understanding, speech recognition, translation, TTS, and audio generation in one MoE model. It keeps the text intelligence of its Nemotron-Cascade-2 backbone with marginal regression. The post NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier