NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text Intelligence of Its Backbone
NVIDIA just unified audio, speech, and text in one 30B model. Here's why the MoE architecture matters for inference costs.

Why it matters
NVIDIA's Audex represents a significant capability expansion in multimodal LLMs—consolidating five audio/speech tasks into a single backbone while maintaining text performance. For enterprises, this signals a path toward unified inference endpoints that reduce deployment complexity and inference overhead.
The key facts
6 to knowModel: Nemotron-Labs-Audex-30B-A3B
Architecture: MoE (Mixture of Experts)
Capabilities: audio understanding, speech recognition, translation, TTS, audio generation
Backbone: Nemotron-Cascade-2
Text performance: marginal regression (claimed minimal loss)
Publication date: July 8, 2026
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: NVIDIA's Nemotron-Labs-Audex-30B-A3B unifies audio understanding, speech recognition, translation, TTS, and audio generation in one MoE model. It keeps the text intelligence of its Nemotron-Cascade-2 backbone with marginal regression. The post NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A…