NVIDIA and the University of Maryland Researchers Released Audio Flamingo Next (AF-Next): A Super Powerful and Open Large Audio-Language Model
Audio just caught up to vision. NVIDIA and University of Maryland released AF-Next, an open large audio-language model that changes what's possible with speech, environmental sounds, and music reasoning.

Why it matters
Audio-language models have lagged behind vision-language models in capability and scale. AF-Next represents a significant step toward closing that gap with an open-source multimodal model, addressing a real frontier in AI reasoning that enterprises are beginning to deploy.
The key facts
6 to knowNVIDIA and University of Maryland collaboration
Audio Flamingo Next (AF-Next) released
Open-source model
Multimodal: speech, environmental sounds, music reasoning
Addresses long-context audio understanding
Audio-language models historically lagged vision-language models in scaling
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Understanding audio has always been the multimodal frontier that lags behind vision. While image-language models have rapidly scaled toward real-world deployment, building open models that robustly reason over speech, environmental sounds, and music — especially at length — has remained quite hard.…