The Agent RaceJuly 23, 2026via The Decoder
Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs
Why it matters
Black Forest Labs shipped a multimodal foundation model that generates video with synchronized audio for the first time, positioning itself ahead of Sora 2.0 in capability benchmarks. The move signals a shift toward coherent video-audio synthesis and early robotics applications—territory that will reshape both the creative tools and embodied AI markets.
Key signals
- Flux 3 generates native audio video up to 20 seconds
- Multimodal foundation model trained on images, video, and audio
- Internal benchmarks show Flux 3 ahead of Sora 2.0
- Independent benchmark results not yet available
- Black Forest Labs testing Flux 3 on robotics tasks
- Company stated long-term goal: build a world model
The hook
Native audio in video generation. Flux 3 just changed the game—and it's already beating Sora's closest competitor.
Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can generate video with native sound for the first time. BFL's own tests put it just ahead of market leader Seedance 2.0, though independent results aren't yet available. The company ultimately wants to build a world model and is already testing Flux 3 on robotics tasks.