FrontierAugust 1, 2026via The Decoder
ByteDance's Seedance 2.5 generates 30-second video clips with built-in audio
Why it matters
A multimodal model capability leap: native audio-video generation and 3x context length vs. the frontier competition signals where the lab race is moving. Practitioners building video tools need to reset their baseline.
Key signals
- Seedance 2.5 produces up to 30-second video clips with integrated audio generation
- 3x longer output than Google's Gemini Omni Flash
- Accepts dozens of images, videos, and audio files as reference inputs
- Native audio-video generation (not post-processed)
- ByteDance product — direct competition with Google, OpenAI multimodal capabilities
The hook
ByteDance's Seedance 2.5 generates 30-second video with native audio — 3x longer than Gemini Omni Flash.
ByteDance just shipped Seedance 2.5, an AI video model that produces video and audio together in one go. Each clip runs up to 30 seconds, three times what Google's Gemini Omni Flash puts out. Users can feed in dozens of images, videos, and audio files as reference. For ad teams, this could kill the …