FrontierSeptember 19, 2026via The Decoder
Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks
Why it matters
A new multimodal model from Alibaba's Qwen achieves near-parity with Google's flagship on audio-video tasks while undercutting its pricing, intensifying competition in the agent-native model race and raising questions about Google's cost positioning.
Key signals
- Qwen3.8-Omni-Flash is Qwen's first multimodal model designed for AI agents
- Processes audio and video together natively
- Performs tool use independently (editing vlogs, translating clips, summarizing movies)
- Nearly matches Gemini 3.8 Flash on audio-video benchmarks
- Significantly lower API pricing than Gemini Flash
- Multimodal capability comparison: audio + video processing
The hook
Qwen3.8-Omni-Flash matches Gemini Flash on multimodal benchmarks at a fraction of the API cost—another sign the frontier is flattening.
Qwen3.8-Omni-Flash is Qwen's first multimodal model designed for AI agents. It processes audio and video together and independently uses tools to edit vlogs, translate clips, or summarize movies. On audio-video benchmarks, it nearly matches Gemini 3.8 Flash at a fraction of the API cost.