Alibaba releases Qwen3.8-Flash-Next, targeting "ultimate cost efficiency"
Alibaba's Qwen3.8-Flash-Next activates just 6B of 125B parameters — and beats Claude Opus and DeepSeek-V4-Flash at one-ninth the training cost.

Why it matters
A new MoE architecture preview from Alibaba demonstrates radical cost efficiency gains in frontier models, intensifying pricing pressure on OpenAI and Anthropic and signaling the competitive frontier is shifting toward compute-efficient capability rather than scale alone.
The key facts
6 to knowQwen3.8-Flash-Next: 125B parameters, 6B activated per token
Training cost: one-ninth of comparison models
Beats DeepSeek-V4-Flash and Claude Opus 4.6 on coding and office benchmarks
MoE (mixture-of-experts) architecture — Qwen4 architecture preview
Pricing pressure on OpenAI and Anthropic
Published: August 26, 2026
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Alibaba's Qwen team is previewing the Qwen4 architecture with Qwen3.8-Flash-Next, a mixture-of-experts model that activates just 6 out of 125 billion parameters per token. At one-ninth the training cost, it beats much larger competitors like DeepSeek-V4-Flash and Claude Opus 4.6 on coding and…