FrontierThe story, in brief

Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

Alibaba's Qwen3.8-Flash-Next: 125B multimodal MoE, 6B active params, 1/9 the training cost of Qwen3.7-Plus. The Qwen4 architecture preview is here.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Open-weight model release with novel MoE and attention architecture, significant training efficiency gains, and concrete multimodal capabilities. This is Alibaba signaling architectural parity with frontier labs while competing on efficiency — practitioners need to benchmark this against Claude/GPT equivalents.

The key facts

7 to know
  1. 125B backbone + 51B N-gram embedding + 4B multi-token prediction = 180B total parameters

  2. 6B active parameters per token

  3. 1/9 training cost vs. Qwen3.7-Plus

  4. 172.78 GiB FP8 checkpoint size (self-hosting requirement)

  5. Four architectural innovations: Gated DeltaNet + Qwen Sparse Attention hybrid, Gated Residual, N-gram Embedding, Muon optimizer

  6. Multimodal capability confirmed

  7. Open-weight release (benchmark results mentioned but specific scores not extracted)

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier