Qwen3.8 27B addition in words
Alibaba's Qwen3.8 27B model now handles multi-token reasoning in a single forward pass—a capability shift that changes how open-weight models compete on inference speed.

Why it matters
Qwen3.8 27B adds native multi-token generation ('addition in words'), reducing inference latency for open-weight deployments. This is a meaningful capability gap-closer against proprietary reasoning models, with direct implications for enterprise inference budgets and on-prem deployment economics.
The key facts
6 to knowQwen3.8 27B parameter model
Multi-token generation capability added
Open-weight release (implied)
Published by Alibaba via Simon Willison's blog post
Inference architecture / capability addition (not a benchmark claim)
Likely reduces per-token latency in reasoning workloads
Go to the source
Simon Willisonsimonwillison.net