Alibaba's Qwen team makes AI models think deeper with new algorithm
Alibaba's Qwen team just cracked the reasoning bottleneck that's been holding back AI models.

Why it matters
This algorithmic breakthrough addresses a fundamental limitation in how AI models learn to reason, potentially giving Alibaba's Qwen models a significant competitive advantage in complex reasoning tasks.
The key facts
3 to knowDoubles the length of thought processes
New algorithm weights each step based on downstream impact
Addresses reinforcement learning limitation where every token gets same reward
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Reinforcement learning hits a wall with reasoning models because every token gets the same reward. A new algorithm from Alibaba's Qwen team fixes this by weighting each step based on how much it shapes what comes next, doubling the length of thought processes in the process.