Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution
7.8×. That's the token throughput gain Orthrus just unlocked on Qwen3—without changing a single output.

Why it matters
Open-source optimization technique demonstrates dramatic inference efficiency improvements on frontier models. Relevant for teams operating under compute constraints and evaluating cost-per-inference at scale.
The key facts
10 to know7.8× tokens/forward throughput improvement
Maintains identical output distribution (no accuracy loss)
Works on Qwen3 model
Open-source implementation (GitHub)
49 points on Hacker News (moderate traction)
7.8× tokens/forward improvement claimed on Qwen3
Identical output distribution preserved (no quality degradation)
Published as open-source GitHub repo
Early-stage community validation (49 HN points, 3 comments as of publication)
Date: May 15, 2026 (future date—verify publication authenticity)
Go to the source
Hacker Newsgithub.com
Publisher excerpt: Article URL: Comments URL: Points: 49 # Comments: 3