VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO
3B parameters. Beats Claude Opus 4.5 on reasoning. Here's how VibeThinker did it.

Why it matters
A smaller, open model claims superior reasoning performance on established benchmarks using novel training techniques (SFT+GRPO), challenging the narrative that scale and closed models dominate frontier reasoning capabilities.
The key facts
11 to knowVibeThinker: 3B parameter model
Claims to beat Anthropic Claude Opus 4.5 on reasoning benchmarks
Training approach: SFT (Supervised Fine-Tuning) + GRPO (novel training method)
Published on arXiv (academic venue, not peer-reviewed yet)
Low engagement: 10 HN points, 0 comments as of publication
UNVERIFIED: Claims not independently validated; benchmark details not specified in summary
Model: VibeThinker, 3B parameters
Benchmark claim: Beats Opus 4.5 on reasoning tasks
Source: arXiv preprint (2606.16140), not peer-reviewed or industry-validated
Community signal: Low engagement (10 HN points, 0 comments) suggests limited validation
UNVERIFIED: No corroborating benchmarks, no comparison methodology details provided in headline
Go to the source
Hacker Newsarxiv.org
Publisher excerpt: Article URL: Comments URL: Points: 10 # Comments: 0