FrontierThe story, in brief

BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost

37.2% fewer thinking tokens. BottleCap AI's fine-tune of Qwen3.8-27B trades 0.86pp accuracy for major efficiency gains—a drop-in replacement on vLLM and SGLang.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Open-weight reasoning models are getting leaner. This fine-tune demonstrates how to optimize inference cost on thinking-based models without rebuilding from scratch—relevant for practitioners deploying reasoning models at scale.

The key facts

11 to know
  1. ThinkingCap-Qwen3.8-27B reduces thinking tokens by 37.2% across 12 benchmarks

  2. Macro accuracy: 86.65% → 85.79% (0.86pp cost)

  3. Long-context AA-LCR improves by 2.25pp

  4. Drop-in vLLM and SGLang compatibility

  5. Builds available: FP8, NVFP4, GGUF, MLX

  6. Based on Qwen3.8-27B (open-weight model)

  7. 37.2% reduction in thinking tokens across 12 benchmarks

  8. Long-context AA-LCR improves 2.25pp

  9. Drop-in compatible with vLLM and SGLang

  10. Available in FP8, NVFP4, GGUF, MLX builds

  11. Based on Qwen3.8-27B open-weight model

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: BottleCap AI has released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B that spends 37.2% fewer thinking tokens across 12 benchmarks. Macro accuracy moves from 86.65% to 85.79%, and long-context AA-LCR improves by 2.25pp. The model is a drop-in replacement on vLLM and SGLang, with FP8, NVFP4,…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier