BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost
37.2% fewer thinking tokens. BottleCap AI's fine-tune of Qwen3.8-27B trades 0.86pp accuracy for major efficiency gains—a drop-in replacement on vLLM and SGLang.

Why it matters
Open-weight reasoning models are getting leaner. This fine-tune demonstrates how to optimize inference cost on thinking-based models without rebuilding from scratch—relevant for practitioners deploying reasoning models at scale.
The key facts
11 to knowThinkingCap-Qwen3.8-27B reduces thinking tokens by 37.2% across 12 benchmarks
Macro accuracy: 86.65% → 85.79% (0.86pp cost)
Long-context AA-LCR improves by 2.25pp
Drop-in vLLM and SGLang compatibility
Builds available: FP8, NVFP4, GGUF, MLX
Based on Qwen3.8-27B (open-weight model)
37.2% reduction in thinking tokens across 12 benchmarks
Long-context AA-LCR improves 2.25pp
Drop-in compatible with vLLM and SGLang
Available in FP8, NVFP4, GGUF, MLX builds
Based on Qwen3.8-27B open-weight model
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: BottleCap AI has released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B that spends 37.2% fewer thinking tokens across 12 benchmarks. Macro accuracy moves from 86.65% to 85.79%, and long-context AA-LCR improves by 2.25pp. The model is a drop-in replacement on vLLM and SGLang, with FP8, NVFP4,…