Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Bonsai 2 27B achieves 9x compression with near-lossless performance — what this means for edge deployment and inference costs.

Why it matters
A new distilled model demonstrates that frontier-scale capability can fit in a fraction of the usual footprint without measurable quality loss. Practitioners deploying on-device or latency-sensitive workflows gain a materially smaller target; the compression technique itself advances the lab race around efficient model design.
The key facts
5 to knowBonsai 2 27B achieves ~9x compression ratio
Claimed near-lossless performance retention
Model size reduction to 27B parameters from an implied larger parent
Relevance to edge deployment, inference latency, and compute cost
Published Sep 17, 2026 via Simon Willison (secondary reporting)
Go to the source
Simon Willisonsimonwillison.net