Bonsai 27B is a full open reasoning model that fits on an iPhone
27B parameters. Under 4GB. Running on iPhone. PrismML just compressed away the on-device AI gap.

Why it matters
Model compression is becoming a competitive capability battleground. PrismML's Bonsai 27B demonstrates that reasoning-class models can run locally with minimal performance loss, forcing a reckoning on where inference happens—and who controls the hardware-to-model fit.
The key facts
6 to knowBonsai 27B compressed to under 4GB from 27B parameters
Maintains 90% of original performance in compression
Math and coding benchmarks barely affected by compression
Apple reportedly testing the technology
Enables full reasoning models to run natively on iPhone
Compression approach could shift competitive dynamics in on-device AI
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: PrismML has compressed a 27-billion-parameter AI model to under 4 GB, small enough to run on an iPhone. In the company's own benchmarks, the smallest version keeps 90 percent of the original performance, with math and coding scores barely affected. Apple is reportedly already testing the…