Apple in talks with startup that shrinks AI models to run on an iPhone - CNBC
Not a lab demo. Apple is in talks to shrink 27B parameter models down to iPhone hardware—on-device AI just got real.

Why it matters
Apple's pursuit of on-device model compression represents a critical infrastructure shift: moving AI inference from cloud to edge reduces latency, privacy exposure, and compute dependency. This signals the smartphone market is ready for lightweight, deployable foundation models.
The key facts
5 to knowPrismML's Bonsai 27B: 1-bit and ternary quantization of Qwen 3.6-27B
Model claimed to run on iPhones and laptops—largest model of its size on mobile hardware
Apple in active talks with PrismML (Khosla-backed)
Implication: on-device inference without cloud dependency
Use case: privacy-first, low-latency AI on consumer devices
Go to the source
Reuters Technologynews.google.com
Publisher excerpt: Apple in talks with startup that shrinks AI models to run on an iPhone CNBC Khosla-Backed Startup Claims Breakthrough With Largest-Ever AI Model on an iPhone The Information Apple looks to shrink AI models for iPhones Baton Rouge Business Report PrismML Releases Bonsai 27B: 1-bit and Ternary Builds…