Google LiteRT-LM Speeds Up Local Inference Up to 2.2x With Gemma 4 Multi-Token Prediction
2.2x faster. Google's LiteRT-LM just shipped native Gemma 4 Multi-Token Prediction—making on-device AI inference dramatically cheaper to run.

Why it matters
Google is expanding LiteRT-LM with Gemma 4 MTP support and new language bindings (Swift, JavaScript), directly enabling developers to deploy faster, more efficient local inference. This matters because on-device AI performance directly impacts app latency and battery life—key bottlenecks for consumer and enterprise adoption.
The key facts
4 to knowLiteRT-LM achieves up to 2.2x faster inference with Gemma 4 Multi-Token Prediction drafters
Native MTP drafter support now included in framework
New API support: Swift and JavaScript (in addition to existing Kotlin and C++)
Published June 5, 2026
Go to the source
InfoQ AI/MLinfoq.com
Publisher excerpt: LiteRT-LM brings native support for Gemma 4 Multi-Token Prediction (MTP) drafters, enabling up to 2.2x faster inference. The framework is expanding beyond Kotlin and C++ adding support for new Swift and a JavaScript APIs. By Sergio De Simone