The Agent RaceJune 28, 2026via MarkTechPost
Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inference
Why it matters
Liquid AI's smallest model demonstrates that parameter count isn't destiny—efficiency-focused architectures can beat larger competitors on edge devices, reshaping the on-device inference landscape for founders building latency-critical, privacy-first applications.
Key signals
- LFM2.5-230M is open-weight
- 213 tok/s on Galaxy S25 Ultra
- 42 tok/s on Raspberry Pi 5
- Beats Qwen3.5-0.8B on instruction following
- Beats Gemma 3 1B on instruction following
- Built on LFM2 architecture
- Supports llama.cpp, MLX, vLLM, SGLang, ONNX
- Targets tool use and data extraction
The hook
230M parameters. That's all Liquid AI needs to outperform Qwen and Gemma on a Raspberry Pi.
Liquid AI released LFM2.5-230M, its smallest model yet. The 230M-parameter, open-weight model runs on-device at 213 tok/s on a Galaxy S25 Ultra and 42 on a Raspberry Pi 5. Built on the LFM2 architecture, it targets tool use and data extraction, beating larger models like Qwen3.5-0.8B and Gemma 3 1B on instruction following.