Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inference
230M parameters. That's all Liquid AI needs to outperform Qwen and Gemma on a Raspberry Pi.

Why it matters
Liquid AI's smallest model demonstrates that parameter count isn't destiny—efficiency-focused architectures can beat larger competitors on edge devices, reshaping the on-device inference landscape for founders building latency-critical, privacy-first applications.
The key facts
8 to knowLFM2.5-230M is open-weight
213 tok/s on Galaxy S25 Ultra
42 tok/s on Raspberry Pi 5
Beats Qwen3.5-0.8B on instruction following
Beats Gemma 3 1B on instruction following
Built on LFM2 architecture
Supports llama.cpp, MLX, vLLM, SGLang, ONNX
Targets tool use and data extraction
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Liquid AI released LFM2.5-230M, its smallest model yet. The 230M-parameter, open-weight model runs on-device at 213 tok/s on a Galaxy S25 Ultra and 42 on a Raspberry Pi 5. Built on the LFM2 architecture, it targets tool use and data extraction, beating larger models like Qwen3.5-0.8B and Gemma 3 1B…