The Agent RaceJune 28, 2026via MarkTechPost

Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inference

Why it matters

Liquid AI's smallest model demonstrates that parameter count isn't destiny—efficiency-focused architectures can beat larger competitors on edge devices, reshaping the on-device inference landscape for founders building latency-critical, privacy-first applications.

Key signals

  • LFM2.5-230M is open-weight
  • 213 tok/s on Galaxy S25 Ultra
  • 42 tok/s on Raspberry Pi 5
  • Beats Qwen3.5-0.8B on instruction following
  • Beats Gemma 3 1B on instruction following
  • Built on LFM2 architecture
  • Supports llama.cpp, MLX, vLLM, SGLang, ONNX
  • Targets tool use and data extraction

The hook

230M parameters. That's all Liquid AI needs to outperform Qwen and Gemma on a Raspberry Pi.

Liquid AI released LFM2.5-230M, its smallest model yet. The 230M-parameter, open-weight model runs on-device at 213 tok/s on a Galaxy S25 Ultra and 42 on a Raspberry Pi 5. Built on the LFM2 architecture, it targets tool use and data extraction, beating larger models like Qwen3.5-0.8B and Gemma 3 1B on instruction following.

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.