Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Model Outputs
3.18x faster decoding. Liquid AI's draft models show speculative decoding can scale without model changes.

Why it matters
Speculative decoding—using smaller draft models to speed inference—is moving from research to production. Liquid AI's DSpark drafters maintain identical outputs while cutting latency, a key efficiency win for practitioners deploying LFM2.5 at scale.
The key facts
12 to knowLiquid AI released LFM2.5-DSpark draft models
Three ~300M parameter drafters
Up to 3.18x faster decoding reported
Identical greedy output—no model behavior changes
Speculative decoding technique
Published August 20, 2026
3.18x faster decoding measured
Three ~300M parameter draft models released
Identical greedy outputs (no quality loss)
Technique: speculative decoding
Model: LFM2.5 (Liquid Foundational Model 2.5)
Published: August 20, 2026
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Three ~300M drafters bring speculative decoding to LFM2.5, delivering up to 3.18x faster decoding with identical greedy output.