Liquid AI has released speculative decoding draft models for its LFM2.5 family, boosting inference speed up to 3.18x on high-end GPUs and 2.87x on Apple Silicon laptops. Weights are free on Hugging Face, with day-one support for llama.cpp and SGLang.
Liquid AI has released a set of draft model checkpoints for its LFM2.5 model family that dramatically accelerate inference speed without changing the quality of outputs. The new models — branded DSpark — are available free on Hugging Face for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B, and come with day-one support for two of the most widely used open-source inference frameworks: llama.cpp and SGLang.
The headline numbers are hard to ignore. On a single H100 80GB GPU, throughput improves by up to 3.18x. On an M4 Max MacBook Pro running local inference through llama.cpp with Metal acceleration, speedup reaches up to 2.87x. Both figures are achieved with no change to the model’s actual outputs — the text generated with DSpark enabled is mathematically identical to what you’d get running the base model alone.
How Speculative Decoding Actually Works
The core trick is called speculative decoding, and the intuition is straightforward once you understand where LLM inference time goes. Generating tokens one at a time is slow not because the math is hard, but because each step requires loading the model’s weights from memory — a bandwidth-bound bottleneck. The bigger the model, the more weight traffic and the slower each token comes out.
DSpark sidesteps this by pairing the large target model with a small, fast companion called a draft model. The draft model — roughly 300 million parameters in Liquid AI’s implementation — rapidly proposes a block of candidate tokens in a single forward pass. The larger target model then verifies the entire block at once, accepting tokens that match its own distribution and discarding ones that don’t. Because the verification happens in one pass instead of token-by-token, the cost of loading weights gets spread across many tokens simultaneously. The result is more tokens per second at the same hardware cost.
DSpark itself is a recently published technique from researchers at Peking University and DeepSeek, building on two earlier speculative decoding approaches. Autoregressive drafters like EAGLE-3 generate each candidate conditioned on the previous one — accurate but slow, since the process is still sequential. Parallel drafters like DFlash generate the entire block in one shot, which is faster, but each token ignores the others in the block, causing the tail to drift and get rejected. DSpark combines both ideas: a parallel backbone produces hidden states for all draft tokens in one pass, a lightweight sequential head adds inter-token dependency to raise acceptance rates at later positions, and a confidence-scheduled verifier prunes the block when verification would cost more than it saves. Compared to DFlash alone, DSpark improves average accepted token length by 16 to 18 percent on comparable target models.
DeepSeek already applied DSpark to its own production models and open-sourced the technique. Liquid AI is the first to release DSpark-compatible draft checkpoints as open weights for a separate model family, lowering the barrier for anyone who wants to run it themselves.
The Caveats Matter
Speedup with speculative decoding is not a flat multiplier — it tracks the acceptance rate of the draft model on a given task, not the size of the target model. That variance is significant. The LFM2.5-8B-A1B model, for example, accepts 8.27 out of 10 proposed tokens per step on MATH500 but only 4.02 on GSM8K, which is the difference between a 3.18x speedup and a 1.29x speedup on the same hardware running the same model. Writers and developers should be careful: the headline figure is a best-case result on favorable benchmarks, not a guaranteed baseline across all workloads.
The 1.2B-Instruct model shows similar variance, with speedup swinging by as much as 52% depending on the text distribution. The 8B-A1B model also underperforms expectations on Apple Silicon despite a solid acceptance rate, because verifying a block of tokens activates more experts in the mixture-of-experts architecture, increasing weight traffic rather than reducing it. That’s a known limitation of the current MoE implementation in llama.cpp’s Metal backend, not a flaw in the technique itself.
The standout result is the 2.6B model in function-calling scenarios. Across multi-tool benchmarks, DSpark cuts latency by 57% on average for LFM2.5-2.6B — and on a MacBook, that model pushes throughput above roughly 140 tokens per second, which Liquid AI says exceeds what most proprietary cloud APIs deliver to end users.
Why This Is Relevant Right Now
The local inference ecosystem has matured quickly in 2026. MLX is stable, Ollama ships an MLX backend, and llama.cpp runs efficiently on a wide range of consumer hardware. But almost none of the popular open-weight models running locally today have speculative decoding draft companions. Liquid AI’s release is notable partly because it ships in GGUF format with llama.cpp support on day one. GGUF has become the de facto standard for distributable quantized models — when a new model drops on Hugging Face, the first versions that appear are almost always GGUF. By meeting developers where they already are, Liquid AI dramatically lowers the setup cost compared to a framework-specific integration that requires migration.
What It Means for Students
The practical implications for students are specific and worth spelling out. Because inference runs locally, data never leaves the device and each run costs nothing beyond electricity. That changes the economics of iterative experimentation — running a model 500 times to test a research hypothesis or debug an agentic pipeline is free, rather than a bill that accumulates against a student API budget.
The LFM2.5-2.6B model is the one to watch here. Liquid AI describes it as capable of planning, calling tools, and working through multi-step tasks on phones, laptops and embedded hardware. With DSpark cutting function-calling latency by 57%, students building tool-using agents as capstone or portfolio projects can now run those pipelines at interactive speeds on a MacBook without cloud round-trips.
For students entering the AI job market, the ability to demonstrate a locally deployed, speculative-decoding-accelerated inference setup signals something that a ChatGPT wrapper does not: hands-on systems knowledge of how inference actually works. Being able to articulate why draft models speed up token generation, how acceptance rates govern real-world speedup, and how to configure llama.cpp or SGLang for speculative decoding is the kind of detail that distinguishes a candidate who has built things from one who has only called APIs.
All three draft model checkpoints — in both Safetensors and GGUF formats — are available now on Hugging Face. The full release details, benchmark tables and setup instructions are published in Liquid AI’s announcement on Hugging Face.
Source: Liquid AI
Additional research sources
- https://www.marktechpost.com/2026/08/20/liquid-ai-releases-lfm2-5-dspark-draft-models-that-deliver-up-to-3-18x-faster-decoding/
- https://x.com/Marktechpost/status/2090516870185255346
- https://www.marktechpost.com/2026/01/06/liquid-ai-releases-lfm2-5-a-compact-ai-model-family-for-real-on-device-agents/
- https://www.marktechpost.com/2026/08/06/liquid-ai-lfm2-5-2-6b-on-device-agentic-model/
- https://www.deeplearning.ai/the-batch/deepseeks-dspark-gains-velocity
