Next upAI x Bio Pitch Contest
News

Liquid AI brings DSpark speculative decoding to LFM2.5-VL-3B

Liquid AI released an experimental DSpark drafter for LFM2.5-VL-3B, bringing speculative decoding to vision-language inference in SGLang, MLX-VLM and llama.cpp.

D
Sep 27, 2026 · 2 min read

Liquid AI released an experimental DSpark draft model for LFM2.5-VL-3B on Sept. 24, extending its speculative-decoding approach to vision-language inference. The drafter is available for SGLang, MLX-VLM and llama.cpp, offering a deployment option intended to accelerate token generation without changing the target model.

In company-run tests detailed in the model card, Liquid AI reported decode speedups of 2.04x to 2.66x on one 80GB NVIDIA H100, 2.30x to 3.13x on an Apple M5 Max, and 1.57x to 2.14x on an M3 Ultra. End-to-end gains were lower in the same company tests: 1.64x to 2.27x, 1.56x to 2.62x and 1.30x to 1.77x, respectively. Those figures are Liquid AI’s measurements and have not been independently reproduced in the sources reviewed for this story.

Speculative decoding uses a smaller draft model to propose several tokens, then has the full target model verify them. Liquid AI says the vision-language system can reuse this method because image patches and text tokens are projected into a shared representation before the target layers that supply hidden states to the drafter. The released drafter is a four-layer, attention-only model with a Markov head and a confidence head, according to the company.

Liquid AI says it tested six tasks modeled on the MMSpec benchmark, using batch size 1 and temperature 0. The H100 tests used SGLang with BF16 precision, while the Apple tests used FP16 weights with MLX-VLM on the M5 Max and llama.cpp on the M3 Ultra. The underlying MMSpec preprint contains 600 samples across six multimodal task categories and cautions that throughput speedup alone does not reliably describe latency.

The gain applies to decoding rather than image encoding or prompt prefill. Liquid AI says those earlier stages therefore limit end-to-end improvements, particularly on edge hardware. The company’s 279.5 million-parameter figure for the drafter represents an 8.9% increase in deployed parameter count; it is not a measured RAM or VRAM overhead figure, and the company did not publish added memory use in bytes for the tested systems.

Liquid AI also claims the method is exact under greedy decoding because the target model verifies every proposed token, producing the same accepted generation as the target alone. It further says matched sampling settings at nonzero temperatures preserve the target model’s output distribution. Those output-equivalence statements were not independently tested for this release.

Upstream support was merged before the announcement: SGLang on Sept. 22, llama.cpp on Sept. 23 and MLX-VLM on Sept. 18. Liquid AI provides the drafter in Safetensors and GGUF formats. The llama.cpp option also provides a point of comparison with other recent local AI model work.

The first-party materials contain two unresolved discrepancies. The blog prints a 20.4x lower bound for H100 decode speed, while the detailed model-card table lists 2.04x; the latter is used here. The blog also says both measured configurations used block size 8, whereas the model card says the H100 used block size 9 and the Apple systems used block size 8.

More news