Home/Events/Up to 3.2x Faster Inference with LFM2.5-DSpark

Up to 3.2x Faster Inference with LFM2.5-DSpark

Confirmed
Confidence
90%
Impact: 80%
Updated 5d ago

Consensus Brief

Hugging Face has released draft model checkpoints for three models in the LFM2.5 family, which include LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. These models utilize a new speculative decoding method called DSpark, resulting in significant improvements in inference speed without compromising output quality.

Sourced from
Primary: Hugging Face

What Changed Since Last Update

5d ago

New official source added: Hugging Face published an update on Thu, 24 Se ("Accelerating vision-language models with LFM2.5-VL-DSpark").

Claim Ledger

4 claims tracked across sources

Confirmed Fact

DSpark achieves up to 3.2x throughput improvement on a GPU and up to 2.87x on-device.

Confirmed Fact

Cuts function-calling latency by 57% on average for LFM2.5-2.6B.

Confirmed Fact

The draft models are relatively small, with each around ~300M parameters.

Confirmed Fact

Day-one support for llama.cpp and SGLang is provided.

Role-Based Impact Analysis