Home/Events/LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

Confirmed
Confidence
90%
Impact: 80%
Updated 1h ago

Consensus Brief

Hugging Face has released LFM2.5-VL-3B, a vision-language model designed for on-device applications, which significantly improves understanding of screens, grounding objects, and function calling. The model is pre-trained on 34 trillion tokens and features enhanced capabilities over its predecessors.

Sourced from
Primary: Hugging Face

What Changed Since Last Update

1h ago

LFM2.5-VL-3B introduces four major improvements: enhanced screen understanding, improved grounding and object detection, multi-image input reasoning, and stronger function calling capabilities.

Claim Ledger

5 claims tracked across sources

Official Claim

LFM2.5-VL-3B is the most capable vision-language model that can run on personal hardware.

Confirmed Fact

The model is pre-trained on about 34 trillion tokens with four times more vision data than previous models.

Official Claim

LFM2.5-VL-3B leads its size class on real-world image tasks and excels in reading digital content.

Confirmed Fact

Inference speed on CPU is 228 tokens/s on an M5 Max and 116 tokens/s on a Ryzen AI Max+ 395.

Confirmed Fact

LFM2.5-VL-3B is the fastest model tested, reaching about 11K tokens per second at high concurrency.

Role-Based Impact Analysis

Source Timeline

1 source corroborating