OFICIAL Hugging Face Blog

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

What happened
Based on Hugging Face Blog · Aug 12, 2026

Hugging Face released LFM2.5-VL-3B, a compact vision-language model designed for on-device and edge applications, supporting real-time document, screen, and object understanding with tool-calling capabilities.

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
Hugging Face Blog — Hugging Face
Key points
·
How we trained its most capable vision-language model Benchmark results Inference speed on CPU and GPU How to use LFM2.5-VL-3B LFM2.5-VL-3B demo Get Started Citation LFM2.5-VL-3B is its most capable vision-language model you can run on your own hardware.
·
It understands documents and screens alike, grounds objects, and can call tools.
·
It answers directly instead of reasoning, so responses stay fast in real-time and on-device apps.
·
LFM2.5-VL-3B pairs a SigLIP2 400M NaFlex vision encoder with the same pre-trained backbone as its LFM2.5-2.6B text model.
Key numbers
·
The model combines a SigLIP2 400M NaFlex vision encoder with a pre-trained text backbone, trained on approximately 34 trillion tokens including four times more vision data than prior iterations.
·
On text-only benchmarks, it demonstrates improved instruction following and tool-use capabilities, matching performance with models like Gemma-4-E2B and Qwen3.
·
5-2B in tool-calling scenarios.

Hugging Face introduced LFM2.5-VL-3B, a 3-billion-parameter vision-language model optimized for local hardware deployment. The model combines a SigLIP2 400M NaFlex vision encoder with a pre-trained text backbone, trained on approximately 34 trillion tokens including four times more vision data than prior iterations. Its expanded 128K-token vocabulary supports non-Latin scripts without retraining the tokenizer from scratch. Post-training involves supervised fine-tuning with knowledge distillation and multi-reward reinforcement learning to enhance performance across diverse tasks.

The model was evaluated on benchmarks covering multilingual visual comprehension, instruction following, document and screen understanding, object detection, and tool use. Results show LFM2.5-VL-3B leads its size class in real-world image tasks while maintaining strong performance in reading digital content such as documents, charts, and UI elements. On text-only benchmarks, it demonstrates improved instruction following and tool-use capabilities, matching performance with models like Gemma-4-E2B and Qwen3.5-2B in tool-calling scenarios.

LFM2.5-VL-3B achieves real-time inference speeds of 228 tokens per second on an Apple M5 Max and 116 tokens per second on an AMD Ryzen AI Max+ 395, with a memory footprint of about 3 GB. On mobile devices, it processes 20 tokens per second on a Galaxy S26 Ultra, enabling fully on-device operation. For GPU deployment, the model maintains low latency and delivers approximately 11,000 tokens per second at high concurrency, outperforming larger 4B-class models and smaller 2B-class models in output throughput.

The model is available immediately on Hugging Face with day-one support across inference frameworks including llama.cpp, MLX, vLLM, SGLang, and ONNX. Documentation provides examples for multi-image inputs, grounding, OCR, and tool calling, while a browser demo demonstrates its capabilities in a vision-capable chat interface. Hugging Face positions LFM2.5-VL-3B as a strong general-purpose solution for high-volume on-device workloads requiring real-time visual and textual processing.

Original source → Deals on Clipraptor.com →