OFICIAL Hugging Face Blog

LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation

What happened
Based on Hugging Face Blog · Aug 19, 2026

Hugging Face released QAD Q4_0 GGUF checkpoints for four LFM2.5 models, enabling 4-bit quantization with minimal performance loss while improving decode speed on edge hardware.

LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
Hugging Face Blog — Hugging Face
Key points
·
Benchmark results Speed and size on real edge hardware How to use QAD GGUFs Get Started with QAD GGUFs Citation Today, we release QAD Q4_0 GGUFs.
·
These are updated 4-bit checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B.
·
We also add one scale-appropriate math evaluation: GSM8K for LFM2.5-230M and LFM2.5-350M, and AIME25 for LFM2.5-1.2B-Instruct and LFM2.5-2.6B.
·
Across all four models, QAD substantially improves the Q4_0 checkpoint.
Key numbers
·
5 models, including versions 230M, 350M, 1.
·
These checkpoints maintain 96.
·
5% to 97.

Hugging Face has introduced QAD Q4_0 GGUF checkpoints for LFM2.5 models, including versions 230M, 350M, 1.2B-Instruct, and 2.6B. These checkpoints maintain 96.5% to 97.4% of the BF16 baseline performance while operating in 4-bit quantization, addressing the typical quality trade-off associated with post-training quantization methods.

Performance benchmarks across reasoning, instruction-following, tool use, and agentic capabilities—including GPQA Diamond, MMLU-Pro, and IFBench—showed consistent improvements with QAD Q4_0. The checkpoints were evaluated against BF16 and F16 baselines, with additional math-specific tests like GSM8K and AIME25 for smaller and larger models, respectively.

Decode throughput tests on edge devices such as MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, and Raspberry Pi 5 revealed that QAD Q4_0 checkpoints achieved 3% to 33% higher throughput compared to standard Q4_0 or Q5_K_M configurations. The 230M and 350M models matched Q5_K_M quality, while the 1.2B and 2.6B models matched Q4_K_M quality.

The QAD Q4_0 GGUF files are now available on Hugging Face for use with llama.cpp or any GGUF-compatible runtime. These checkpoints are positioned as competitive alternatives to Unsloth's UD-Q4_K_XL for applicable models, offering developers optimized performance for edge deployment.

Original source → Deals on Clipraptor.com →