LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
Hugging Face released QAD Q4_0 GGUF checkpoints for four LFM2.5 models, enabling 4-bit quantization with minimal performance loss while improving decode speed on edge hardware.
Hugging Face has introduced QAD Q4_0 GGUF checkpoints for LFM2.5 models, including versions 230M, 350M, 1.2B-Instruct, and 2.6B. These checkpoints maintain 96.5% to 97.4% of the BF16 baseline performance while operating in 4-bit quantization, addressing the typical quality trade-off associated with post-training quantization methods.
Performance benchmarks across reasoning, instruction-following, tool use, and agentic capabilities—including GPQA Diamond, MMLU-Pro, and IFBench—showed consistent improvements with QAD Q4_0. The checkpoints were evaluated against BF16 and F16 baselines, with additional math-specific tests like GSM8K and AIME25 for smaller and larger models, respectively.
Decode throughput tests on edge devices such as MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, and Raspberry Pi 5 revealed that QAD Q4_0 checkpoints achieved 3% to 33% higher throughput compared to standard Q4_0 or Q5_K_M configurations. The 230M and 350M models matched Q5_K_M quality, while the 1.2B and 2.6B models matched Q4_K_M quality.
The QAD Q4_0 GGUF files are now available on Hugging Face for use with llama.cpp or any GGUF-compatible runtime. These checkpoints are positioned as competitive alternatives to Unsloth's UD-Q4_K_XL for applicable models, offering developers optimized performance for edge deployment.