OFICIAL Hugging Face Blog

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

What happened
Based on Hugging Face Blog · Oct 01, 2026

Hugging Face released Olmo-core 3, an open framework designed to scale mixture-of-experts (MoE) training to trillion-parameter models while maintaining computational efficiency and reducing costs.

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Hugging Face Blog — Hugging Face
Key points
·
Olmo-core 3 scales MoE training to trillion-parameter models while maintaining computational efficiency
·
A 47-billion-parameter MoE achieved 2.7x higher throughput with Olmo-core 3 on NVIDIA B300 GPUs
·
The framework supports MXFP8, reducing memory usage and improving training throughput in benchmarks
Key numbers
·
A 47-billion-parameter MoE processed 52,000 tokens per second per GPU using the new stack, compared to 19,400 with the previous implementation—a 2.
·
7x throughput increase.
·
2 trillion total parameters, achieving 858 TFLOP/s/GPU in system performance tests.

Hugging Face unveiled Olmo-core 3, an upgraded open framework for developing large language models with a redesigned mixture-of-experts (MoE) training system. The update is built to scale MoE training into the trillion-parameter range while preserving computational efficiency, serving as a core system for the next generation of Olmo. The framework aims to lower barriers for academic researchers and smaller labs by improving training infrastructure and reducing energy and compute costs associated with large AI model development.

Olmo-core 3 introduces techniques to distribute large MoEs across GPU clusters without requiring every GPU to store the entire model in memory. Key optimizations include rowwise expert parallelism, GPU-resident routing, and grouped GEMM, which collectively reduce communication and coordination overhead. The framework also supports MXFP8, a lower-precision number format that can improve throughput and reduce memory usage when applied selectively.

Benchmark tests on NVIDIA B300 GPUs showed significant performance gains with Olmo-core 3. A 47-billion-parameter MoE processed 52,000 tokens per second per GPU using the new stack, compared to 19,400 with the previous implementation—a 2.7x throughput increase. The framework has been tested at scales up to 1.2 trillion total parameters, achieving 858 TFLOP/s/GPU in system performance tests.

Olmo-core 3 is fully open-source, allowing researchers and developers to train their own MoEs, adapt the system to different hardware, and experiment with routing and parallelism. The framework underpins Hugging Face’s next-generation Olmo model, which will use an MoE architecture trained on a larger dataset with an extended context window. The release reflects Hugging Face’s commitment to open model development by making both model weights and training infrastructure publicly accessible.

Original source → Deals on Clipraptor.com →