OFICIAL Hugging Face Blog

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

What happened
Based on Hugging Face Blog · Oct 07, 2026

Hugging Face reports Nemotron models achieved gold-level results in IOI 2026 and IMO 2026 through fine-tuning and specialized inference pipelines, releasing open resources for reproducibility.

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
Hugging Face Blog — Hugging Face
Key points
·
Nemotron-3-Nano-CC reached 468 points on IOI 2025 after SFT, RL, and GenCorrect, surpassing the gold threshold of 438.3.
·
Nemotron-3-Ultra-CC scored 535.4 out of 600 on IOI 2026 using the same test-time strategy as the smaller model.
·
IMO 2026 system scored 30 out of 42 points, exceeding the gold threshold, using complementary SFT and RL checkpoints without formal provers.
Key numbers
·
The IOI system leveraged Nemotron-3-Nano-CC and Nemotron-3-Ultra-CC, with the smaller model improving from 130 to 468 points after supervised fine-tuning, reinforcement learning, and the GenCorrect generate-evaluate-refine strategy.
·
4 out of 600 using the same test-time approach.
·
For IMO, teams trained specialists on 414,890 proof problems and 9,597 frontier problems, combining supervised fine-tuning and reinforcement learning checkpoints to score 30 out of 42 points, exceeding the gold threshold.

Hugging Face announced that its Nemotron models reached gold-medal performance in both the International Olympiad in Informatics (IOI) 2026 and the International Mathematical Olympiad (IMO) 2026. The IOI result was achieved under live, prospective constraints mirroring human contestant conditions, while the IMO proofs were evaluated by official graders. Neither result was included in official rankings, as the systems participated unofficially.

The IOI system leveraged Nemotron-3-Nano-CC and Nemotron-3-Ultra-CC, with the smaller model improving from 130 to 468 points after supervised fine-tuning, reinforcement learning, and the GenCorrect generate-evaluate-refine strategy. The Ultra model scored 535.4 out of 600 using the same test-time approach. For IMO, teams trained specialists on 414,890 proof problems and 9,597 frontier problems, combining supervised fine-tuning and reinforcement learning checkpoints to score 30 out of 42 points, exceeding the gold threshold.

The IMO pipeline used complementary strengths from SFT and RL checkpoints, generating candidate proofs, scoring them, producing critiques, and refining attempts without formal provers or external tools. A high-compute selection stage chose the final submission, demonstrating how co-designing model specialization and inference loops drives performance. The approach contrasts with relying solely on fine-tuning or brute-force sampling.

Hugging Face released open resources including Nemotron Labs IMO 2026 checkpoints, datasets, and Nemotron-IMO-Bench, alongside IOI training recipes and inference pipelines in NeMo-Skills. The releases aim to enable reproducible research beyond competitive settings, emphasizing reusable specialization recipes for demanding domains.

Original source → Deals on Clipraptor.com →