OFICIAL Hugging Face Blog

Impactful scheduling for GPU clusters

What happened
Based on Hugging Face Blog · Oct 09, 2026

Ai2 replaced a priority-based GPU scheduler with a budget-driven system to prioritize impactful research while maintaining full cluster occupancy across thousands of NVIDIA GPUs.

Impactful scheduling for GPU clusters
Hugging Face Blog — Hugging Face
Key points
·
Ai2 replaced a priority-based GPU scheduler with a hierarchical fair-share system using GPU time budgets to prioritize impactful research.
·
The new system ensures every GPU time request must be funded by a budget or risks preemption, reducing squatting and priority inflation.
·
Allocation decisions are made by research leads, principal investigators, and program managers, with guaranteed shares of cluster capacity.
Key numbers
·
Demand consistently exceeds supply, with outstanding requests for 2-3 times the available GPU capacity at any given moment.
·
For example, Project A1 receives a 35% claim on total cluster capacity, with allocation decisions made by leads, principal investigators, and program managers.

Ai2’s AI Infrastructure team manages thousands of NVIDIA H100, B200, and B300 GPUs across clusters of 88 to 1024 GPUs, serving about 150 internal researchers in domains like LLM training and robotics reinforcement learning. Demand consistently exceeds supply, with outstanding requests for 2-3 times the available GPU capacity at any given moment. Historically, a priority-based scheduler led to inefficiencies such as GPU squatting, where users reserved idle GPUs with no-op workloads to avoid latency in debugging tasks.

The team identified priority inflation as a core issue, where all workloads eventually used the highest priority setting, starving lower priority levels entirely. Engineers spent excessive time negotiating shutdowns of non-preemptable workloads on hosts with maintenance issues, delaying root cause identification. These challenges highlighted a 'tragedy of the commons' scenario, where individual optimization led to suboptimal global outcomes and resource abuse.

To address this, Ai2 transitioned from assigning GPU monopolies to allocating GPU time budgets, enabling leadership to fund research efforts based on strategic impact assessments. The new system uses a hierarchical fair-share scheduler paired with GPU time budgets, where every request must be funded by a budget or risks preemption. This shifts the debate from operational negotiations to transparent administrative budgeting processes.

The hierarchical system allows managers to proportionally allocate GPU time to projects and researchers, ensuring guaranteed shares of capacity regardless of queue length. For example, Project A1 receives a 35% claim on total cluster capacity, with allocation decisions made by leads, principal investigators, and program managers. The scheduler enforces these allocations, making it costlier to game the system than to engage in budget debates.

Original source → Deals on Clipraptor.com →