OFICIAL Google Cloud Blog

AI21 achieves an 83% reduction in time-to-start for AI workloads with AI Hypercomputer

What happened
Based on Google Cloud Blog · Oct 02, 2026

AI21 Labs reduced AI workload wait times by 83% and eliminated manual scheduling by adopting Google Cloud’s AI Hypercomputer with Kueue and GKE.

AI21 achieves an 83% reduction in time-to-start for AI workloads with AI Hypercomputer
Google Cloud Blog — Google
Key points
·
AI21 reduced high-priority job wait times from 72 hours to 12 after adopting Google Cloud AI Hypercomputer with Kueue.
·
Manual scheduling interventions dropped from 20 per week to zero by replacing Slack-based coordination with automated Kueue scheduling.
·
Fragmentation across AI21’s GPU cluster fell from 15% to 8% after implementing Topology Aware Scheduling and Admission Fair Sharing.
Key numbers
·
Before migrating to AI Hypercomputer, AI21 relied on manual coordination via Slack to secure compute capacity, which became untenable as cluster utilization approached 100%.
·
High-priority training runs requiring large contiguous GPU allocations often waited up to 72 hours, while fragmentation left free GPUs unusable due to scattered allocation across nodes.
·
The transition to Kueue and AI Hypercomputer reduced manual scheduling interventions from 20 per week to zero and cut high-priority job wait times from 72 hours to 12.

AI21 Labs, a developer of foundation models including the Jamba family, adopted Google Cloud’s AI Hypercomputer to address scheduling bottlenecks in its AI workloads. The company operates a shared Google Kubernetes Engine cluster pooling thousands of A3 and A3 Ultra instances with NVIDIA H100 and H200 GPUs, supporting model training and agent optimization across multiple teams.

Before migrating to AI Hypercomputer, AI21 relied on manual coordination via Slack to secure compute capacity, which became untenable as cluster utilization approached 100%. High-priority training runs requiring large contiguous GPU allocations often waited up to 72 hours, while fragmentation left free GPUs unusable due to scattered allocation across nodes.

AI21 evaluated open-source schedulers and selected Kueue for its compatibility with standard Kubernetes and minimal integration overhead. Paired with AI Hypercomputer’s flexible operations on GKE, Kueue enabled automatic scheduling across reserved and elastic capacity, including Spot VMs and Dynamic Workload Scheduler, eliminating manual negotiations.

The transition to Kueue and AI Hypercomputer reduced manual scheduling interventions from 20 per week to zero and cut high-priority job wait times from 72 hours to 12. Fragmentation decreased from 15% to 8%, and the cluster’s “zombie job” problem was resolved without increasing total cost.

Original source → Deals on Clipraptor.com →