OFICIAL Google Cloud Blog

Global AI routing with <1% overhead on multi-cluster GKE Inference Gateway

What happened
Based on Google Cloud Blog · Sep 21, 2026

Google Cloud introduced a multi-cluster GKE Inference Gateway that pools globally distributed AI accelerators into a single virtual fleet, achieving near-linear throughput gains with less than 1% routing overhead and a 99.9% success rate under heavy load.

Global AI routing with <1% overhead on multi-cluster GKE Inference Gateway
Google Cloud Blog — Google
Key points
·
Google Cloud’s multi-cluster GKE Inference Gateway pools 17,000 nodes across three regions into a single virtual fleet for AI inference.
·
The Gateway routes traffic based on real-time KV-cache utilization, achieving a 99.9% success rate under heavy concurrency with less than 1% overhead.
·
Scaling from one to three clusters delivered near-linear throughput gains while maintaining 99.5% of local cluster throughput.
Key numbers
·
9% success rate even under heavy multi-client concurrency, proving its reliability for production-scale AI workloads.
·
Routing overhead remained below 1%, with the Gateway delivering 99.
·
5% of the throughput of a direct, local cluster call.

Google Cloud’s new multi-cluster GKE Inference Gateway addresses the challenge of fragmented AI infrastructure by treating geographically dispersed accelerator clusters as a unified resource. The system uses real-time telemetry, including KV-cache utilization metrics from inference engines, to dynamically route requests based on actual memory and compute pressure rather than traditional round-robin methods. This approach ensures high utilization of expensive hardware, reducing idle capacity and improving cost efficiency across regions.

In a benchmark involving 17,000 compute nodes across three GKE clusters in the US and Europe, the Gateway demonstrated its scalability by delivering a near-linear throughput increase when expanding from one to three clusters. The deployment, serving a leading Mixture of Experts model via SGLang, maintained a 99.9% success rate even under heavy multi-client concurrency, proving its reliability for production-scale AI workloads.

The Gateway’s architecture integrates with existing distributed inference engines, such as those using tensor parallelism, by routing traffic to leader pods while respecting regional service topologies. This design ensures that cross-region load balancing aligns with underlying multi-node configurations, avoiding inefficiencies that typically arise in distributed AI deployments.

Routing overhead remained below 1%, with the Gateway delivering 99.5% of the throughput of a direct, local cluster call. The system’s ability to detect and respond to memory pressure—such as spilling traffic when KV-cache utilization exceeds 40%—eliminates the need for manual intervention, simplifying operations while maximizing accelerator fleet performance.

Original source → Deals on Clipraptor.com →