OFICIAL Google Cloud Blog

Accelerating agentic RL and evaluation research velocity with 45x faster GKE Agent Sandbox

What happened
Based on Google Cloud Blog · Sep 29, 2026

Google Cloud introduces GKE Agent Sandbox and an RL orchestration SDK to reduce idle GPU time in agentic reinforcement learning by eliminating CPU sandbox cold starts and optimizing image hydration.

Accelerating agentic RL and evaluation research velocity with 45x faster GKE Agent Sandbox
Google Cloud Blog — Google
Key points
·
GKE Agent Sandbox eliminates CPU sandbox cold starts using SandboxWarmPool for pre-initialized environments in agentic RL workloads.
·
Agent Sandbox RL orchestration SDK integrates with Gymnasium, NVIDIA NeMo Gym, and OpenHands for simplified RL training setup.
·
Mistral AI uses the system to orchestrate over 30,000 sandboxes on a single cluster, reducing idle GPU time in RL training.
Key numbers
·
1x during high-concurrency bursts.

Google Cloud addresses infrastructure bottlenecks in agentic reinforcement learning (RL) where GPU clusters remain idle during CPU sandbox cold starts and multi-gigabyte image pulls. The new GKE Agent Sandbox, optimized for RL workloads, and the Agent Sandbox RL orchestration SDK are now generally available to streamline large-scale agentic RL training and evaluations. The solution integrates with popular RL tools and leverages Kubernetes primitives to support rapid, parallel agent execution without manual configuration overhead.

Agentic RL workloads require isolated CPU sandboxes for action execution, but scaling to tens of thousands of parallel rollouts exposes critical bottlenecks: cold-start delays, image hydration overhead, and control-plane instability. GKE Agent Sandbox mitigates these issues with SandboxWarmPool, which pre-initializes environments to eliminate cold starts, and GKE Image Streaming to handle high-cardinality image sets efficiently. The system also supports checkpointing and sandbox forking for error recovery and parallel exploration.

The Agent Sandbox RL orchestration SDK provides an async Python API with pluggable warm-pooling strategies, reducing the need for custom Kubernetes configurations. Native integrations with RL tools like Gymnasium, NVIDIA NeMo Gym, and OpenHands simplify deployment. Testing on a 10-node gVisor sandbox pool demonstrated performance gains across workloads ranging from 500 to 18,312 concurrent sandboxes, with image hydration and control-plane pacing identified as key contributors to efficiency improvements.

Google Cloud’s solution prioritizes accelerator utilization by moving image hydration off the critical path and implementing rate controls in the Agent Sandbox Controller to stabilize the control plane. The SDK’s in-place pod recycling strategy further reduces scheduling overhead, cutting pod creations by 3.1x during high-concurrency bursts. Mistral AI has adopted the technology to orchestrate over 30,000 sandboxes on a single cluster, enabling faster model training and iteration cycles.

Original source → Deals on Clipraptor.com →