Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
NVIDIA’s Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than GB300 NVL72 on agentic AI workloads, according to measured performance data using the SemiAnalysis AgentX workload.
Agentic AI workloads, such as those performing investment research or software development, require significantly more tokens than simple chat requests due to iterative reasoning and tool calls. Each step in an agentic workflow accumulates context, often reaching hundreds of thousands of tokens, which demands efficient long-context handling. NVIDIA reports that Vera Rubin NVL72 systems achieve up to 30x higher throughput per megawatt than GB300 NVL72 on these workloads, as measured using the SemiAnalysis AgentX workload with real-world agentic coding sessions preserved.
The performance advantage translates to 30x more agentic work for the same energy footprint in power-constrained AI factories. Vera Rubin NVL72’s efficiency is attributed to extreme codesign across hardware and software layers, including fifth-generation Tensor Cores, third-generation Transformer Engine, and NVFP4 quantization. The platform’s NVL72 scale-up domain enables high-bandwidth inter-GPU communication via sixth-generation NVLink technology, supporting techniques like distributed KV-caching and expert parallelism.
NVIDIA’s software stack, including TensorRT LLM and Dynamo, is codesigned with Vera Rubin’s hardware to optimize inference performance. The platform also incorporates purpose-built components such as the Vera CPU, Groq 3 LPU, and BlueField-4 DPU, all engineered for AI factories deploying agentic workloads at scale. Early results indicate Vera Rubin NVL72 delivers up to 35x lower cost per million tokens than GB300 NVL72, improving profit margins for AI factories.
The Vera Rubin platform is in full production and scaling across the ecosystem. NVIDIA DSX MaxLPS technologies manage power across GPU, rack, and workload levels, enabling up to 40% more GPUs within the same megawatt budget. Performance improvements are expected to continue with ongoing software optimizations for both Vera Rubin NVL72 and GB300 NVL72.