AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories
NVIDIA unveiled Vera Rubin and DSX platform updates at the AI Infra Summit, highlighting energy efficiencies in AI factories by optimizing tokens per watt, with new benchmarks and grid-responsive power management.
At the AI Infra Summit in Santa Clara, NVIDIA’s Ian Buck emphasized the shift in AI infrastructure metrics from raw performance to validated agentic tokens per megawatt, addressing demands from agentic AI workloads. The company introduced DSX MaxLPS, which delivers up to 1.4x more tokens per megawatt through factory-wide power optimization, while NVLink unites large-scale accelerated computing into a single high-performance system. The platform is designed to optimize system performance, infrastructure scaling, and software efficiency, enabling enterprises to run diverse AI workloads on unified infrastructure.
New MLPerf Inference v6.1 results demonstrated NVIDIA’s advantages, with the Vera Rubin NVL72 system delivering up to 3.7x higher throughput than the GB300 NVL72. A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, showcasing near-linear throughput growth. Software optimizations alone improved performance by up to 1.6x from v6.0 to v6.1, with additional gains post-submission, underscoring the impact of continuous innovation.
Silicon Valley Power and Emerald AI collaborated to demonstrate automated load reduction, using NVIDIA DSX Flex to dynamically adjust AI factory energy consumption based on real-time grid signals. DSX Flex prioritizes critical workloads while temporarily pausing lower-priority jobs, enabling AI factories to act as flexible grid resources. This approach helps reduce grid demand during peak periods while maintaining performance for high-priority AI tasks.
Lambda validated NVIDIA DSX MaxLPS on Blackwell servers, showing a 24% increase in cluster-wide token throughput—from 4 million to 5 million tokens per second—while improving performance per watt by 23%. For Vera Rubin NVL72 systems, MaxLPS can enable up to 40% more GPU capacity within the same megawatt budget, maximizing AI factory productivity and economic value.