From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production
NVIDIA’s DSX platform enables AI factories to dynamically adjust power use, boosting efficiency without new infrastructure, as demonstrated in a Silicon Valley Power pilot.
On a summer evening in Silicon Valley, an AI factory automatically reduced its power demand from four megawatts to three in response to a utility signal, without interrupting critical workloads. The adjustment was orchestrated by Emerald AI’s Conductor platform, which prioritized lower-value tasks while maintaining high-priority inference services. The demonstration marked the first deployment across thousands of NVIDIA GPUs, with Silicon Valley Power later sending over 200 demand signals, all executed successfully. The pilot highlighted how AI factories can act as flexible grid resources, avoiding the need for new transmission lines.
At the AI Infra Summit, NVIDIA’s Ian Buck emphasized AI factory efficiency, while Lambda released validation results showing a 24% increase in token throughput within a fixed power budget. Lambda’s cloud services president, Dave Ward, stated that NVIDIA DSX MaxLPS converts stranded capacity into productive compute, improving performance per watt by 23%. The findings were based on a five-rack, 19-node cluster running NVIDIA HGX B200 GPU servers, demonstrating practical gains in real-world deployments.
NVIDIA DSX MaxLPS dynamically reallocates power across nodes based on workload type, optimizing allocation for both training and inference. The software monitors GPU and rack-level consumption, recovering capacity that static provisioning would leave unused. Lambda’s deployment achieved 24% more cluster-wide token throughput—from roughly 4 million to 5 million tokens per second—while operating within the same power budget as 16 nodes at full power.
NVIDIA founder and CEO Jensen Huang has noted that expanding AI factory capacity through physical power increases is unsustainable, advocating instead for system-level optimization. Introduced at GTC Taipei in May, NVIDIA DSX addresses this by integrating power management with networking, cooling, and facility design. Early deployments, including the Santa Clara pilot, prove the platform’s ability to unlock stranded capacity, with projections suggesting up to 40% more GPU capacity for future AI factories like Vera Rubin NVL72.