NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
NVIDIA’s Vera Rubin NVL72 system achieved up to 3.7x higher throughput than GB300 NVL72 in MLPerf Inference v6.1 benchmarks, demonstrating improved AI inference economics through hardware-software codesign.
NVIDIA’s Vera Rubin NVL72 system delivered leading performance in MLPerf Inference v6.1 benchmarks, focusing on DeepSeek-R1 and Qwen3-VL workloads. The results highlight Vera Rubin’s ability to generate more tokens per rack, improving revenue potential while reducing cost per token. Performance gains were achieved using NVIDIA’s Dynamo open source inference framework and TensorRT-LLM library, emphasizing software optimization as a key driver of efficiency.
The Vera Rubin NVL72 leverages enhanced Tensor Cores and Transformer Engine to accelerate both prefill and decode stages of inference. Disaggregated serving and large-scale expert parallelism further boost throughput, particularly for mixture-of-experts models like DeepSeek-R1 and Qwen3-VL. The system’s sixth-generation NVLink and NVLink Switch provide 10x higher packet rates and 3x lower latency than Ethernet, enabling effective rack-scale scaling.
NVIDIA’s partner Nebius also submitted Vera Rubin NVL72 preview results, demonstrating strong performance in agentic inference workloads. In SemiAnalysis AgentX benchmarks, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72. The upcoming MLPerf Endpoints benchmark will standardize measurements for agentic inference, reflecting evolving workload demands beyond traditional throughput metrics.
Scaling efficiency was a critical focus, with Vera Rubin NVL72 achieving 99% scaling efficiency on DeepSeek-R1 in offline scenarios. The system scaled from 72 to 288 GPUs with throughput growing nearly in proportion to hardware additions. Continuous software improvements, including lower KV cache precision and kernel optimizations, contributed to performance gains in MLPerf v6.1.