Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment
NVIDIA outlines how AI factories maximize ROI through throughput, durability, and workload flexibility, citing Vera Rubin systems and CUDA-X libraries.
AI factories require multi-million-dollar investments, with megawatt-scale facilities costing roughly $60 million each. Operators prioritize return on investment, which depends on sustained high throughput, durable performance, and broad workload compatibility. Engineering codesign across the full stack—from models to hardware—optimizes efficiency, while standardized architectures ensure accessibility for all operators.
Power efficiency is critical, as tokens per second per megawatt determine earning capacity. NVIDIA Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt than GB300 NVL72 systems, with up to 45x lower cost per million tokens on the DeepSeek V4 Pro model. These gains stem from extreme codesign across the full stack, including compute, networking, and memory.
Demand for compute does not shrink with cheaper tokens, as new use cases emerge that consume more tokens than efficiency gains save. Older generations retain economic value, with systems like the A100 GPU still in commercial service six years after launch. Depreciation schedules have extended, reflecting hardware longevity and resale value.
NVIDIA AI factories support every type of AI model and workload, from data processing to inference, across diverse environments. CUDA-X libraries enable broad compatibility, with over 1,000 ready-made libraries covering applications from deep neural networks to climate modeling. General-purpose accelerated computing keeps utilization high, ensuring sustained returns.