OFICIAL NVIDIA Newsroom

How XPUs Meet a World-Class AI Factory

What happened
Based on NVIDIA Newsroom · Aug 24, 2026

NVIDIA introduced NVLink Fusion to integrate custom XPUs with its AI infrastructure, reducing deployment complexity and accelerating time to market for hyperscalers and AI-native companies.

How XPUs Meet a World-Class AI Factory
NVIDIA Newsroom — NVIDIA
Key points
·
To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime.
·
That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators.
·
At AI factory scale, this path is complex and costly, and represents a fundamental obstacle to getting XPUs to market quickly.
·
Breaking the constraint means combining custom XPUs with proven, mature infrastructure — allowing builders to focus innovation where it matters most while harnessing established technology for the rest.
Key numbers
·
NVLink Fusion integrates XPUs into NVIDIA’s NVLink scale-up domain, offering high-bandwidth, low-latency networking across up to 72 XPUs.
·
Sixth-generation NVLink reduces XPU-to-XPU latency by 3x compared to Ethernet alternatives and increases packet rates by 10x.
·
The technology also includes NVIDIA NVLink-C2C for connecting XPUs to Vera CPUs, improving energy efficiency by up to 6x over PCIe interfaces.

AI factories require continuous operation and depend on metrics like tokens per second and cost per token, making infrastructure design critical. Hyperscalers building custom XPUs face challenges in developing full AI platforms, including networking, rack architecture, and software. NVLink Fusion addresses this by combining custom XPUs with NVIDIA’s mature infrastructure, enabling faster deployment and reduced risk for semi-custom AI factories. The solution allows builders to focus innovation on XPU design while leveraging established technology for other components.

NVLink Fusion integrates XPUs into NVIDIA’s NVLink scale-up domain, offering high-bandwidth, low-latency networking across up to 72 XPUs. Sixth-generation NVLink reduces XPU-to-XPU latency by 3x compared to Ethernet alternatives and increases packet rates by 10x. Systems like the NVIDIA GB300 NVL72 deliver higher throughput and interactivity, with future configurations supporting up to 1,152 accelerators. The technology also includes NVIDIA NVLink-C2C for connecting XPUs to Vera CPUs, improving energy efficiency by up to 6x over PCIe interfaces.

Custom XPU developers often underestimate the effort required to deploy AI infrastructure, including platform design, integration, and supply chain management. NVLink Fusion provides a unified architecture that supports rapid development, integration, and deployment, with ecosystem partners spanning ASIC design, CPU, and optical interconnects. Customers can select CPU architectures, performance levels, and software capabilities tailored to their workloads, as noted by Intel’s Tim Wilson.

NVLink Fusion aligns with NVIDIA’s DSX reference architecture for AI factories, enabling pre-construction validation of infrastructure through digital twin modeling. The solution supports shared rack footprints, networking, cooling, and power delivery for XPU and GPU systems, allowing operators to defer precise silicon choices. Serviceability features, such as liquid-cooled trays, ensure continuous operation during maintenance, reducing downtime and costs.

Original source → Deals on Clipraptor.com →