AI factories are purpose-built infrastructure platforms designed to produce intelligence at scale. Their economics depend on measurable output, including tokens per second, tokens per watt, cost per token, utilization, and uptime.
Achieving these goals requires an AI infrastructure platform designed and built as a complete factory—not as a collection of individual accelerators.
Hyperscalers and AI-native companies developing custom XPUs must consider more than XPU design. They also need to develop the complete AI platform, including scale-up and scale-out networking, rack-scale architecture, factory software, power and cooling systems, and a resilient supplier ecosystem.
At AI factory scale, building every component from the ground up is complex, expensive, and time-consuming. These challenges can create significant barriers to bringing custom XPUs to market quickly.
The solution is to combine custom XPUs with proven, mature infrastructure. This approach allows developers to focus innovation where it delivers the most value while relying on established technologies for the rest of the platform.
NVLink Fusion addresses this need by connecting custom XPUs to NVIDIA’s AI infrastructure. The platform is designed to improve performance, accelerate time to market, and reduce risk for semi-custom AI factories.
Unlock XPU performance with high-speed scale-up networking
Modern AI workloads—including trillion-parameter models, mixture-of-experts architectures, and agentic AI—depend on fast communication between accelerators. If the scale-up fabric cannot keep pace with the compute platform, utilization declines and cost per token increases.
An effective scale-up networking solution must deliver excellence in three critical areas:
- Performance: End-to-end network performance, in-network computing capabilities, and mature software integration.
- AI factory resilience: High uptime, continuous health monitoring, detailed telemetry, and component-level serviceability during factory operations.
- Platform maturity: A proven technology stack with a track record of large-scale deployments, helping reduce operational risk and improve return on investment.
NVLink Fusion uses the NVIDIA NVLink scale-up fabric. Sixth-generation NVLink provides high-bandwidth, low-latency networking across 72-XPU domains. End-to-end latency for transfers between XPUs is up to three times lower, while packet rates are up to 10 times faster than alternative solutions based on off-the-shelf Ethernet.
In end-to-end performance evaluations, the NVIDIA GB300 NVL72 system delivers higher throughput and improved interactivity compared with configurations without NVL72. Future NVLink roadmap configurations are expected to support up to 1,152 accelerators and co-packaged optics domains.
NVLink Fusion also includes NVIDIA NVLink-C2C for connecting XPUs to the NVIDIA Vera CPU or other ecosystem CPUs. This technology can deliver up to six times the energy efficiency of PCIe interfaces, helping reduce the communication barrier between system control and compute resources.
A proven AI infrastructure stack and ecosystem
Companies developing custom XPUs often underestimate the complexity of converting silicon innovation into a reliable data center deployment. A complete AI platform may require:
- Integration of high-speed CPUs and scale-up interfaces
- Sourcing and validation of scale-up networking solutions
- Compute tray and switch tray design
- Rack architecture design and validation, including power and cooling
- Security and storage integration
- Management of complex manufacturing and supplier ecosystems
The ideal AI infrastructure platform provides these capabilities out of the box. Teams can then focus on targeted XPU innovation while relying on proven solutions for system integration, deployment, and operations.
NVLink Fusion is supported by an ecosystem built for rapid development, integration, and deployment across ASIC design, CPU, intellectual property, and optical interconnect partners.
“NVLink Fusion allows customers to choose the CPU architecture, performance level, and software features that best meet the needs of their workloads,” said Tim Wilson, vice president and general manager of Data Center Silicon Engineering at Intel.
Companies adopting NVLink Fusion can also use the NVIDIA MGX rack-scale architecture and the supply chain supporting MGX-based systems, including the NVIDIA Vera Rubin NVL72. Manufacturing partners can manage system design and integration, while MGX suppliers provide rack infrastructure, cooling, power delivery, and platform building blocks for emerging AI products, including 800 VDC designs.
“With Vera Rubin NVL72, we’re looking at automating nearly 100% of system construction on the manufacturing line,” said Jack Luo, head of products and solutions at QCT and Quanta Computer. “With NVLink Fusion, XPUs can leverage most of these investments.”
Reduce AI factory risk through infrastructure standardization
AI factory planning begins well before the final silicon configuration is available. Power procurement, facility design, cooling infrastructure, rack layouts, and network architecture must often be finalized months or years in advance. Designing a data center around a single chip or accelerator can therefore introduce significant schedule and deployment risk.
Different workloads may require different types of accelerators, including XPUs, GPUs, CPUs, and LPUs. GPU-based systems can operate alongside semi-custom platforms for training, post-training, inference, search, and AI serving.
“The value of the NVLink Fusion program is that enterprises can deploy rack-level solutions using NVIDIA GPUs, then decouple XPU development and move at a different pace,” said Vince Hu, corporate senior vice president and general manager of the Data Center and Computing business group at MediaTek.
NVLink Fusion addresses these challenges with a unified rack-scale architecture. XPU- and GPU-based systems, including the Vera Rubin NVL72, can share rack footprints, networking, cooling, power delivery, and management systems.
This flexibility enables operators to begin building infrastructure while deferring the final silicon mix. Capacity can then be reprovisioned as workload requirements, silicon availability, and business priorities change.
“NVLink Fusion allows hyperscalers and custom ASIC designers to integrate their own custom CPUs and XPUs, bridging NVIDIA technology and third-party processes to create a unified rack-scale architecture,” said Lie-Szu Juang, chairman and chief strategy officer at GUC.
Design, validate, and operate the AI factory as a complete system
Factory assembly is expensive, and design errors can lead to costly rework and deployment delays. AI infrastructure must be modeled and verified before construction begins.
NVLink Fusion works with the NVIDIA DSX AI Factory Reference Architecture to support the co-design of buildings, power systems, cooling, compute, and networking. The NVIDIA Omniverse DSX AI Factory Blueprint provides a digital twin and open reference design for gigawatt-scale AI factories, enabling partners to model facilities and technology together before deployment.
At the rack level, serviceability is an essential part of performance. The reference compute tray uses 100% liquid cooling without fans, cables, or hoses, and can be removed while the rest of the rack remains operational. The NVLink switch tray is also water-cooled to help support continuous operation during maintenance.
“NVLink Fusion allows us to accelerate time to market using the proven NVL72 rack design and access multiple suppliers to get more products into the hands of our customers,” said CC Lee, senior hardware development manager at Annapurna Labs, an Amazon company.
Software completes the AI factory. NVIDIA NCCL supports distributed workloads, while NVIDIA Dynamo and NIXL help manage distribution and data movement. NVIDIA Mission Control provides cluster management, telemetry, and debugging capabilities, allowing operators to run mixed AI infrastructure as a coordinated system.
NVLink Fusion makes custom XPUs compatible with advanced AI platforms, enabling hyperscalers and AI-native companies to build unified, semi-custom AI factories. By combining the strengths of multiple technology and manufacturing partners, organizations can create infrastructure that would be difficult to build independently.
Learn more about NVLink Fusion.
Source: blogs.nvidia.com


