NVIDIA NVLink Fusion and NVHBM Accelerate Next-Generation AI Infrastructure
The next generation of artificial intelligence is creating unprecedented demands on data center infrastructure.
As AI agents and trillion-parameter models become more common, AI performance depends on more than compute power alone. Compute, memory, storage, networking and software must work together as an integrated system to deliver efficient, scalable AI performance.
To help hyperscalers and AI innovators develop the next generation of semi-custom AI infrastructure, NVIDIA has announced NVIDIA NVLink Fusion and NVIDIA NVHBM.
NVIDIA NVHBM is a next-generation high-bandwidth memory technology designed to deliver greater memory performance and efficiency for XPUs. Validated and supplied by leading memory partners, NVHBM will also be available to customers using NVLink Fusion.
In traditional HBM architectures, the memory controller is located on the XPU die. This consumes valuable silicon area that could otherwise be dedicated to compute resources. NVHBM takes a different approach by integrating NVIDIA’s custom memory controller directly into the 3D HBM stack rather than placing it on the XPU.
By moving the memory controller into the HBM stack, NVHBM can provide:
- Up to 30% greater memory bandwidth for improved AI workload performance.
- Up to 15% lower HBM power consumption for greater energy efficiency.
- Up to 25% more available area on the XPU compute die compared with standard HBM4E designs.
NVIDIA has established a standardized NVHBM implementation that can be provided by multiple memory manufacturers. This approach reduces the engineering, integration and qualification work required to support memory from different suppliers, helping NVLink Fusion customers bring custom AI chips to market faster.
Amazon’s Annapurna Research Institute will be the first company to adopt NVHBM as part of a broader collaboration with NVIDIA focused on NVLink Fusion and next-generation AI infrastructure.
AWS and NVIDIA Expand Their NVLink Fusion Collaboration
Amazon’s Annapurna Research Institute is collaborating with NVIDIA on NVHBM technology and NVLink scale-up architectures to improve the performance, scalability and efficiency of AI workloads.
The collaboration builds on AWS’s previously announced support for NVLink Fusion. Annapurna Labs plans to support NVLink Fusion in next-generation Trainium chips, beginning with Trainium4. This will enable Amazon’s custom AI chips and NVIDIA GPUs to operate together within a common rack-scale architecture.
“NVHBM represents a new architectural approach to improving the performance and efficiency of high-bandwidth memory,” said Nafea Bshara, vice president of Annapurna Labs at Amazon. “We look forward to this technology collaboration, which will benefit future AWS infrastructure designs.”
Vertical Integration and an Open Ecosystem
NVIDIA NVLink Fusion enables partners to connect custom XPUs and CPUs to NVIDIA rack-scale computing platforms.
Through the platform, partners gain access to NVIDIA NVLink chiplets, NVLink-C2C, NVLink switches, NVIDIA MGX systems and racks, as well as a broad ecosystem of CPU providers, ASIC designers, system manufacturers and technology companies.
Delivered with each generation of NVIDIA rack-scale system architectures, NVLink Fusion gives hyperscalers and AI-native companies access to a proven technology stack for scale-up and scale-out networking, rack-scale systems and AI software.
This allows companies to focus their engineering resources on XPU innovation while relying on established NVIDIA technologies for connectivity, systems and software. As a result, NVLink Fusion provides a faster and lower-risk path to deploying semi-custom AI infrastructure.
Learn more about NVIDIA NVLink and NVLink Fusion.
Source: blogs.nvidia.com


