AI has officially entered the gigascale era.
The world’s most advanced AI factories now harness the power of hundreds of thousands of GPUs and CPUs to train frontier models, drive agent AI, and generate intelligence on an unmatched scale. This significant advancement enables the network to double its computing power, crucial for effective token generation.
In celebration of this networking milestone, NVIDIA Spectrum-6, a revolutionary 102.4 Tbit/s Ethernet switch system, is being deployed in gigascale AI factories worldwide, delivering double the capacity of its predecessors and integral to the NVIDIA Vera Rubin platform.
Spectrum-6 represents the next generation of the NVIDIA Spectrum-X Ethernet Platform, providing the bandwidth, scalability, and intelligence essential for operating AI factories as a cohesive end-to-end computing system.
Leading AI Builders Take Charge
Top AI infrastructure builders, including CoreWeave, Microsoft, Nevius, Space XAI, and Tesla, are among the first to implement Spectrum-6, rapidly accelerating their AI operations.
For cloud service providers, Spectrum-6 enhances computing capabilities, creating a unified, high-performance resource that enables quicker model training and faster deployment of inference services.
“CoreWeave is designed for the most demanding AI workloads, where networking is critical for delivering performance at scale,” says Min Jun, Director of Network Products at CoreWeave. “With NVIDIA Spectrum-6 and our liquid-cooled Spectrum-X Ethernet infrastructure, we can offer our customers the bandwidth, reliability, and efficiency they need for swift frontier model training and inference deployment.”
“At gigascale, tuning performance is vital: ensuring all GPUs stay synchronized to prevent one slow link from derailing an entire job,” states Laurel Roseman, Vice President of Global Partnerships at Nebius. “NVIDIA Spectrum-6 embodies this capability, which is why we were early adopters: it provides a fabric that remains robust and fast, even as our customers’ workloads grow.”
For AI innovators establishing their own infrastructure, Spectrum-6 facilitates synchronized operation of more GPUs, enhancing efficiency during intensive collective tasks and ensuring greater resilience for prolonged jobs.
For every scenario, Spectrum-6 promotes faster time-to-results and outstanding economies of scale.
CoreWeave, Microsoft, and Nevius will be among the first to deploy NVIDIA Vera Rubin-based infrastructure utilizing Spectrum-6, broadening access to this platform for a large community of developers, startups, and enterprises.
AI Performance is a Network Concern
Relying solely on peak GPU performance is no longer sufficient for assessing AI factory performance.
Large-scale training and inference ventures depend on thousands of accelerators that must consistently exchange data. Collective communication is vital for synchronizing operations across GPUs, generating substantial east-west traffic from multiple systems at once.
Traditional Ethernet was designed primarily for enterprise applications and north-south data traffic between users, servers, and storage. It lacks the capabilities for the synchronized, high-volume communication patterns necessary for gigascale AI.
NVIDIA Spectrum-X Ethernet changes that paradigm. Specifically crafted for AI, it transforms Ethernet into a high-performance scale-out fabric that ensures all GPUs are consistently fed with data.
Next Generation Spectrum-X Ethernet
Combining Spectrum-6 switch chips with NVIDIA ConnectX-9 SuperNICs, the next-generation Spectrum-X Ethernet is integrated into the NVIDIA Vera Rubin architecture. Explore more about it here. The system consists of NVIDIA Vera CPUs, Rubin GPUs, NVLink 6 switches, ConnectX-9 SuperNICs, BlueField-4 DPUs, and the Spectrum-6 Ethernet switch.
Spectrum-6 accommodates both pluggable and co-packaged optical form factors. Moreover, it supports liquid cooling, allowing for a comprehensive end-to-end cooling solution for the entire AI factory while enhancing network power efficiency.
In contrast to generic Ethernet setups, the NVIDIA Spectrum-X Ethernet networking platform for AI factory scale-out integrates intelligent switches, NVIDIA ConnectX-9 SuperNICs, and comprehensive networking software, all tailored for AI. This platform continually optimizes traffic flow through the network fabric.
Furthermore, NVIDIA Spectrum-X technology adeptly balances traffic across available pathways, ensuring rapid failure avoidance and accurate recovery when data en route cannot reach its target. Additionally, with support for open network operating systems and diverse RDMA transport models, AI builders enjoy flexibility without sacrificing performance.
The Spectrum-X Ethernet platform exemplifies NVIDIA’s integrated, open approach, featuring co-designed silicon, systems, and software across computing platforms, while supporting standard Ethernet, open network operating systems, open protocols, and a vast ecosystem of cloud providers, system manufacturers, and infrastructure partners. Customers receive a complete AI factory platform designed to achieve the fastest training times and lowest token costs, eliminating the need for piecing together components and optimizing them later.
For demonstration purposes, Spectrum-X Ethernet achieves up to 1.6 times higher AI networking performance than standard Ethernet and maintains an impressive 95% network efficiency during deployments involving over 100,000 GPUs.
The hardware-accelerated Spectrum-X multiplane topology reduces the number of required switches in your data center by 1.7 times, thereby enhancing network efficiency. Building on Spectrum-X Ethernet photonics’ advantages, such as 5 times power efficiency and 10 times average time-to-incident improvements, ongoing network fabric enhancements are set to accelerate workloads beyond component-level optimization.
For additional information, visit the NVIDIA Spectrum-X Ethernet platform page.
Source: blogs.nvidia.com


