Nvidia Vera Rubin is revolutionizing AI infrastructure with monumental capabilities.
The production of the Vera Rubin NVL72 is powered by cloud leaders like CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure. With over 350 factory locations in 30 countries, Vera Rubin boasts the largest and most mature rack-scale supply chain to cater to the soaring computing needs of its clientele.
Crafted for optimal efficiency, this core weave benchmarking highlights a 10x output per megawatt compared to the Grace Blackwell NVL72, directly targeting the metrics crucial for AI factories constrained by power.
Enhancing Performance Through Collaborative Design
This innovation stems from unprecedented co-design among seven chips and five rack trays (Vera Rubin NVL72, Vera CPU Rack, Groq 3 LPX, Spectrum-6 SPX, and Vera BlueField-4 STX) engineered as a cohesive system rather than merely assembled from standard components.
At its core, the NVIDIA Vera CPU redefines CPU capabilities. Designed for the era of agents, the custom Olympus core boosts single-threaded performance by 2x, inter-core bandwidth by 3x, and reduces memory latency by 40% compared to competitive chiplet architectures, making it the most adept CPU for demanding agent workloads.
Accelerate Your AI Factory with a Dedicated Network
On the networking front, the 6th generation platform, NVLink scales up to yield over 2x the throughput, 3x lower latency, and 10x the packet rate for intricate workloads compared to conventional Ethernet.
Industry giants like CoreWeave, Microsoft, SpaceX AI, and Tesla are among the early adopters deploying Spectrum-6 switches to enhance their AI factories. The NVIDIA Photonics system showcases the industry’s first high-volume switch with 5x lower power consumption and 10x improved MTBI over traditional transceivers.
The Spectrum-XGS Ethernet enables performance across multiple sites with a 1.9x uplink throughput, addressing the needs of gigascale AI solutions.
Additionally, NVLink Fusion expands the NVIDIA infrastructure to include third-party XPUs, facilitating a rapid market entry for partners leveraging the established NVLink ecosystem.
Reduce Setup Time and Water Usage
NVIDIA’s innovative three-generation rack-scale co-design has led to the creation of the Vera Rubin NVL72 system, eliminating the need for cables, fans, or hoses, which shortens assembly time from hours to just one minute.
With a liquid cooling inlet temperature designed for 45 degrees Celsius, this system allows dry cooler operation without requiring a chiller, ultimately conserving millions of gallons of water per megawatt yearly through its high-temperature dry cooling solution.
Tuesday, July 21, 8 a.m. Pacific Time 🔗
NVIDIA Vera Rubin Drives Open Model Innovation in Europe
The Vera Rubin platform is set to deliver revolutionary performance for Europe’s AI infrastructure.
This initiative underpins a significant partnership between Microsoft and Mistral, aiming to enhance AI capabilities in the region by pairing Europe’s open model with a customer-managed environment, enabling localized AI deployment.
Central to this alliance is a multi-billion dollar agreement designed to bolster AI infrastructure in Europe. Mistral utilizes thousands of NVIDIA Vera Rubin GPUs to enhance GPU capacity, ensuring widespread availability of AI computing resources for clients.
NVIDIA Vera Rubin integrates seven co-designed chips into a single system, leveraging tens of thousands of GPUs to empower cutting-edge Mistral computing and bolster Microsoft’s European AI framework.
Europe aims to establish a formidable AI infrastructure, designed to operate under its own regulations and entirely within its jurisdiction. The bar is set high: Agent systems can utilize up to 15x more tokens than standard AI applications, necessitating a focus on scalable, efficient infrastructure that respects local governance and data protection needs.
Vera Rubin serves as the backbone of Europe’s open model ecosystem, combining accelerated computing, advanced networking, and robust software to seamlessly support the full spectrum of AI functionalities from training to deployment.
An Open Model Tailored for Enterprise AI
In collaboration, Microsoft, Mistral, and NVIDIA are set to deliver sovereignty-ready AI across both public and private cloud environments, ensuring the operational flexibility and scalability that modern enterprises demand.
Mistral Medium 3.5 and OCR 4 are now deployed on Microsoft Foundry, with Mistral models integrated into Microsoft Copilot Studio. Through Azure Local and Foundry Local, customers can utilize the same models and tools across cloud-hosted and customer-driven environments.
AI Tailored to European Standards
Government bodies can leverage AI for sensitive tasks while healthcare and finance sectors can align operations with local regulations. Manufacturers may also process data in real-time for swift decision-making. Compared to the NVIDIA GB200 NVL72, the Vera Rubin NVL72 provides up to 10x more tokens per megawatt, reducing costs significantly while ensuring enhanced performance.
Sovereign AI should empower organizations to innovate without compromising control or costs. With Vera Rubin as the foundational computing platform alongside Microsoft and Mistral’s sovereign cloud offerings, Europe stands to achieve all three objectives.
For further details, read the press release.
Tuesday, July 21, 8 a.m. Pacific Time 🔗
NVIDIA Vera Rubin NVL72 Delivers 10x More Tokens per Megawatt with CoreWeave

In collaboration with NVIDIA, CoreWeave ushers in an unparalleled performance leap with the Vera Rubin NVL72.
To harness the next-generation NVIDIA accelerated computing powers for your AI factory, partnership with leaders like CoreWeave is essential.
Following extensive engineering collaboration, CoreWeave has become the first AI cloud provider to successfully launch and validate the Vera Rubin NVL72, providing initial performance data from operational hardware.
CoreWeave’s benchmark of the DeepSeek-R1 on the Vera Rubin NVL72 revealed a 10x improvement in tokens per megawatt against the Grace Blackwell NVL72.
Tokens per megawatt is a crucial metric measuring the profitability of AI infrastructure scalability, indicating that more intelligent outputs can be achieved within the same power constraints. AI firms, including Jane Street, are leveraging the Vera Rubin platform to expand their AI capabilities using the CoreWeave cloud.
Overcoming Bottlenecks with NVIDIA Spectrum-X
The DeepSeek-R1 utilizes a MoE architecture mandating extensive GPU communication across networks. Through the Vera Rubin NVL72’s 260 TB/s all-to-all NVLink 6 fabric, this limitation is overcome, enabling seamless operation of the rack as a unified accelerator.
CoreWeave was among the first to adopt NVIDIA Spectrum-X Ethernet SN6600-LD as the switching fabric for use with the Vera Rubin NVL72.
Equipped with the 102.4 Tb/s Spectrum-6 switch chip and a liquid-cooled design, CoreWeave now boasts high-density switching racks offering 1.64 Pb/s per rack, doubling capacity over prior air-cooled models, along with a fully non-blocking, multiplane, multirail fabric that connects Vera Rubin NVL72 GPUs without oversubscription.
Tuesday, July 21, 8 a.m. Pacific Time 🔗
NVIDIA Vera Rubin NVL72 Powers Google Cloud A5X Instances for Unprecedented Intelligence

The NVIDIA Vera Rubin NVL72 drives Google Cloud’s newly launched A5X instance, currently operational at the London startup, Ineffable Intelligence.
Ineffable Intelligence is at the forefront of creating “super learner” systems capable of continuous learning through experiences, leading to breakthroughs across numerous domains.
Rather than relying on static datasets, the agents at Ineffable Intelligence learn dynamically from their environment, employing reinforcement learning to derive experiences through massively parallel simulations, rapidly converting those into actionable policy updates.
“The next era of research mandates cutting-edge hardware,” stated Lasse Espeholt, co-founder of Ineffable Intelligence. “We are grateful for the collaboration with NVIDIA and Google Cloud, which granted us early access to Vera Rubin, allowing us to expedite our infrastructure tests for super learner systems.”
The NVIDIA Vera Rubin NVL72 is specifically engineered for such advanced agent training needs, providing predictable latency, high utilization rates, and significantly enhancing intelligence per dollar compared to previous generations.
The A5X instances, introduced at Google Cloud Next, are bare-metal instances based on the NVIDIA Vera Rubin NVL72 rack-scale system, offering up to 10x lower inference costs per token and 10x greater token throughput per megawatt than earlier versions.
Utilizing NVIDIA ConnectX‑9 SuperNICs combined with Google’s innovative Virgo networking, A5X instances can scale to handle tens of thousands of NVIDIA Rubin GPUs within a single data center and approach nearly 1 million GPUs across multiple sites, providing a comprehensive AI optimization journey from training to deployment, all while balancing performance, cost, and sustainability.
This infrastructure is poised to unlock the next wave of reinforcement learning advancements, paving the way for breakthroughs in superlearning and superintelligence.
Tuesday, July 21, 8 a.m. Pacific Time 🔗
NVIDIA Vera CPUs Enhance Speeds for DeepInfra AI Cloud

Benchmark results from DeepInfra confirm that NVIDIA Vera CPUs outperform traditional CPUs, offering over 2x the speed and enabling greater simultaneous AI agent support.
Cloud platform DeepInfra, an early participant in the NVIDIA Open AI Ecosystem, executed the benchmark with its AI agent infrastructure processing nearly 5 trillion tokens each week, approximately 30% of which are handled by agent systems.
These benchmarks illustrate that NVIDIA Vera CPUs can support up to 1.6x more concurrent AI agents without compromising service quality and achieve up to 2.2x faster orchestration compared to standard CPUs, while simultaneously enhancing infrastructure efficiency and cost-effectiveness.
As AI agents navigate increasingly complex inference, planning, and data flow, CPU performance is critical for effective coordination of model processes.
This joint design effort from NVIDIA prioritizes agent workloads, showcasing how NVIDIA Vera CPUs can assist cloud services in optimizing resource use, lowering costs, and supporting a higher number of simultaneous AI agents with the same level of service quality.
For additional information, visit the NVIDIA Vera Rubin Platform.
Source: blogs.nvidia.com


