NVIDIA Vera Rubin NVL72 Delivers Up to 3.7× Higher AI Inference Performance in MLPerf v6.1
System performance, infrastructure scaling, and continuous software optimization are key factors in the economics of AI inference. Higher performance enables more tokens to be generated and can increase revenue. Efficient scaling helps throughput grow as hardware is added, while ongoing optimization delivers more value from existing infrastructure investments.
Underlying all three is platform fungibility. From training and inference to recommender systems, language models, and video workloads, the same infrastructure can run different models and workloads while maintaining high utilization.
The NVIDIA platform is designed to optimize performance, scaling efficiency, and software value, as demonstrated by the MLPerf Inference v6.1 results released today.
- NVIDIA Vera Rubin NVL72 makes its debut: In an initial MLPerf Inference preview submission, Vera Rubin NVL72 delivered up to 3.7× higher throughput than GB300 NVL72.
- NVIDIA GB300 NVL72 scales efficiently: A submission spanning 288 GPUs across four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput increasing almost linearly from a single-rack baseline.
- Software optimization continues to improve performance: Software improvements in NVIDIA’s MLPerf Inference v6.1 submission delivered up to 1.6× higher performance than v6.0. Additional optimizations after the v6.1 submission produced further gains.
For organizations planning AI infrastructure, performance, scaling efficiency, and software optimization are important factors in determining long-term inference economics.
Vera Rubin NVL72 makes its MLPerf Inference debut
NVIDIA submitted Vera Rubin NVL72 preview results for two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL.
On Qwen3-VL, Vera Rubin NVL72 delivered up to 3.7× higher throughput than GB300 NVL72 across offline, server, and interactive scenarios. These results used vLLM with the NVIDIA Dynamo open-source inference framework. On DeepSeek-R1, Vera Rubin NVL72 delivered up to 2.5× higher throughput than GB300 NVL72 using the NVIDIA TensorRT-LLM library.
These early results demonstrate NVIDIA’s pace of innovation and the impact that continuous software optimization can have on AI inference performance.
Higher performance means each Vera Rubin NVL72 rack can deliver more tokens than a GB300 NVL72 rack, serve more users, and potentially lower the cost per token while increasing revenue.
The results reflect full-stack co-design across hardware and software. Vera Rubin’s enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages of inference. NVFP4 precision reduces model-weight, attention, and overall KV-cache memory requirements, increasing throughput with minimal loss in output quality.
Vera Rubin’s submission also made extensive use of decoupled serving to separate prefill and decode, along with large-scale expert parallelism to improve system efficiency for mixture-of-experts models such as DeepSeek-R1 and Qwen3-VL.
With sixth-generation NVIDIA NVLink and NVLink switches, the NVL72 scale-up domain delivers 10× higher packet rates and 3× lower latency than off-the-shelf Ethernet. This provides the interconnect foundation required to make rack-scale AI inference effective.
The full-stack approach also extends to NVIDIA’s partner ecosystem. Nebius submitted preview results for Vera Rubin NVL72, demonstrating strong performance.
AI agents that reason, plan, and act across multiple steps are changing how inference performance is measured. In benchmarks designed to capture these workloads, including SemiAnalysis AgentX, Vera Rubin NVL72 achieved a 30× performance result in preview testing. The GB300 also delivered stronger performance than the NVL72 in preview testing. The upcoming MLPerf Endpoints benchmark is expected to provide standardized measurements for agent inference workloads beyond traditional throughput benchmarks.
GB300 NVL72 achieves 99% AI infrastructure scaling efficiency
Scaling efficiency measures how effectively additional GPUs increase throughput. NVIDIA achieves this through high-bandwidth, low-latency scale-up interconnects within each rack, high-bandwidth networking between racks, and efficient request orchestration across nodes.
NVIDIA’s DeepSeek-R1 submission scaled from one GB300 NVL72 rack with 72 GPUs to four racks with 288 GPUs. In offline scenarios, the system achieved 99% scaling efficiency, with throughput increasing almost linearly as hardware was added.

Scaling efficiency matters because adding GPUs does not automatically increase throughput proportionally. If nearly twice as much hardware produces only a single-digit percentage increase in throughput, infrastructure costs can outweigh the performance gains. Architecture, interconnects, and software must scale together.
The GB300 NVL72 also demonstrated rack-scale efficiency on the WAN 2.2 text-to-video benchmark. It reached 0.65 720p videos per second at 5.7 seconds per video, delivering 9× higher throughput and 7.5× lower latency than a single node.
Continuous software optimization improves AI inference economics
The NVIDIA platform undergoes continuous software development to improve performance and functionality across AI workloads.
In MLPerf Inference v6.1, GB300 NVL72 performance on Qwen3-VL improved by approximately 1.6× compared with v6.0 results. The improvement was driven by lower KV-cache accuracy, kernel fusion, improved kernels, and distributed services using vLLM and NVIDIA Dynamo.
Software optimization continued after the v6.1 submission deadline. Post-submission results for GPT-OSS-120B and DLRMv3 have not yet been verified by MLCommons, but they show additional performance improvements.

AI inference performance from the edge to the data center
In addition to results for the NVIDIA Grace Blackwell and Vera Rubin NVL72 platforms, NVIDIA submitted results for Jetson AGX Thor using NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B.
NVIDIA’s partner ecosystem was broadly represented, with 19 partners, including eight partners on multi-node Blackwell NVL72 systems, demonstrating strong performance. The partners included ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro, and Wiwynn.
NVIDIA continues to improve performance across its technology stack, from compact edge devices to large AI factories, while advancing the software and ecosystem needed to deliver AI inference at scale.
For more information, see the NVIDIA Vera Rubin Platform.
Source: blogs.nvidia.com


