According to OpenRouter data, agentic AI workloads consume up to 15 times more tokens than simple chat requests. The reason is that AI agents perform multiple steps, use tools, call subagents and continuously process expanding context before completing a task.
For example, an AI agent researching companies for investment decisions may query financial databases, search news and regulatory filings, ask subagents to conduct peer comparisons, evaluate financial models and combine the results into a final recommendation. Each step adds new context and may trigger additional reasoning, tool calls and model responses.
The tokens generated at one stage become input for the next. As a result, long-context processing and efficient token generation are central to agentic AI performance.
The same pattern appears across agentic AI applications, including software development, customer service, research and complex business analysis.
As organizations deploy AI agents in production, the infrastructure supporting these workloads must generate and process tokens efficiently while operating within real-world power and cost constraints.
New performance measurements show that NVIDIA Vera Rubin NVL72 systems deliver up to 30 times higher throughput per megawatt than NVIDIA GB300 NVL72 systems on agent workloads. NVIDIA measured this inference performance using the SemiAnalysis AgentX workload. AgentX consists of recorded real-world coding sessions that preserve context augmentation, tool calls and subagent generation.
For a power-constrained AI factory, this performance difference could enable up to 30 times more agent work using the same amount of energy. These early Vera Rubin NVL72 results highlight the rapid pace of innovation in AI inference hardware and software. Ongoing software optimization is expected to improve the performance of both Vera Rubin NVL72 and GB300 NVL72 systems.
NVIDIA Vera Rubin NVL72 Delivers 30x Higher Throughput per Megawatt and 35x Lower Token Cost
Agent workloads differ significantly from chat and document summarization. Traditional input and output sequences often range from 1,000 to 8,000 tokens. During an agent session, however, context accumulates across multiple steps and can reach hundreds of thousands of input tokens. Request lengths can also vary widely, making a single-request benchmark insufficient for measuring real-world agentic AI performance.

For this reason, performance benchmarks must evaluate the complete agent workflow rather than measuring only one inference request. The results below reflect performance measured on real-world agent coding trajectories.
In SemiAnalysis AgentX testing, the NVIDIA Blackwell platform demonstrated strong performance across multiple agent models, including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro.
For example, NVIDIA GB300 NVL72 delivers up to 15 times higher throughput per megawatt than NVIDIA Hopper on the DeepSeek V4 Pro model. This provides customers with a high-performance foundation for deploying agentic AI workloads. The improvement is enabled by a larger GPU domain and software co-designed to increase inference efficiency.
NVIDIA Vera Rubin extends this advantage across the performance and efficiency curve, delivering up to 30 times higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model. These initial results were measured with the SemiAnalysis AgentX workload and are currently pending SemiAnalysis review. They also do not yet include Vera CPU performance for tool calls.

NVIDIA DSX MaxLPS technology manages power across the GPU, rack and workload levels. By optimizing power distribution, it can provision up to 40% more GPUs within the same megawatt budget, further increasing throughput per megawatt at AI factory scale.
Throughput per megawatt also directly influences the cost of generating tokens. NVIDIA Vera Rubin NVL72 can deliver up to 35 times lower cost per million tokens than GB300 NVL72, helping organizations run AI agents continuously and at production scale across a wide range of workloads.

For power-constrained AI factories, throughput per megawatt helps determine potential AI factory revenue. The cost per million tokens then influences the return generated from that revenue.
Collaborative Hardware and Software Design for Agentic AI at Scale
Modern inference optimization uses multiple techniques that are especially important for agentic AI. NVIDIA Vera Rubin NVL72 combines hardware and software co-design across the platform to improve long-context processing, token generation and overall inference efficiency.
- Disaggregated serving separates context processing, known as prefill, from response generation, known as decoding, allowing each stage to scale independently.
- Rate matching synchronizes prefill and decode workloads so GPUs can generate tokens efficiently and avoid unnecessary idle time.
- Massive expert parallelism distributes the subnetworks of mixture-of-experts models across GPU domains to support larger and more efficient model execution.
- Distributed KV caching extends memory across scale-up GPU domains. KV cache offload moves less-active contexts to host memory and storage, enabling previously processed context to be reused without recomputation.
- KV-aware routing directs incoming requests to GPUs that already contain the relevant cached context, reducing redundant computation during long agent sessions.
- Fused CUDA kernels, including MegaMoE, combine compute and GPU-to-GPU communication operations into a unified execution path. This keeps the GPU active instead of waiting for data transfers.
The fifth-generation tensor cores in NVIDIA Vera Rubin GPUs and the third-generation Transformer Engine accelerate both prefill and decode stages. NVFP4 quantization compresses model weights to 4-bit precision, reducing memory requirements and increasing throughput while maintaining output quality.
The NVL72 scale-up domain used by Vera Rubin and Grace Blackwell enables high-bandwidth, low-latency GPU-to-GPU communication. This connectivity is essential for massive expert parallelism, distributed KV caching and other long-context inference techniques.
NVIDIA NVLink interconnect technology and NVLink switches, now in their sixth generation, are designed for this scale-up architecture. They deliver up to 10 times higher packet rates and three times lower latency than standard Ethernet alternatives.
The platform also includes optimized CUDA kernels, inference runtimes such as NVIDIA TensorRT-LLM and service frameworks such as NVIDIA Dynamo. Together, these components form a software stack co-designed with the hardware to improve agentic AI inference performance.
The results above reflect performance from the current Vera Rubin NVL72. The complete platform is a seven-chip architecture that also includes an NVIDIA Vera CPU, Groq 3 LPX, NVLink 6 switches, BlueField-4 DPUs, Spectrum-6 switches and ConnectX-9 SuperNICs. These components are purpose-built for AI factories deploying intelligent agents at scale.
This hardware and software co-design extends to collaboration between NVIDIA and its ecosystem partners. Vera Rubin is designed to support the next generation of high-performance, energy-efficient agentic AI infrastructure.
For more information, visit the NVIDIA Vera Rubin Platform.
Source: blogs.nvidia.com


