NVIDIA Highlights AI Factory Efficiency, Power Optimization and Agentic AI Performance
Ian Buck, NVIDIA’s vice president of hyperscale and high-performance computing, discussed AI factory efficiency Tuesday at the AI Infrastructure Summit at the Santa Clara Convention Center.
The event drew more than 8,000 attendees this year, up from 3,500 last year. Buck highlighted new collaborations across NVIDIA platforms, along with performance and efficiency results from NVIDIA partners.
- Amazon’s Annapurna Research Institute: Working with NVIDIA on NVHBM custom high-bandwidth memory technology.
- d-Matrix: Integrating the d-Matrix Raptor XPU with the NVIDIA NVLink Fusion platform.
- Emerald AI: Demonstrating NVIDIA’s Commercial AI Factory Flexible Load program with Silicon Valley Power.
- Lambda: Reporting a 23% improvement in performance per watt with NVIDIA DSX MaxLPS.
- Pinterest: Bringing conversational AI to visual discovery with the NVIDIA Blackwell platform and NVIDIA Dynamo inference software.
The announcements come as agentic AI drives a new class of workloads that requires greater performance, efficiency and scale from AI infrastructure.
NVIDIA’s full-stack AI Factory platform includes the Vera Rubin system, Dynamo inference software, NeMo libraries and NVIDIA NVLink for scale-up computing. It also includes Spectrum-X Ethernet and ConnectX SuperNICs for connecting thousands of nodes, along with BlueField-powered context memory storage and BlueField DPUs for infrastructure security.
AI infrastructure metrics are increasingly moving beyond peak performance toward verified agent tokens per megawatt. AI factories must now be co-designed from silicon to the power grid. NVIDIA says DSX MaxLPS can deliver up to 1.4x more tokens per megawatt through factory-wide power optimization, while NVLink helps consolidate large-scale accelerated computing into a single high-performance system.
The goal is an AI infrastructure platform that generates more tokens, improves efficiency and helps customers get more value from every megawatt of power.
Emerald AI and NVIDIA Demonstrate Flexible AI Factory Load Management
Silicon Valley Power operates a flexible load interconnection program designed to help AI factories support grid flexibility. Emerald AI worked with NVIDIA to demonstrate automatic load shedding through the program.
The system responded to hundreds of request signals from Silicon Valley Power while protecting the performance of AI workloads.
Emerald AI plans to use NVIDIA DSX Flex for grid-aware power management. The software can dynamically adjust an AI factory’s energy consumption based on real-time grid signals and hybrid energy sources.
DSX Flex can autonomously balance workloads by throttling the power used by low-priority AI jobs and returning those jobs to normal operation when conditions allow.
The software receives signals such as load-shedding requests, demand-response events and pricing changes. It then operates within predefined workload hierarchies: high-priority jobs continue running, while other workloads can be temporarily paused and resumed later.
This capability allows AI factories to act as flexible grid resources. Facilities can reduce demand when the grid needs relief, protect critical AI workloads and demonstrate how controllable AI loads can help free grid capacity for further growth.
Learn more about the Emerald AI partnership with Silicon Valley Power.

Lambda Reports 24% Higher Token Throughput With NVIDIA DSX MaxLPS
AI cloud provider Lambda presented results from the first validation of NVIDIA DSX MaxLPS on NVIDIA Blackwell servers at the AI Infrastructure Summit.
DSX MaxLPS continuously monitors power consumption across GPUs and racks. It dynamically shifts available power to the areas where it is most needed and reclaims capacity that would otherwise remain unused under static power provisioning.
Because training and inference workloads have different power profiles, DSX MaxLPS optimizes power allocation across an AI factory running mixed workloads.
Lambda ran 19 nodes within the same power budget typically allocated to 16 full-power nodes. The company reported a 24% increase in cluster token throughput, from approximately 4 million to 5 million tokens per second, along with a 23% improvement in performance per watt.
With the next-generation NVIDIA Vera Rubin NVL72 AI Factory, DSX MaxLPS can enable up to 40% more GPU capacity within the same megawatt budget in the right deployment.
For AI factory operators, the results could mean greater AI capacity and token throughput within the same power envelope, increasing productivity and economic value per megawatt.
Learn more about Lambda’s results.

NVIDIA Vera Rubin and Groq 3 LPX Convert More Power Into Tokens
Power is a central constraint for AI factories. As a result, AI infrastructure operators are focused on generating more tokens from every megawatt.
NVIDIA is addressing this challenge with a full-stack AI factory built on Vera Rubin NVL72. The platform combines systems, networking, software and power management. At the factory level, NVIDIA DSX MaxLPS dynamically shifts power between racks as demand changes.
The platform is designed to deliver:
- Up to 40% more GPUs within the same site power envelope.
- Up to a 35% increase in token throughput without requiring new power lines.
Inside each rack, intelligent power-smoothing software and enhanced energy buffering absorb short spikes and keep systems operating closer to sustained demand. This turns unused power headroom into productive computing capacity.
Agentic AI workloads can involve chained inference steps and tool calls, causing latency and context length to grow quickly. NVIDIA Groq 3 LPX adds deterministic, ultra-low-latency inference to Vera Rubin and complements DSX MaxLPS.
The unified platform delivers up to 35x higher token throughput per megawatt than GB200 NVL72 for models with more than 2 trillion parameters in long-context workloads.
On the Qwen 3.8 27B workload with 100,000-token context, Groq 3 LPX achieved 2,529 output tokens per second per user. This additional headroom allows agents to perform more inference steps and tool calls within the same response budget as workloads grow.
Learn more about how NVIDIA Vera Rubin NVL72 and Groq 3 LPX generate more tokens with less power.
Vera Rubin NVL72 Delivers Up to 30x Higher Agent Throughput per Megawatt
Rather than measuring a single request, SemiAnalysis AgentX evaluates inference using recorded real-world agent coding sessions. The benchmark maintains real-world context, tool invocation delays and subagent generation.
According to the SemiAnalysis AgentX Dashboard, NVIDIA Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on DeepSeek V4 Pro models.

Agent workloads differ from traditional request-and-response inference benchmarks. As agents reason, call tools and launch subagents, hundreds of thousands of input tokens can accumulate during a single session. That is approximately 15 times the number of tokens in a simple chat request, while input and output lengths can vary significantly.
AgentX is designed to measure the complete agent trajectory and complements standardized suites such as MLPerf Inference, which measure per-request performance across models and scenarios.
These efficiency gains directly affect AI factory economics. Up to 30x higher throughput per megawatt means up to 30x more agent work from the same energy usage. AgentX results also show up to 45x lower cost per million tokens.
For power-constrained deployments, throughput per megawatt affects the revenue generated by an AI factory, while cost per million tokens affects its margin.
The results are based on a combination of technologies, including an NVL72 scale-up domain, sixth-generation NVLink interconnect, NVFP4 precision on fifth-generation Tensor Cores, and an inference stack spanning NVIDIA TensorRT LLM and NVIDIA Dynamo. NVIDIA says Vera Rubin is fully operational and scaled across the ecosystem, with continued software optimization expected to improve performance.
Learn more about the platform architecture behind these results.
Startups Report Strong NVIDIA Vera CPU Performance for Agent Workloads
Several startups are testing NVIDIA Vera CPUs for agent workloads and other demanding applications.
Perplexity: Perplexity benchmarked Vera CPUs on its SPACE secure sandbox platform for agentic AI and reported that sandbox startup is now 1.9x faster. View the results.
Daytona: Daytona applied agent workloads to Vera CPU and reported significant benefits from carrying out executions by agents. Read the comment.
ClickHouse: ClickHouse shared Vera CPU results from ClickBench, an open benchmark for analytical databases. Vera was the fastest CPU the team had measured, which ClickHouse described as a strong sign of future CPU performance in data-intensive workloads. View the results.
Deep Infra: Deep Infra reported that Vera CPUs delivered superior performance across its measured metrics, including 2.2x faster orchestration-step latency. Read the benchmark.
Prime Intellect: Prime Intellect reported that Vera CPUs maintained high bandwidth and low, consistent memory latency as more workloads ran in parallel.
Redpanda: Redpanda reported that Vera CPUs delivered 5.5x lower latency and 73% higher throughput than other CPUs. Read the benchmark.
Starburst: Starburst reported that Vera CPUs achieved three times faster query throughput than other CPUs. Read the benchmark.
Kinetica: Kinetica observed 2.7x faster analytical query performance on Vera CPUs compared with traditional CPUs. Read the results.

NVIDIA NVLink 6 Improves AI Factory Reliability at Scale
As AI factories scale to hundreds of thousands of GPUs, reliability becomes a critical part of performance.
Training and inference workloads depend on continuous operation. Transient errors, signal degradation and hardware failures are inevitable challenges at this scale. NVIDIA NVLink 6 addresses them with a multi-tier resiliency architecture designed to detect, contain and recover from failures before they affect applications.
At the physical layer, custom forward error correction, physical-layer retry and universal physical-layer recovery help minimize latency impacts and maintain a lossless fabric.
At the network layer, credit-based flow control, dynamic routing and link rebalancing contain failures locally. These technologies are designed to prevent cascading stalls that could reduce AI factory throughput.
Source: blogs.nvidia.com


