Power management is a critical challenge for AI infrastructure. The ability to generate electricity within a defined power budget directly influences the profitability of an AI factory. Thus, optimizing performance per watt—an irrefutable metric measurable through genuine results—is foundational to any AI Factory.
As agent AI escalates the demand for tokens, the infrastructure strategies organizations implement today will dictate who scales successfully in a power-limited environment.
Currently, most cutting-edge AI models are designed using a mixture of experts (MoE) architecture. To effectively serve these large-scale models, the scale of GPU deployment (the number of GPUs linked through ultra-fast interconnections) is crucial: larger configurations are more effective.
The NVIDIA Hopper generation has set benchmarks in the 8-GPU space, but today’s advanced AI implementations have surpassed that threshold. Achieving MoE in a GPU environment necessitates comprehensive code design and operational expertise derived from deploying these models under actual production loads.
The NVIDIA Blackwell NVL72 Platform provides a robust and tested foundation, delivering optimal performance per watt to enhance revenue and minimize token costs to increase profit margins. This foundation is bolstered by the NVIDIA Vera Rubin platform which further advances rack-scale energy efficiency.
Maximize Performance Per Watt with Frontier AI
Each new generation of Frontier models introduces architectural innovations that enhance intelligence while necessitating new optimizations for efficient scaling.
Among the latest generation of top-tier open models, NVIDIA GB300 NVL72 achieves up to 25x more performance per watt compared to the NVIDIA Hopper generation. This demonstrates that expanding the GPU domain from 8 GPUs to 72 GPUs significantly improves MoE performance. These figures showcase the current capabilities of Blackwell, which continue to evolve.
A single metric only provides a partial view. Various workloads necessitate distinct operational points—some prioritize latency, while others emphasize throughput and cost-efficiency. Many workloads need to transition between these priorities.
To accurately illustrate these operational points, NVIDIA provides a Pareto curve for each model instead of relying on a single figure, and offers tools such as: Dynamo Sim that empowers teams to identify the optimal point on the Pareto frontier before consuming any GPU resources for validation.


The impressive performance per watt of NVIDIA Blackwell results from a meticulous co-design of all rack-scale system components, from silicon to software, engineered to maximize token throughput for AI inference tasks. This collaborative design integrates all layers of the technology stack.
For example, the NVIDIA NVLink Switch plays a pivotal role in rack-scale performance, purpose-built to support large-scale GPU domains and enabling in-network computing directly within the switch, thus relieving the GPU of some workload. Now in its sixth generation, powered by the Vera Rubin platform, it is specifically tailored for AI tasks such as SHARP.
The NVIDIA Inference Software Stack is designed to execute numerous optimizations, including NVIDIA Dynamo and TensorRT LLM, alongside SGLang and vLLM, featuring NVFP4 quantization, disaggregated serving, massive expert parallelism, KV-aware routing, and KV cache offloading. These integrations multiply the performance from each GPU. Additionally, the software continually improves performance over time. DeepSeek V4 has recorded up to a 5x increase in performance per watt within a month.
In an AI factory, typically only 60% of power drawn from the grid translates into meaningful AI computations, due to losses from cooling and inefficiencies at the rack level. NVIDIA DSX MaxLPS optimizes power efficiency. NVIDIA DSX not only bridges this gap by dynamically reallocating power between the GPU and the rack but also supports hot-water liquid cooling and utilizes innovations like power steering to enhance performance, enabling operators to run up to 40% more GPUs within the same power budget.
It’s All About Production
Fostering rack-scale reliability at AI factory scale is a challenging task. Rack-scale systems introduce unique failure modes not encountered in single-node deployments, necessitating rigorous engineering and substantial production time to tackle these issues effectively.
NVIDIA Blackwell NVL72 systems continue to set the industry benchmark across a range of models and production use cases, ensuring sustained performance, rack-level reliability, and economics that can withstand real-world operational conditions.
That’s why leading AI laboratories like Anthropic, OpenAI, and SpaceXAI leverage NVIDIA Blackwell NVL72 systems for their inference operations.
Moreover, various inference service providers and AI-centric organizations utilize the Blackwell platform to implement open models in production environments.
CoreWeave’s deployment of K2.6 on the NVIDIA GB300 NVL72 combines NVFP4 quantization and EAGLE3 speculative decoding to optimize inference performance.
Perplexity runs Quen 3 235B and Qwen 3.5-397B-A17B to power its AI agent platform, processing millions of queries daily with the requisite latency and reliability.
Fireworks AI is implementing GLM 5.2 on the NVIDIA Blackwell platform to support customer production deployments including Cursor and Factory AI.
This accumulated production knowledge, derived from multiple generations of frontier models and real-world deployments, provides NVIDIA Vera Rubin with a significant advantage in innovation.
Discover more about the NVIDIA Vera Rubin platform. Read our technology blog for additional insights, NVIDIA DSX AI Factory Scale Platform and DSX MaxLPS.
Source: blogs.nvidia.com


