How NVIDIA DSX Helps AI Factories Produce More Compute With the Same Power
As AI factories grow, power is becoming one of the industry’s most important constraints. NVIDIA DSX is designed to help data center operators increase useful computing capacity, respond to grid conditions and improve throughput without simply demanding more megawatts.
An early example came during a sweltering August evening in Silicon Valley. As the sun set and air-conditioning demand increased, Silicon Valley Power sent a signal to an AI factory asking it to adjust its electricity consumption.
Varun Sivaram of Emerald AI watched the response on Zoom with approximately 40 members of his team. In a conference room in San Francisco, data center engineers and power company representatives monitored the event. No one had to intervene manually.
Emerald AI’s Conductor platform received information about the state of the grid and automatically adjusted flexible computing workloads. Lower-priority work was slowed or rescheduled while high-priority services continued operating.
The objective was to reduce demand when the power grid was constrained without disrupting critical AI workloads. When the savings appeared on screen, the team cheered.
“We watched with bated breath,” Sivaram said. “This was the first time we were deploying across thousands of NVIDIA GPUs.”
Mansi Shah, Emerald AI’s head of product, described the moment as feeling “like a SpaceX rocket launch.”
Since then, more than 200 demand-response signals have been sent to the AI factory, and the system has responded successfully each time.
This production demonstration shows how flexible AI factories could help unlock additional computing capacity without waiting years for new transmission infrastructure. The approach also illustrates the broader goal of NVIDIA DSX: optimizing the entire AI factory rather than isolated components.
At the AI Infrastructure Summit, Ian Buck, NVIDIA’s vice president of hyperscale and high-performance computing, placed AI factory efficiency at the center of his infrastructure keynote.
Initial validation from cloud provider Lambda showed that intelligent power management can support 24% more token throughput within a fixed power budget.
“With our proof of concept, we believe we have surpassed the limitations of fixed power budgets,” said Dave Ward, president of cloud services at Lambda. “NVIDIA DSX MaxLPS paves the way for reclaiming pent-up capacity and converting it to real-world use while delivering significantly higher compute density in the same footprint.”
During the Silicon Valley demonstration, Conductor operated against a predefined workload hierarchy. The lowest-priority jobs were paused or rescheduled, high-priority inference continued running and power consumption fell from 4 megawatts to 3 megawatts—all automatically.
Why AI Factory Efficiency Matters
In the AI factory economy, the key metric is useful work per gigawatt. Data center operators are focused on efficiency at every layer, from facility design and cooling to rack-level power conversion and workload scheduling.
NVIDIA DSX extends that discipline to AI workloads. Smarter rack provisioning delivers power where workloads need it. Operational intelligence, including improved scheduling, faster restarts and leaner checkpointing, helps keep GPUs working instead of waiting.
“A 1-gigawatt factory will never become a 2-gigawatt factory,” said NVIDIA founder and CEO Jensen Huang.
When a system reaches physical limits, the answer is to stop optimizing individual parts and start designing the entire system. Announced at GTC Taipei in May, NVIDIA DSX applies that approach to AI factories through a platform covering networking, cooling, water efficiency, facility design and power management.
The following developments highlight how the platform is being applied to power management and grid participation.
DSX MaxLPS: More Compute Within the Same Power Budget
Lambda’s results presented at the AI Infrastructure Summit represent an initial validation of DSX MaxLPS on NVIDIA HGX B200 GPU servers.
DSX MaxLPS monitors GPU and rack-level power consumption, reallocates available power between nodes according to workload type and recovers capacity that could otherwise be lost through static provisioning. Because training and inference workloads consume power differently, dynamic allocation can improve utilization in AI factories running both.
Lambda, a GPU cloud provider serving more than 10,000 customers ranging from AI-native startups to hyperscalers, tested the software on a five-rack, 19-node cluster.
By running 19 nodes at 85% power and staying within the same power budget as 16 nodes operating at full power, Lambda increased cluster token throughput by 24%—from approximately 4 million tokens per second to 5 million tokens per second. The deployment also delivered 23% more performance per watt.
Based on NVIDIA projections, DSX MaxLPS could increase GPU capacity in next-generation Vera Rubin NVL72 AI factories by up to 40% within the same megawatt power budget in suitable deployments.
Automated Demand Response in Production
The Santa Clara demonstration was not a DSX Flex installation. It was an earlier commercial-scale proof that the underlying concept works.
NVIDIA’s Eos AI Factory operates as a participant in Emerald AI Conductor and Silicon Valley Power’s Flexible Load Interconnect Program. The program is designed to treat AI factories as dispatchable grid resources.
When Silicon Valley Power sends a signal, Conductor responds within a minute. Flexible factories can potentially operate at a larger scale by reducing demand when the grid needs relief and restoring workloads when conditions allow.
As Emerald AI Conductor is integrated into DSX Flex, this operating model is intended to become more common. The first dedicated DSX Flex commercial deployment will build on five previous demonstrations across two continents at the 96-megawatt Vera Rubin AI Factory at NVIDIA’s AI Factory Research Center in Manassas, Virginia.
800V DC Power Architecture for Denser AI Racks
The efficiency gains available in today’s AI factories are ready to deploy. The next challenge is powering denser and faster computing racks.
As AI factories expand, traditional low-voltage power paths can add conversion complexity and create power distribution constraints.
NVIDIA’s 800V DC architecture is designed to reduce conversion complexity, improve power delivery efficiency and support higher-density accelerated computing racks. NVIDIA DSX incorporates 800V DC into its reference design.
Optimizing the Whole AI Factory
No single component can optimize an AI factory by itself. A faster GPU may be limited by the network. Improper provisioning can leave power capacity unused. Cooling overhead can consume additional energy, and a GB200 NVL72 rack using direct liquid cooling can carry up to 120 kilowatts of heat that must be managed before it reaches the compute system.
The reliable path to increasing tokens per megawatt is to optimize the entire factory. DSX Sim can be used before the first rack is installed. DSX OS and DSX Exchange support operations after deployment, while DSX Reference Designs give builders a validated architecture to work from rather than requiring them to start from scratch.
Building a New Standard for Gigawatt Infrastructure
The central question is simple: how many useful jobs can an AI factory produce for every megawatt it consumes?
NVIDIA DSX provides infrastructure builders with reference designs, simulation tools, operational software and power-management technologies intended to help them meet next-generation AI infrastructure requirements.
When the power grid needed relief, the factory provided it without dropping critical jobs or requesting additional power. With NVIDIA DSX, that is the proposed new baseline for AI factories.
Source: blogs.nvidia.com


