NVIDIA AI Factory Economics: How Productivity, Durability and Flexibility Maximize Revenue
AI factories are built at the megawatt—and increasingly gigawatt—scale. Each megawatt of capacity costs approximately $60 million, so operators commit capital only after evaluating the expected return on investment.
Three factors determine AI factory revenue:
- Earning power: How much revenue a factory could generate in a year if it sold all the tokens it could produce.
- Service life: How long the AI hardware will remain profitable.
- Demand: How much demand exists for the tokens the factory produces.
No single factor can fully compensate for weakness in another. High production capacity means little if a factory sells only a fraction of its output. Strong demand matters less if the factory becomes unprofitable after one year. These factors are also connected: factories that support more workloads can attract more demand and continue generating revenue as the market evolves.
The NVIDIA AI Factory platform is designed to maximize all three:
- Productive: High throughput per megawatt and low cost per token maximize earning potential.
- Durable: NVIDIA GPUs and systems can continue generating revenue years after shipment, extending their service life.
- Fungible: NVIDIA platforms run many types of AI workloads, across every stage and location, as well as workloads that do not involve AI. This broadens and deepens demand.
Engineering co-design across the full stack maximizes AI factory throughput, while continuous software optimization keeps installed hardware productive years after shipment. NVIDIA CUDA-X libraries support accelerated workloads, while standardized architecture makes validated reference designs accessible to operators.
Productivity: More Tokens per Megawatt at Lower Cost
Power is a binding constraint on AI factories. That makes tokens per second per megawatt a critical measure of earning potential. The more tokens a factory produces within a fixed power envelope, the more revenue it can generate. A lower cost per token can also increase margins.
According to SemiAnalysis AgentX, the NVIDIA Vera Rubin NVL72 system delivers more than 30 times higher throughput per megawatt than the NVIDIA GB300 NVL72. The DeepSeek V4 Pro model also achieves up to 45 times lower cost per million tokens.
These gains come from coordinated optimization across the full stack, including models, workloads, software, compute, networking and memory.

Does lower token pricing reduce demand for computing? The answer is no. As tokens become cheaper, more use cases become economically viable, and those use cases can consume more tokens than the efficiencies save.
This leads to a second question: If every new generation is significantly more capable, what happens to earlier generations?
Durability: Continued Profitability from Installed Hardware
Not every workload requires the newest system. The right platform depends on the complexity and geometry of the workload, which allows earlier generations to remain economically valuable after newer systems arrive.
The NVIDIA A100 GPU shipped in 2020 and remains in commercial service six years later. CoreWeave recently extended reservations for units first introduced in 2020 until 2029. Major operators have also extended depreciation schedules for their servers, although those schedules are only a proxy for when hardware will stop generating revenue and be relocated or retired. Sprout Analysis tracks how data center GPU lifespans have changed across major operators.

Based on resale values, the estimated useful life is five to six years for an eight-GPU H100 system and nine to 10 years for the GB300 NVL72, according to Barkr. Silicon Data reports that a six-year-old A100 GPU remains worth approximately one-quarter of its original cost, compared with a value of zero under a five-year depreciation schedule more than a year earlier.
ORN found that the market pays 80% of the same rental price for A100 GPUs on a five-year contract as on a one-month contract.
CUDA runs across multiple GPU generations. When a new architecture arrives, operators can continue using existing hardware, while ongoing software and kernel optimization improves the capabilities of installed systems.
The same platform supports machine learning, deep learning, generative AI, inference, agentic AI and physical AI. Each new category of work can run on already installed hardware, increasing the number of opportunities to keep systems productive.
Fungibility: One Platform for AI and Other Workloads
A factory built for one type of work is betting that demand for that work will continue. A flexible factory that runs many workloads can remain useful and generate revenue as operations and market demand change.
NVIDIA AI Factory platforms run open and proprietary AI models across language, vision, biology, physics and robotics. They support the full AI lifecycle, from data processing and pre-training to post-training and inference. They can also be deployed in hyperscale data centers, AI clouds, sovereign programs, enterprise data centers and at the edge.

AI model development is only one use case. The same infrastructure can support data processing, scientific computing, simulation, graphics and more.
AI and non-AI workloads rely on parallel computation. NVIDIA GPUs are designed to perform that computation across thousands of cores. CUDA enables one chip to simulate light, fold proteins and predict the next token. More than 1,000 CUDA-X libraries cover applications ranging from deep neural networks and computational lithography to quantum circuit simulation, vector search and climate modeling. More than 10 million developers build on these libraries.
This makes NVIDIA GPUs general-purpose accelerated computing platforms rather than custom ASICs designed for a single workload. Tensor Cores and Transformer Engine provide AI-optimized hardware within a programmable architecture, combining specialization with flexibility. An architecture that supports many workloads can help maintain utilization and the revenue associated with it.
Examples of NVIDIA AI Factory Workloads
- Lily: Builds and runs protein, small-molecule and genomics models on an on-premises cluster of 1,016 GPUs, along with chatbots and agent workflows for internal teams.
- Pinterest: Post-trains and deploys vision-language models on a hyperscale cloud across 14,000 GPUs spanning NVIDIA Blackwell, Hopper and earlier architectures.
- Revolut: Processes billions of transaction records using NVIDIA cuDF to train and deploy models to the AI cloud.
- Runway: Trains world models on NVIDIA Hopper and serves them on the NVIDIA Blackwell platform using cloud infrastructure.
- Texas A&M University: Runs molecular simulations and AI drug discovery workloads on supercomputers, achieving 95% to 98% utilization across 26 projects and seven institutions.
The same flexibility extends beyond AI. Dassault Systèmes uses NVIDIA technology for virtual-twin simulations supporting aircraft certification in Wichita and vehicle design at Lucid Motors. Unilever uses digital twins to create product images, reducing production costs by half.
Why AI Factory Flexibility Matters
These examples are not exhaustive. They illustrate why a productive, durable and fungible platform can support more workloads, remain useful for longer and serve a broader market.
The NVIDIA AI Factory is designed to maximize profitable token production through three connected advantages: more tokens per megawatt, longer hardware service life and the flexibility to run a wide range of AI and non-AI workloads.
Hear from NVIDIA Founder and CEO Jensen Huang and learn more about the NVIDIA AI Factory at the GTC Berlin keynote on Wednesday, October 21, at 11 a.m. CET.
Source: blogs.nvidia.com


