THE APEX TIMES
NVIDIA argues AI factory returns hinge on power efficiency, software longevity, and workload “fungibility”
In a new post, NVIDIA says the economics of mega-watt-scale AI data centers depend less on headline performance and more on how many useful tokens a site can produce per unit of power over many years, across many types of workloads.
NVIDIA is framing the next phase of the AI buildout around economics, not benchmarks. In a company blog post published Oct. 1, the chipmaker argues that AI factories are capital-intensive enterprises, built at a scale measured in megawatts and priced accordingly, and that returns come from a specific mix of performance, durability and flexibility.
The post puts a point estimate on the magnitude of the bet. NVIDIA says each megawatt AI factory costs roughly $60 million, so operators will commit capital only if they can see a credible return on investment. NVIDIA then outlines three factors it says jointly shape those returns: the maximum earning capacity of the hardware, whether demand keeps the systems running at high utilization rather than idling after a short cycle, and whether the factory can be used for many different workloads instead of one narrow use case.
Under NVIDIA’s logic, “strength” in one area cannot fully compensate for weakness elsewhere. A factory that can produce a lot, but sells only part of that output, will not generate strong returns. Likewise, even if demand is high at first, the economics deteriorate if the system stops being competitive at full capacity after a relatively short period.
To engineer for those tradeoffs, NVIDIA says its AI factory approach is grounded in codesign across the full stack, from model and workloads down to compute, networking and memory, plus continuous software optimization after deployment. The company points to CUDA-X libraries, describing them as enabling a factory to run accelerated workloads broadly, and it ties its standardized architecture to a “validated reference design” that can be deployed by different operators.
Power, NVIDIA says, is the binding constraint on an AI factory. That turns attention to what it calls tokens per second per megawatt, a metric that links output directly to the energy envelope that limits scale. In the same framework, NVIDIA argues that higher throughput per megawatt increases revenue potential within a fixed power footprint, and that lower cost per token supports margins on each token produced.
NVIDIA’s post includes comparative performance and cost claims attributed to SemiAnalysis AgentX data, comparing systems it describes as NVIDIA Vera Rubin NVL72 and NVIDIA GB300 NVL72. NVIDIA says the Vera Rubin configuration delivers over 30 times higher throughput per megawatt than the GB300 NVL72, and it says it can cut cost per million tokens by up to 45 times on DeepSeek V4 Pro. The post attributes those gains to “extreme” end-to-end codesign across the model and workload levels, software, and compute and system components.
The post then turns to two policy questions operators face: if each generation makes tokens dramatically cheaper, will demand for compute shrink, and what happens to older systems when new ones arrive? NVIDIA answers both by arguing that cheaper tokens expand the addressable set of use cases, which in turn increases total token demand rather than reducing it. On durability, it says not every workload requires the newest platform, so older generations can continue earning by matching the “complexity and shape” of the work to the appropriate system.
To illustrate durability, NVIDIA points to the NVIDIA A100 GPU, which it says shipped in 2020 and is still in commercial service six years later. It also cites an example of CoreWeave extending bookings for units introduced in 2020 through 2029, and it references third-party analysis of how operators extend depreciation schedules over time. NVIDIA also includes market estimates of useful life and resale or leasing value, stating that Barkr assesses useful life at five to six years for an eight-GPU H100 system and nine to 10 years for GB300 NVL72, while additional data providers it names suggest older A100 value can remain significant and that rental pricing depends on contract length.
NVIDIA ties its durability thesis to software portability. The post says CUDA, NVIDIA’s GPU programming platform, runs across generations, so existing purchased hardware is not stranded when architectures change. It also argues that continuous kernel and software optimization can improve what existing hardware does, and it emphasizes “fungibility” as the longer-term utilization engine: the more kinds of AI and non-AI workloads a system can run, the longer it stays valuable as demand shifts. The company says its platform supports everything from language, vision and reasoning models to agentic AI and physical AI, spanning data processing, pretraining, post-training and inference, and that the same infrastructure also supports non-AI uses such as simulation, graphics and scientific computing.
What remains less specific in NVIDIA’s post is the measured impact of these ideas across different operator business models, such as how quickly each operator’s mix of workloads adapts, or how utilization evolves when new generations enter the market. NVIDIA does not provide new, company-specific financial guidance tied to the framework, and its performance comparisons cite third-party analysis rather than presenting audited results under a single test protocol. Still, for operators planning multi-year deployments, the central message is clear: the industry’s ROI problem is inseparable from power-constrained throughput, software-managed longevity, and the ability to sell multiple kinds of computation from the same installed base.
Why It Matters
- If power-constrained throughput and token economics dominate ROI, future buying decisions may lean harder on system-level efficiency than on peak single-model performance.
- A durable installed base could reduce the perceived risk of multi-year data center investments, particularly if operators can keep selling older hardware into workloads that fit their capabilities.
- “Fungibility,” or the ability to run many workload types, may become a differentiator for AI factories as demand changes across model families, modalities and deployment environments.
- The framework may influence how operators design contracts, depreciation schedules, and hardware refresh cycles, since utilization longevity becomes a key driver of returns.
Sources
Key Facts
- NVIDIA says each megawatt AI factory costs roughly $60 million, pushing operators to require a clear return on investment.
- The company argues AI factory returns depend on earning capacity, the ability to sustain demand without early capacity loss, and the capacity to run multiple workloads rather than one.
- NVIDIA says power is the binding constraint and frames tokens per second per megawatt as the metric that governs earning capacity in a fixed energy envelope.
- NVIDIA cites SemiAnalysis AgentX data claiming its Vera Rubin NVL72 delivers over 30x higher throughput per megawatt than its GB300 NVL72, and up to 45x lower cost per million tokens on DeepSeek V4 Pro.
- The post argues durability improves because not every workload needs the newest system, and because CUDA and CUDA-X support cross-generation software reuse.
- NVIDIA points to the A100 GPU, shipped in 2020, as still in commercial service years later, and it cites CoreWeave extending bookings for units introduced in 2020 through 2029.
Technology Related
Oracle Shares Drop After Reports of OpenAI Revenue Gap, Sparking Broader AI Stock Worry
Market coverage tying the move to a reported $20 billion gap in OpenAI revenue highlights how quickly expectations for AI spending can ripple across software and chip-related names. Oracle, NYSE:ORCL, is the latest example of a stock reacting to sentiment around the AI buildout.
Anaconda and Intel expand push to help enterprises move AI projects toward production
The software tools company says its expanded collaboration with Intel is aimed at making Intel-optimized platforms easier for developers to use as they deploy AI in real-world environments.
Google Maps’ “Fan-Favorite Dining List” turns a year of restaurant data into a city-by-city food trend guide
Alphabet’s Google is using Google Maps engagement outlines, including reviews, ratings, and direction requests, to surface what diners are eating across 10 cities and to highlight budget-friendly favorites.
Stuut’s $52.5M Series B highlights a push toward AI agents, with Microsoft tied to the “revenue layer” thesis
A new funding round for Stuut is being framed by investors and industry observers as evidence that the next phase of enterprise AI is shifting from co-pilots that assist users to AI agents that help complete tasks, billed and managed as software products. The round’s timing also feeds a broader narrative around Microsoft’s strategy in the agent economy.
Reported Google AI Pact with SpaceX could be worth as much as $29B, but Alphabet investors may face execution risk
A widely discussed agreement tied to SpaceX’s expanding AI push highlights Alphabet’s growing role in frontier computing, while also underscoring how hard it can be to convert large contract headlines into durable, protected revenue.
Paramount-Warner’s New Combination Rises as Netflix’s Biggest Streaming Rival, With About $70 Billion in Sales and $82 Billion of Debt, Report Says
A newly combined media company, formed from Paramount and Warner Bros. assets, is described as outpacing Netflix in annual sales while also carrying a large debt load. Analysts and investors will likely weigh content-scale advantages against balance-sheet risk.
Microsoft’s Nvidia Partnership Expansion Puts AI Momentum in Focus, While Valuation Raises the Bar
A market note highlights how a deeper Nvidia collaboration could support Microsoft’s artificial intelligence growth, but the question for investors is whether the stock’s premium pricing leaves enough upside.
Oracle weighs logistics shift for Project Jupiter: trucking compressed natural gas in New Mexico, report says
A Yahoo Finance market chatter item says Oracle is considering shipping compressed natural gas to its Project Jupiter data center site in New Mexico, highlighting how power supply and fuel delivery are becoming strategic planning issues for large data centers.
Google expands Google Maps dining discovery with Ask Maps, trending lists, and “know before you go” tips
Alphabet’s Google is rolling out three Google Maps features aimed at helping users find restaurants, decide what to order, and plan trips with community-supplied guidance.
NVIDIA rolls out a slate of new titles for GeForce NOW in October, including Control Resonant rewards and The Witcher 3 remastered
GeForce NOW members get 25 new games across October, with six titles available immediately. New this season: streaming of Control Resonant for Performance and Ultimate subscribers and The Witcher 3: Wild Hunt – Remastered joining the cloud library.