THE APEX TIMES
NVIDIA argues “performance per watt” should drive AI data center design as power becomes the limiting factor
In a new technical blog, NVIDIA frames power efficiency as the metric that determines who can scale AI services profitably, pointing to Blackwell systems, rack-level software, and switch-and-network co-design as key levers.
Power has become the hard constraint behind AI infrastructure, NVIDIA says, and it is reshaping how data center operators plan their “AI factories.” In a technical blog published today, the company argues that organizations should measure progress not just by raw compute, but by performance per watt, the amount of AI work that can be delivered within a fixed power budget.
The core premise is straightforward, NVIDIA writes: the number of tokens an AI system can generate under a given power allocation influences revenue and profitability. As AI workloads increasingly shift toward agentic systems that can consume more tokens, the infrastructure choices made today, the company says, will decide which providers can scale when power is limited.
NVIDIA also claims that today’s frontier models are overwhelmingly based on mixture-of-experts (MoE), an architecture that can improve model quality while changing how inference is executed. Serving MoE efficiently at rack scale, it says, requires co-design across the system hardware and the software stack, plus operational maturity gained from running under real production load.
On that foundation, NVIDIA highlights its Blackwell NVL72 platform, describing it as delivering “the highest performance per watt” and lower token cost for inference. It then adds that the Vera Rubin platform is intended to build on this rack-scale base to further improve energy efficiency. NVIDIA’s blog emphasizes that each new wave of frontier models introduces architectural changes that unlock more intelligence while also demanding fresh optimizations to run efficiently at scale.
For comparative context, NVIDIA cites performance-per-watt gains between generations. It says that among the newest generation of leading open models, its GB300 NVL72 delivers up to 25x performance per watt versus the previous Hopper generation, while noting these numbers are a starting point because performance continues to improve over time.
To reflect that different deployments operate at different points, NVIDIA argues against relying on a single headline number. It says teams often need to choose between latency, throughput, and cost, and it showcases “Pareto curves” rather than one operating point. NVIDIA also points to tools such as DynoSim to help teams find a preferred point on the efficiency frontier before investing GPU-hours in validation.
The company attributes those efficiency gains to what it calls “extreme codesign.” In the networking portion of the rack, for example, NVIDIA says NVLink Switch is purpose-built for scale-up GPU domains rather than adapted from general-purpose networking. It also describes work on its sixth-generation NVLink Switch within the Vera Rubin platform, including functions it says are designed for AI workloads, such as SHARP (a capability described by NVIDIA as enabling in-network computing in the switch to offload work from GPUs).
NVIDIA then connects the rack-level performance story to its software stack for inference. It lists components including NVIDIA Dynamo and TensorRT LLM, as well as SGLang and vLLM, and says they support a broad set of optimizations. Examples it names include NVFP4 quantization (a reduced-precision technique meant to cut compute and memory costs), disaggregated serving, large-scale expert parallelism, KV-aware routing, and KV cache offloading. NVIDIA adds that software improvements can compound quickly, citing that performance per watt improved by up to 5x in a single month on DeepSeek V4 after optimization work.
Beyond raw compute efficiency, NVIDIA argues that the rest of the power chain matters. It states that in AI factories, only about 60% of electricity drawn from the grid turns into useful AI work because of losses in cooling and rack-level inefficiencies. It says DSX MaxLPS, the power-and-efficiency software in its NVIDIA DSX platform, helps close that gap by shifting power between GPUs and racks in real time, supporting warm-water liquid cooling and using techniques it calls power steering to extract more performance. NVIDIA claims this approach can let operators run up to 40% more GPUs within the same power budget, while also emphasizing that rack-scale reliability requires additional engineering because rack-level systems introduce failure modes not seen in single-node deployments.
In terms of real-world usage, NVIDIA says the Blackwell NVL72 platform is used by leading AI labs such as Anthropic and OpenAI for inference, and that a range of inference service providers and AI “natives” run open models in production on the platform. It provides a few deployment examples, including CoreWeave deploying Kimi K2.6 on NVIDIA GB300 NVL72 with NVFP4 quantization and EAGLE3 speculative decoding, Perplexity running Qwen variants on GB200 NVL72 for an agent platform serving millions of queries daily, and Fireworks AI deploying GLM 5.2 on the Blackwell platform to enable production deployments for customers such as Cursor and Factory AI. The company ends by suggesting that this accumulation of production experience underpins the Vera Rubin platform’s head start.
Why It Matters
- As power and cooling become limiting factors for AI buildouts, efficiency metrics such as performance per watt can increasingly determine which operators scale at sustainable economics.
- NVIDIA’s emphasis on rack-level co-design suggests that future model capability upgrades will need to be matched by infrastructure and software optimizations, not just faster chips.
- If providers treat Pareto tradeoffs as a planning framework, procurement and deployment decisions may shift from single-number benchmarks toward workload-specific efficiency targets.
- The blog’s focus on software-driven power management and token cost indicates that margins in inference businesses could hinge on optimization pipelines as much as on hardware selections.
Sources
Key Facts
- NVIDIA says performance per watt, not just raw compute, should be the key metric for AI infrastructure because token output under a fixed power budget affects revenue and profitability.
- The company argues that MoE (mixture-of-experts) models require hardware-software co-design to be served efficiently at rack scale, along with production operational depth.
- NVIDIA positions its Blackwell NVL72 platform as delivering the highest performance per watt and lower token cost for inference, with Vera Rubin intended to further elevate rack-scale energy efficiency.
- NVIDIA claims GB300 NVL72 can deliver up to 25x performance per watt versus Hopper-generation systems for some leading open models, while stressing this is a starting point that improves over time.
- NVIDIA highlights tools and methods like Pareto curves and DynoSim to help teams select operating points across latency, throughput, and cost tradeoffs.
- NVIDIA says its inference software stack includes components such as NVIDIA Dynamo, TensorRT LLM, SGLang, and vLLM, and supports optimizations like NVFP4 quantization and KV cache offloading.
- NVIDIA claims DSX MaxLPS helps reduce AI-factory power losses and can enable up to 40% more GPUs within the same power budget, after stating that only about 60% of grid electricity becomes useful AI work due to cooling and rack inefficiencies.
Technology Related
Google spotlights XR storytelling projects at Venice, using Gemini and spatial film tools
Google’s 100 ZEROS program is backing three extended-reality projects premiering at the 83rd Venice International Film Festival, all built to run on Android XR and to combine spatial experiences with Gemini-powered conversational interactions.
AMD says Instinct AI systems are now operating in Saudi Arabia, highlighting a potential ramp tied to additional data-center power
A recent market report frames AMD’s Instinct deployments in Saudi Arabia as a move from plan to production, and points to how incremental data-center capacity, measured in megawatts, could influence investor expectations.
Salesforce says AI-driven revenue momentum is building as Agentforce adoption spreads
In a recent market update circulated by Yahoo Finance, Salesforce management pointed to expanding use of its AI offerings, including agentic workflows and consumption-style pricing, as the company positions its next growth phase.
Salesforce backs HiBob to bolster workforce AI, and adds a new AgentExchange email tool
Salesforce said it is supporting HR-analytics and talent-workforce platform HiBob as part of efforts to connect enterprise data with “powered AI.” The company also announced an AgentExchange email tool aimed at expanding what business agents can do inside everyday workflows.
EverPass Media expands NFL distribution via multi-year Netflix deal for 2026 slate
EverPass Media says it has added Netflix’s five NFL games for the 2026 season to its NFL distribution offering, including the first-ever Thanksgiving Eve game, plus “NFL Honors.”
Broadcom leans harder into VMware AI with a push aimed at enterprise rivals
Broadcom’s VMware AI push is tied to the latest VCF 9.1 release, as the company’s messaging positions it against Nutanix and Microsoft in hybrid cloud and enterprise AI rollouts.
Yahoo Finance points to “buy zones” for Microsoft, Palantir, Shopify and ServiceNow
A market-readout from Yahoo Finance flagged several software and AI-linked names, including Palantir (PLTR), as trading in or near so-called buy zones. The note is framed as technical or timing-oriented, with limited company-specific detail.
Oracle Shares Fall as Investors Focus on Cash Flow Gap and Rising Borrowing Costs
A reported $23.7 billion cash shortfall over Oracle’s last fiscal year and $43 billion in borrowing are drawing attention to the company’s interest-rate exposure, a factor that can quickly change sentiment when Treasury yields are elevated.
Adobe’s next report faces a split view: Citi still expects a beat, but flags lingering risks
After Adobe lowered its annual revenue outlook, one analyst said the company can still deliver a beat-and-raise in fiscal third-quarter results, even as concerns remain.
Palantir’s commercial growth may overtake government revenue sooner than expected, according to a new market model
A widely watched growth-math forecast argues Palantir’s commercial revenue could surpass its government revenue before 2027, driven by a widening gap in the companies’ growth rates.