Business Wire
BusinessGoogle spotlights XR storytelling projects at Venice, using Gemini and spatial film toolsThe Apex TimesBusinessAMD says Instinct AI systems are now operating in Saudi Arabia, highlighting a potential ramp tied to additional data-center powerThe Apex TimesBusinessCostco and Old Navy promotions, Apple leadership change, and other retail and tech themes surfaced in a market roundupThe Apex TimesBusinessDeere named among stocks making notable moves in late-Thursday trading recapThe Apex TimesBusinessSalesforce says AI-driven revenue momentum is building as Agentforce adoption spreadsThe Apex TimesBusinessSalesforce backs HiBob to bolster workforce AI, and adds a new AgentExchange email toolThe Apex TimesBusinessEverPass Media expands NFL distribution via multi-year Netflix deal for 2026 slateThe Apex TimesBusinessUnitedHealth shares rise as it moves to drop prior-authorization checks for about 30% of servicesThe Apex TimesBusinessBerkshire Hathaway CEO Greg Abel to Appear on TV in Rare Interview, With Focus Likely on Insurance and BNSFThe Apex TimesBusinessCoinbase expands Webull crypto trading footprint into CanadaThe Apex TimesBusinessBroadcom leans harder into VMware AI with a push aimed at enterprise rivalsThe Apex TimesBusinessModerna shares jump after GSK advances a rival mRNA flu vaccine to Phase IIIThe Apex TimesBusinessGoogle spotlights XR storytelling projects at Venice, using Gemini and spatial film toolsThe Apex TimesBusinessAMD says Instinct AI systems are now operating in Saudi Arabia, highlighting a potential ramp tied to additional data-center powerThe Apex TimesBusinessCostco and Old Navy promotions, Apple leadership change, and other retail and tech themes surfaced in a market roundupThe Apex TimesBusinessDeere named among stocks making notable moves in late-Thursday trading recapThe Apex TimesBusinessSalesforce says AI-driven revenue momentum is building as Agentforce adoption spreadsThe Apex TimesBusinessSalesforce backs HiBob to bolster workforce AI, and adds a new AgentExchange email toolThe Apex TimesBusinessEverPass Media expands NFL distribution via multi-year Netflix deal for 2026 slateThe Apex TimesBusinessUnitedHealth shares rise as it moves to drop prior-authorization checks for about 30% of servicesThe Apex TimesBusinessBerkshire Hathaway CEO Greg Abel to Appear on TV in Rare Interview, With Focus Likely on Insurance and BNSFThe Apex TimesBusinessCoinbase expands Webull crypto trading footprint into CanadaThe Apex TimesBusinessBroadcom leans harder into VMware AI with a push aimed at enterprise rivalsThe Apex TimesBusinessModerna shares jump after GSK advances a rival mRNA flu vaccine to Phase IIIThe Apex TimesBusinessGoogle spotlights XR storytelling projects at Venice, using Gemini and spatial film toolsThe Apex TimesBusinessAMD says Instinct AI systems are now operating in Saudi Arabia, highlighting a potential ramp tied to additional data-center powerThe Apex TimesBusinessCostco and Old Navy promotions, Apple leadership change, and other retail and tech themes surfaced in a market roundupThe Apex TimesBusinessDeere named among stocks making notable moves in late-Thursday trading recapThe Apex TimesBusinessSalesforce says AI-driven revenue momentum is building as Agentforce adoption spreadsThe Apex TimesBusinessSalesforce backs HiBob to bolster workforce AI, and adds a new AgentExchange email toolThe Apex TimesBusinessEverPass Media expands NFL distribution via multi-year Netflix deal for 2026 slateThe Apex TimesBusinessUnitedHealth shares rise as it moves to drop prior-authorization checks for about 30% of servicesThe Apex TimesBusinessBerkshire Hathaway CEO Greg Abel to Appear on TV in Rare Interview, With Focus Likely on Insurance and BNSFThe Apex TimesBusinessCoinbase expands Webull crypto trading footprint into CanadaThe Apex TimesBusinessBroadcom leans harder into VMware AI with a push aimed at enterprise rivalsThe Apex TimesBusinessModerna shares jump after GSK advances a rival mRNA flu vaccine to Phase IIIThe Apex TimesBusinessGoogle spotlights XR storytelling projects at Venice, using Gemini and spatial film toolsThe Apex TimesBusinessAMD says Instinct AI systems are now operating in Saudi Arabia, highlighting a potential ramp tied to additional data-center powerThe Apex TimesBusinessCostco and Old Navy promotions, Apple leadership change, and other retail and tech themes surfaced in a market roundupThe Apex TimesBusinessDeere named among stocks making notable moves in late-Thursday trading recapThe Apex TimesBusinessSalesforce says AI-driven revenue momentum is building as Agentforce adoption spreadsThe Apex TimesBusinessSalesforce backs HiBob to bolster workforce AI, and adds a new AgentExchange email toolThe Apex TimesBusinessEverPass Media expands NFL distribution via multi-year Netflix deal for 2026 slateThe Apex TimesBusinessUnitedHealth shares rise as it moves to drop prior-authorization checks for about 30% of servicesThe Apex TimesBusinessBerkshire Hathaway CEO Greg Abel to Appear on TV in Rare Interview, With Focus Likely on Insurance and BNSFThe Apex TimesBusinessCoinbase expands Webull crypto trading footprint into CanadaThe Apex TimesBusinessBroadcom leans harder into VMware AI with a push aimed at enterprise rivalsThe Apex TimesBusinessModerna shares jump after GSK advances a rival mRNA flu vaccine to Phase IIIThe Apex Times
Back to front
NVIDIA argues “performance per watt” should drive AI data center design as power becomes the limiting factor
The Apex Times

THE APEX TIMES

Business/The Apex Times/Jul 14, 11:25 AM EDT

NVIDIA argues “performance per watt” should drive AI data center design as power becomes the limiting factor

In a new technical blog, NVIDIA frames power efficiency as the metric that determines who can scale AI services profitably, pointing to Blackwell systems, rack-level software, and switch-and-network co-design as key levers.

Power has become the hard constraint behind AI infrastructure, NVIDIA says, and it is reshaping how data center operators plan their “AI factories.” In a technical blog published today, the company argues that organizations should measure progress not just by raw compute, but by performance per watt, the amount of AI work that can be delivered within a fixed power budget.

The core premise is straightforward, NVIDIA writes: the number of tokens an AI system can generate under a given power allocation influences revenue and profitability. As AI workloads increasingly shift toward agentic systems that can consume more tokens, the infrastructure choices made today, the company says, will decide which providers can scale when power is limited.

NVIDIA also claims that today’s frontier models are overwhelmingly based on mixture-of-experts (MoE), an architecture that can improve model quality while changing how inference is executed. Serving MoE efficiently at rack scale, it says, requires co-design across the system hardware and the software stack, plus operational maturity gained from running under real production load.

On that foundation, NVIDIA highlights its Blackwell NVL72 platform, describing it as delivering “the highest performance per watt” and lower token cost for inference. It then adds that the Vera Rubin platform is intended to build on this rack-scale base to further improve energy efficiency. NVIDIA’s blog emphasizes that each new wave of frontier models introduces architectural changes that unlock more intelligence while also demanding fresh optimizations to run efficiently at scale.

For comparative context, NVIDIA cites performance-per-watt gains between generations. It says that among the newest generation of leading open models, its GB300 NVL72 delivers up to 25x performance per watt versus the previous Hopper generation, while noting these numbers are a starting point because performance continues to improve over time.

To reflect that different deployments operate at different points, NVIDIA argues against relying on a single headline number. It says teams often need to choose between latency, throughput, and cost, and it showcases “Pareto curves” rather than one operating point. NVIDIA also points to tools such as DynoSim to help teams find a preferred point on the efficiency frontier before investing GPU-hours in validation.

The company attributes those efficiency gains to what it calls “extreme codesign.” In the networking portion of the rack, for example, NVIDIA says NVLink Switch is purpose-built for scale-up GPU domains rather than adapted from general-purpose networking. It also describes work on its sixth-generation NVLink Switch within the Vera Rubin platform, including functions it says are designed for AI workloads, such as SHARP (a capability described by NVIDIA as enabling in-network computing in the switch to offload work from GPUs).

NVIDIA then connects the rack-level performance story to its software stack for inference. It lists components including NVIDIA Dynamo and TensorRT LLM, as well as SGLang and vLLM, and says they support a broad set of optimizations. Examples it names include NVFP4 quantization (a reduced-precision technique meant to cut compute and memory costs), disaggregated serving, large-scale expert parallelism, KV-aware routing, and KV cache offloading. NVIDIA adds that software improvements can compound quickly, citing that performance per watt improved by up to 5x in a single month on DeepSeek V4 after optimization work.

Beyond raw compute efficiency, NVIDIA argues that the rest of the power chain matters. It states that in AI factories, only about 60% of electricity drawn from the grid turns into useful AI work because of losses in cooling and rack-level inefficiencies. It says DSX MaxLPS, the power-and-efficiency software in its NVIDIA DSX platform, helps close that gap by shifting power between GPUs and racks in real time, supporting warm-water liquid cooling and using techniques it calls power steering to extract more performance. NVIDIA claims this approach can let operators run up to 40% more GPUs within the same power budget, while also emphasizing that rack-scale reliability requires additional engineering because rack-level systems introduce failure modes not seen in single-node deployments.

In terms of real-world usage, NVIDIA says the Blackwell NVL72 platform is used by leading AI labs such as Anthropic and OpenAI for inference, and that a range of inference service providers and AI “natives” run open models in production on the platform. It provides a few deployment examples, including CoreWeave deploying Kimi K2.6 on NVIDIA GB300 NVL72 with NVFP4 quantization and EAGLE3 speculative decoding, Perplexity running Qwen variants on GB200 NVL72 for an agent platform serving millions of queries daily, and Fireworks AI deploying GLM 5.2 on the Blackwell platform to enable production deployments for customers such as Cursor and Factory AI. The company ends by suggesting that this accumulation of production experience underpins the Vera Rubin platform’s head start.

Why It Matters

  • As power and cooling become limiting factors for AI buildouts, efficiency metrics such as performance per watt can increasingly determine which operators scale at sustainable economics.
  • NVIDIA’s emphasis on rack-level co-design suggests that future model capability upgrades will need to be matched by infrastructure and software optimizations, not just faster chips.
  • If providers treat Pareto tradeoffs as a planning framework, procurement and deployment decisions may shift from single-number benchmarks toward workload-specific efficiency targets.
  • The blog’s focus on software-driven power management and token cost indicates that margins in inference businesses could hinge on optimization pipelines as much as on hardware selections.

Sources

Key Facts

  • NVIDIA says performance per watt, not just raw compute, should be the key metric for AI infrastructure because token output under a fixed power budget affects revenue and profitability.
  • The company argues that MoE (mixture-of-experts) models require hardware-software co-design to be served efficiently at rack scale, along with production operational depth.
  • NVIDIA positions its Blackwell NVL72 platform as delivering the highest performance per watt and lower token cost for inference, with Vera Rubin intended to further elevate rack-scale energy efficiency.
  • NVIDIA claims GB300 NVL72 can deliver up to 25x performance per watt versus Hopper-generation systems for some leading open models, while stressing this is a starting point that improves over time.
  • NVIDIA highlights tools and methods like Pareto curves and DynoSim to help teams select operating points across latency, throughput, and cost tradeoffs.
  • NVIDIA says its inference software stack includes components such as NVIDIA Dynamo, TensorRT LLM, SGLang, and vLLM, and supports optimizations like NVFP4 quantization and KV cache offloading.
  • NVIDIA claims DSX MaxLPS helps reduce AI-factory power losses and can enable up to 40% more GPUs within the same power budget, after stating that only about 60% of grid electricity becomes useful AI work due to cooling and rack inefficiencies.

Technology Related

NVIDIA argues “performance per watt” should drive AI data center design as power becomes the limiting factor | The Apex Times