THE APEX TIMES
NVIDIA’s new push for AI “factories” aims to bring compute access, financing and usage economics under one roof
NVIDIA says it is opening accelerated-computing capacity to AI cloud operators through a revenue-sharing and credit-support approach, with early deployments from Sharon AI and Firmus.
NVIDIA is trying to solve a bottleneck that has become central to the next phase of artificial intelligence: not just building models, but running them continuously at scale for real users. In a new strategy, the company is partnering with AI cloud companies to package large-scale, multi-tenant accelerated compute into “AI factories” that can be brought online quickly and kept highly utilized, aligning the economics of production inference with how much customers actually use.
The push reflects an industry shift from training, which is often done in batches, to inference at production scale, where systems generate outputs continuously and must be ready to serve demand that can spike. NVIDIA’s framing is that compute demand is moving toward token-scale services, meaning the business value is increasingly tied to how efficiently companies can generate tokens, the basic units of text or other generated outputs produced by AI models.
According to NVIDIA, many emerging AI companies have struggled to obtain capital-intensive infrastructure even when they can secure long-term commitments. The company says the gap is financing itself, not just interest from customers, arguing that traditional approaches can leave compute access slow or constrained until facilities, power and hardware bring-up are complete.
To address that, NVIDIA says it will “open up compute access” to a broader AI ecosystem that includes startups, model builders, enterprises, research organizations and regional AI players. The intent is to give AI clouds access to large-scale NVIDIA infrastructure while using a model that combines revenue sharing with credit support, so that capacity expansion can happen without requiring each operator to shoulder the full upfront investment burden.
Under the approach described by NVIDIA, AI cloud companies will sell cloud services delivered through NVIDIA DSX AI factories, which NVIDIA characterizes as systems that manufacture tokens at scale. NVIDIA says the structure is meant to accelerate adoption of NVIDIA platforms among cloud customers, while giving NVIDIA a new recurring revenue stream that is linked to usage.
NVIDIA provided early examples of the strategy taking shape. Sharon AI, a company focused on sovereign AI compute infrastructure, says it is deploying up to 40,000 NVIDIA Grace Blackwell GB300 GPUs as part of the collaboration, positioning the rollout as a way to deliver large-scale AI capacity under local control.
Firmus Technologies, another early partner, says it is building a DSX AI factory campus in Batam, Indonesia. Firmus expects the campus to scale to 360 megawatts and up to 170,000 NVIDIA GPUs, describing the project as an energy- and cost-efficient platform intended to let its AI cloud serve more customers as demand grows.
The company also points to what it calls a new commercial reality for AI-native businesses. It says organizations like Baseten, Fireworks AI and Together AI illustrate where compute demand is heading, as they work through model development steps including training, post-training and fine-tuning, alongside high-volume agentic inference, which refers to AI systems that can take actions as well as generate responses. Their customers, NVIDIA adds, need reliable access to accelerated compute as usage scales, but they also need flexibility as products move from pilots to production.
NVIDIA’s described model is also aimed at shortening timelines that typically slow down AI deployments. For inference providers, agent platforms and enterprises scaling AI, NVIDIA says the partnership can provide full-stack accelerated computing faster than waiting through site selection, power procurement, construction and hardware bring-up.
What NVIDIA does not provide in the announcement is the detailed financial structure of the revenue-sharing and credit-support terms, including whether specific customers will have minimum purchase commitments or how long credit backing lasts. The company also does not disclose pricing, capacity reservation mechanisms, or the timeline for broader availability of the “cloud partner” program beyond the stated early participants. For operators in this space, those specifics may determine whether the arrangement meaningfully reduces capital risk or simply reshapes it.
Why It Matters
- If successful, NVIDIA’s structure could make it easier for AI providers to scale inference capacity without waiting for full data center buildouts and financing cycles.
- By tying revenue to usage, the approach aims to match how AI products monetize, potentially reducing mismatch between compute costs and customer demand.
- The strategy could broaden the set of companies that can access large-scale NVIDIA infrastructure, increasing competition among AI cloud operators and inference offerings.
- It may also create a more predictable, recurring revenue mix for NVIDIA as inference volumes grow across many customer workloads.
Sources
Key Facts
- NVIDIA says AI compute demand is shifting from model development toward production inference that runs continuously and generates tokens at scale.
- NVIDIA is introducing a strategy to open compute access to AI cloud operators, startups, enterprises, research organizations and regional AI players.
- The company says AI clouds will sell services delivered through NVIDIA DSX AI factories that manufacture tokens at scale.
- NVIDIA describes the approach as aligning economics through a revenue-sharing and credit-support model.
- Sharon AI says it is deploying up to 40,000 NVIDIA Grace Blackwell GB300 GPUs as part of the collaboration.
- Firmus Technologies says it is building a DSX AI factory campus in Batam, Indonesia, targeting 360 megawatts and up to 170,000 NVIDIA GPUs.
Technology Related
Elon Musk’s chip preference spotlights Nvidia’s edge over AMD, but investors still watch execution
A Yahoo Finance analysis highlighted Nvidia’s faster growth relative to AMD, drawing attention to how high-profile tech users, including Elon Musk, frame the semiconductor race.
Ming-Chi Kuo says Nvidia has revived Rubin CPX after it seemingly vanished from the AI roadmap
The analyst Ming-Chi Kuo says Nvidia’s Rubin CPX accelerator is back, with what he characterizes as a substantial redesign after the chip appeared to be shelved earlier this year.
Apple’s next CEO arrives with a different kind of power: money, and an AI test
A new leadership chapter at Apple, as reported by Yahoo Finance, raises a central question for investors and customers alike: will Apple use its unusual financial profile to change its AI direction, or simply defend its status quo?
ZonPrep buys inbound-inventory software and services, betting on Amazon logistics automation
The Amazon-focused supply chain and FBA prep company says it acquired Wizard-Industries and FNSKU Studio, tools aimed at helping sellers get inventory into Amazon faster and with fewer process steps.
Nvidia pauses part of its AI customer financing after a strong quarter, raising questions about timing
After delivering another heavy AI-related quarter, Nvidia indicated it is stepping back from a portion of its financing approach for customers. Market coverage framed the move as potentially awkward, given investor expectations tied to continued momentum in AI infrastructure spending.
Apple CEO transition hands AI test to John Ternus as AAPL slips
John Ternus takes over as Apple’s chief executive role as Phil Schiller steps back, with market attention focused on how leadership changes could affect ongoing work on artificial intelligence initiatives. Apple shares slid in early trading following the transition reports.
Anthropic reportedly signs $35 billion cloud deal involving Nvidia-backed Lambda and a Texas data-center lease
A Yahoo Finance report says Anthropic has agreed to a long-term cloud-computing arrangement worth $35 billion, with the infrastructure and data-center lease tied to Lambda, an Nvidia-backed provider.
FTC and 22 states sue Amazon, alleging it overcharged advertisers using its retail platform
The U.S. Federal Trade Commission and a coalition of state attorneys general accused Amazon of misleading businesses about pricing tied to advertising on its shopping marketplace, alleging the conduct resulted in billions in gains for the company.
Intel’s push toward on-prem, privacy-focused AI gets a partnership spotlight as Xeon 6 platform work expands
A new extension to Kasm Technologies’ deal work with Intel highlights a market trend toward running large language model workloads locally on enterprise hardware, aiming to reduce data exposure and reliance on GPUs.
Broadcom (AVGO) set to report earnings Wednesday after the bell, with investors focused on guidance and demand outlines
The fabless chip and software maker Broadcom will release its next quarterly results this Wednesday after market close, according to a preview posted by Yahoo Finance.