THE APEX TIMES
NVIDIA says Blackwell GPUs are powering OpenAI’s faster “Astra Ultrafast” GPT-6 mode
NVIDIA claims its Blackwell-based hardware and software stack helped OpenAI deliver up to 8x faster token generation for GPT-6 Astra Ultrafast, aimed at improving latency in coding and tool-using agent workflows.
NVIDIA on Oct. 1 detailed how it says its Blackwell GPUs are being used to accelerate OpenAI’s GPT-6 Astra Ultrafast, a faster-response mode now available through the OpenAI API and to eligible ChatGPT Work and Codex users. The company’s account focuses less on model training and more on inference, the compute that turns a model’s learned parameters into responses in production. NVIDIA says its platform and tooling helped OpenAI deliver a performance jump that is particularly important for applications where systems must repeatedly generate tokens, call external tools, and decide what to do next.
According to NVIDIA’s blog, Astra Ultrafast runs on NVIDIA Blackwell GPUs and is available now. NVIDIA said the Ultrafast mode can generate tokens up to 8x faster than Astra Standard. Tokens are the basic chunks of text (and code) that large language models produce step by step, so faster token generation translates into shorter wait times for users and for software agents that need to iterate quickly.
NVIDIA linked the speed improvement to a practical developer metric: reducing the time spent in edit-test-debug loops for coding agents. The company described workflows in which an agent writes code, uses a tool, checks the result, and then chooses a next action. In that kind of cycle, the delay between tool calls and the time to produce intermediate outputs can compound, making an interactive application feel slow even if the tools themselves respond quickly.
The blog also highlighted latency and throughput, the two core characteristics that affect how quickly and how consistently a deployed AI system can respond to requests. NVIDIA said OpenAI’s models tap into capabilities of the Blackwell architecture to generate “high-performance kernels,” software units that implement parts of the inference pipeline efficiently on the GPU. NVIDIA framed the result as delivering the acceleration needed to improve response times while balancing cost and operational efficiency.
A quote from Philippe Tillet, inference lead at OpenAI, attributed to the work NVIDIA helped enable: he said NVIDIA’s deep investment in tooling and documentation helped OpenAI make its models exceptionally good at programming Blackwell and Rubin GPUs. Tillet added that Astra can turn that knowledge into high-performance kernels that make NVIDIA hardware compelling across the “frontier of latency, throughput and cost,” and that Astra Ultrafast brings faster model responses for agent-based coding and complex tasks.
NVIDIA said the performance work continues after deployment as well. It described OpenAI using its own models to help refine inference software running on NVIDIA GPUs, taking advantage of the platform’s programmability to test and implement improvements over time. In this view, the overall system optimization is not treated as a one-time engineering project, but as an iterative cycle where inference libraries can be tuned as workloads and model behavior evolve.
A second quote came from Uday Ruddarraju, chief technology officer of compute at OpenAI. NVIDIA said Ruddarraju credited internal-model optimization of inference on NVIDIA GPUs and said NVIDIA’s programmability helped deliver the acceleration behind Astra Ultrafast. The message is that the speed gain is not only about having the right GPU hardware, but also about having software infrastructure that lets teams shape how inference executes at runtime.
Beyond the immediate Astra Ultrafast announcement, NVIDIA spent time explaining a broader architectural point: a programmable NVIDIA platform allows teams to reuse infrastructure across training, inference, and reinforcement learning as models change. NVIDIA said that flexibility can help improve utilization and avoid overprovisioning for each workload, an operational concern for data center operators that need to balance capacity and demand.
NVIDIA did not provide additional benchmarks beyond the claimed “up to 8x” token generation improvement, and it did not break out how much of the gain comes from model changes versus inference-engine optimizations. It also did not publish pricing for the mode in the blog post itself, pointing readers to an Ultrafast guide for access and pricing details. As with many performance claims in AI, the practical experience can depend on request patterns, batch sizes, deployment configuration, and how often applications trigger tool calls that shape an agent’s step-by-step flow.
For developers and operators, the main question to watch next is whether the Astra Ultrafast speed gains hold consistently across different application types and at scale. If the mode meaningfully reduces end-to-end latency in agent workflows, it could make tool-using coding assistants feel more interactive, and it may shift engineering priorities toward tighter inference loops. NVIDIA and OpenAI will likely be judged on how quickly they can translate this kind of acceleration into sustained capacity, reliability, and measurable user outcomes once the feature expands beyond initial eligibility groups.
Why It Matters
- Faster token generation can directly reduce perceived wait times in interactive AI tools and agent-driven coding workflows where delays compound across iterations.
- If inference optimization continues post-deployment, performance improvements may become more incremental and ongoing rather than tied solely to new model releases.
- Claims of gains “across latency, throughput and cost” underscore the data center challenge of meeting responsiveness requirements without sacrificing efficiency.
- The emphasis on tool-using agent loops highlights a shift from standalone chat speed to end-to-end workflow latency as a key benchmark.
Sources
Key Facts
- NVIDIA says GPT-6 Astra Ultrafast is running on NVIDIA Blackwell GPUs and is available now through the OpenAI API and for eligible ChatGPT Work and Codex users.
- NVIDIA claims Astra Ultrafast delivers up to 8x faster token generation than Astra Standard.
- The company links faster token generation to improved responsiveness in agent workflows that repeatedly generate outputs, call tools, and decide next steps.
- NVIDIA attributes the acceleration to inference optimizations, including software “kernels” that it says are derived from OpenAI’s ability to program NVIDIA Blackwell hardware.
- NVIDIA says OpenAI is also using internal models to refine inference software on NVIDIA GPUs over time.
- NVIDIA points to its programmable platform as a way to reuse infrastructure across training, inference, and reinforcement learning as workloads change.
Technology Related
Oracle Shares Drop After Reports of OpenAI Revenue Gap, Sparking Broader AI Stock Worry
Market coverage tying the move to a reported $20 billion gap in OpenAI revenue highlights how quickly expectations for AI spending can ripple across software and chip-related names. Oracle, NYSE:ORCL, is the latest example of a stock reacting to sentiment around the AI buildout.
Anaconda and Intel expand push to help enterprises move AI projects toward production
The software tools company says its expanded collaboration with Intel is aimed at making Intel-optimized platforms easier for developers to use as they deploy AI in real-world environments.
Google Maps’ “Fan-Favorite Dining List” turns a year of restaurant data into a city-by-city food trend guide
Alphabet’s Google is using Google Maps engagement outlines, including reviews, ratings, and direction requests, to surface what diners are eating across 10 cities and to highlight budget-friendly favorites.
Stuut’s $52.5M Series B highlights a push toward AI agents, with Microsoft tied to the “revenue layer” thesis
A new funding round for Stuut is being framed by investors and industry observers as evidence that the next phase of enterprise AI is shifting from co-pilots that assist users to AI agents that help complete tasks, billed and managed as software products. The round’s timing also feeds a broader narrative around Microsoft’s strategy in the agent economy.
NVIDIA argues AI factory returns hinge on power efficiency, software longevity, and workload “fungibility”
In a new post, NVIDIA says the economics of mega-watt-scale AI data centers depend less on headline performance and more on how many useful tokens a site can produce per unit of power over many years, across many types of workloads.
Reported Google AI Pact with SpaceX could be worth as much as $29B, but Alphabet investors may face execution risk
A widely discussed agreement tied to SpaceX’s expanding AI push highlights Alphabet’s growing role in frontier computing, while also underscoring how hard it can be to convert large contract headlines into durable, protected revenue.
Paramount-Warner’s New Combination Rises as Netflix’s Biggest Streaming Rival, With About $70 Billion in Sales and $82 Billion of Debt, Report Says
A newly combined media company, formed from Paramount and Warner Bros. assets, is described as outpacing Netflix in annual sales while also carrying a large debt load. Analysts and investors will likely weigh content-scale advantages against balance-sheet risk.
Microsoft’s Nvidia Partnership Expansion Puts AI Momentum in Focus, While Valuation Raises the Bar
A market note highlights how a deeper Nvidia collaboration could support Microsoft’s artificial intelligence growth, but the question for investors is whether the stock’s premium pricing leaves enough upside.
Oracle weighs logistics shift for Project Jupiter: trucking compressed natural gas in New Mexico, report says
A Yahoo Finance market chatter item says Oracle is considering shipping compressed natural gas to its Project Jupiter data center site in New Mexico, highlighting how power supply and fuel delivery are becoming strategic planning issues for large data centers.
Google expands Google Maps dining discovery with Ask Maps, trending lists, and “know before you go” tips
Alphabet’s Google is rolling out three Google Maps features aimed at helping users find restaurants, decide what to order, and plan trips with community-supplied guidance.