THE APEX TIMES
NVIDIA and AWS pair Blackwell GPUs with OpenSearch Serverless to push AI into production
A new collaboration targets the two bottlenecks enterprises face when moving AI from experiments to real systems: low-latency inference and high-performance vector search for retrieval-augmented generation, semantic search and agentic AI.
NVIDIA and Amazon Web Services are rolling out a tighter stack for deploying AI systems at scale, focusing on the parts that tend to break first in production: inference latency, vector search speed, and infrastructure that can scale up or down without adding new operational work for engineering teams.
In an NVIDIA blog post published June 24, the companies say they are integrating NVIDIA’s AI infrastructure across AWS services including Amazon EC2 and Amazon OpenSearch. The goal, as described by NVIDIA, is to help enterprises stand up GPU-accelerated workloads for both compute-heavy inference and retrieval-based AI applications with fewer custom components.
On the compute side, NVIDIA introduced a new AWS instance line, called Amazon EC2 G7, powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs. NVIDIA positioned G7 as an instance type designed for production workloads across AI inference, graphics, video and GPU-accelerated data analytics. The company said G7 is available in one-, two-, four- and eight-GPU configurations and that bare metal options are “coming soon.”
NVIDIA also gave performance and capacity figures to distinguish G7 from the prior generation, G6. The company said G7 delivers up to 4.6x AI inference performance versus G6, up to 2.1x graphics performance, and significantly faster GPU-accelerated data analytics on Amazon EMR for Apache Spark using the NVIDIA cuDF library. It also described the platform’s right-sizing posture, saying customers can scale to the number of GPUs and memory their workloads actually need rather than over-provisioning a dedicated GPU platform.
For teams that rely on large-scale retrieval, NVIDIA said AWS’s next-generation Amazon OpenSearch Serverless is shifting to GPU-accelerated vector indexing using NVIDIA cuVS. cuVS is NVIDIA’s GPU-accelerated library for vector indexing, aimed at speeding up similarity search against embeddings, which are numeric representations of text, images or other data. NVIDIA said cuVS becomes the default compute choice for all vector collections in OpenSearch Serverless, and that this matters for retrieval-augmented generation (RAG), semantic search, recommendation systems and agentic AI applications.
NVIDIA quantified the potential retrieval improvement, saying vector indexing can be up to 10x faster at a quarter of the cost compared with CPU-only vector database builds. The company tied this to a practical deployment milestone, saying billion-scale vector databases could be built in under an hour, framing the update as a way to make production-grade retrieval pipelines more attainable for teams without specialized performance tuning.
Beyond runtime inference and search, the collaboration includes assurances for training workloads. NVIDIA said AWS has achieved NVIDIA Exemplar Cloud status for NVIDIA GB300 on training performance. NVIDIA described Exemplar Clouds as a performance evaluation and benchmark framework, where AWS meets thresholds used to compare cloud performance against NVIDIA’s reference architecture through co-engineering between the companies.
NVIDIA said Exemplar Clouds is meant to give developers more confidence that their training workloads will run on cloud infrastructure with consistent, high performance, while also helping teams improve total cost of ownership and move from planning to production more efficiently. For enterprises, that is often the harder part of the AI procurement cycle, where teams must validate training throughput and cost before they commit to a long-lived platform.
The release also detailed how customers can access the new building blocks through common AWS software paths. NVIDIA said G7 instances can be reached via AWS Deep Learning Amazon Machine Images (AMIs), Amazon Deep Learning Containers, Amazon EMR, Amazon EKS, Amazon ECS, and graphics AMIs, with support coming soon to Amazon SageMaker AI. On the storage and networking side, NVIDIA cited support for up to eight GPUs, total GPU memory up to 256GB, up to 700 Gbps of EFA-enabled networking (Enhanced Fabric Adapter), and up to 7.6TB of local NVMe SSD storage across single- to eight-GPU configurations.
It remains unclear from the blog post exactly when every component is broadly available, which regions it will roll out to first, and what specific application benchmarks NVIDIA expects customers to reproduce in their own environments. NVIDIA also did not provide concrete pricing for EC2 G7 or for OpenSearch Serverless vector collections; it only cited relative performance and cost comparisons and the operational promise of “serverless scaling” when workloads are idle.
For the next phase, the main thing to watch is whether customers can translate these infrastructure changes into measurable deployment outcomes, such as lower latency for inference-driven applications and faster time-to-production for RAG and vector search workloads. If the promised retrieval speedups and the Exemplar Cloud training assurances hold up in customer testing, the pairing of G7 compute plus cuVS-backed OpenSearch Serverless could become a more standard reference architecture for production AI systems built on AWS. However, teams will likely still need to validate performance, security controls, and cost under their own data distributions and workload patterns.
Why It Matters
- Enterprise AI deployments often stall at the point where inference latency and vector search performance become constraints, and this effort targets both with GPU acceleration.
- Making GPU-powered vector indexing the default in OpenSearch Serverless could reduce the amount of specialized engineering work required to run RAG and semantic search at scale.
- The quantified performance improvements for EC2 G7, if confirmed by customers, could change how teams size GPU capacity and manage operational overhead for production systems.
- Training workload validation through Exemplar Cloud status can shorten procurement and platform evaluation cycles, especially for teams that need predictable throughput and cost.
Key Facts
- NVIDIA and AWS say they are integrating NVIDIA AI infrastructure across Amazon EC2 and Amazon OpenSearch to support production-scale AI deployments.
- EC2 G7 instances, powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, target AI inference, graphics, video and GPU-accelerated data analytics workloads.
- NVIDIA claims EC2 G7 delivers up to 4.6x AI inference performance versus EC2 G6, up to 2.1x graphics performance, and faster GPU-accelerated analytics on Amazon EMR using cuDF for Apache Spark.
- NVIDIA said cuVS becomes the default GPU-accelerated vector indexing compute choice in Amazon OpenSearch Serverless, aimed at speeding up retrieval for RAG, semantic search and agentic AI.
- NVIDIA claims vector indexing can be up to 10x faster at a quarter of the cost versus CPU-only vector database builds, making billion-scale vector database builds practical in under an hour.
- AWS achieved NVIDIA Exemplar Cloud status on NVIDIA GB300 for training workloads, based on NVIDIA’s benchmark thresholds and co-engineering efforts.
Technology Related
Elon Musk’s chip preference spotlights Nvidia’s edge over AMD, but investors still watch execution
A Yahoo Finance analysis highlighted Nvidia’s faster growth relative to AMD, drawing attention to how high-profile tech users, including Elon Musk, frame the semiconductor race.
Ming-Chi Kuo says Nvidia has revived Rubin CPX after it seemingly vanished from the AI roadmap
The analyst Ming-Chi Kuo says Nvidia’s Rubin CPX accelerator is back, with what he characterizes as a substantial redesign after the chip appeared to be shelved earlier this year.
Apple’s next CEO arrives with a different kind of power: money, and an AI test
A new leadership chapter at Apple, as reported by Yahoo Finance, raises a central question for investors and customers alike: will Apple use its unusual financial profile to change its AI direction, or simply defend its status quo?
ZonPrep buys inbound-inventory software and services, betting on Amazon logistics automation
The Amazon-focused supply chain and FBA prep company says it acquired Wizard-Industries and FNSKU Studio, tools aimed at helping sellers get inventory into Amazon faster and with fewer process steps.
Nvidia pauses part of its AI customer financing after a strong quarter, raising questions about timing
After delivering another heavy AI-related quarter, Nvidia indicated it is stepping back from a portion of its financing approach for customers. Market coverage framed the move as potentially awkward, given investor expectations tied to continued momentum in AI infrastructure spending.
Apple CEO transition hands AI test to John Ternus as AAPL slips
John Ternus takes over as Apple’s chief executive role as Phil Schiller steps back, with market attention focused on how leadership changes could affect ongoing work on artificial intelligence initiatives. Apple shares slid in early trading following the transition reports.
Anthropic reportedly signs $35 billion cloud deal involving Nvidia-backed Lambda and a Texas data-center lease
A Yahoo Finance report says Anthropic has agreed to a long-term cloud-computing arrangement worth $35 billion, with the infrastructure and data-center lease tied to Lambda, an Nvidia-backed provider.
FTC and 22 states sue Amazon, alleging it overcharged advertisers using its retail platform
The U.S. Federal Trade Commission and a coalition of state attorneys general accused Amazon of misleading businesses about pricing tied to advertising on its shopping marketplace, alleging the conduct resulted in billions in gains for the company.
Intel’s push toward on-prem, privacy-focused AI gets a partnership spotlight as Xeon 6 platform work expands
A new extension to Kasm Technologies’ deal work with Intel highlights a market trend toward running large language model workloads locally on enterprise hardware, aiming to reduce data exposure and reliance on GPUs.
Broadcom (AVGO) set to report earnings Wednesday after the bell, with investors focused on guidance and demand outlines
The fabless chip and software maker Broadcom will release its next quarterly results this Wednesday after market close, according to a preview posted by Yahoo Finance.