NVIDIA and AWS expand AI infrastructure partnership with EC2 G7 Blackwell instances, GPU-accelerated vector search, and Exemplar Cloud certification
NVIDIA and AWS announced a set of coordinated infrastructure expansions this week, including new Amazon EC2 G7 instances powered by NVIDIA's RTX PRO 4500 Blackwell Server Edition GPUs, GPU-accelerated vector indexing on by default in Amazon OpenSearch Serverless, and AWS's achievement of NVIDIA Exemplar Cloud Status for GB300 training workloads.
What's new
Amazon EC2 G7 instances bring Blackwell-architecture GPUs to AWS general availability compute:
- Up to 4.6× faster AI inference versus G6 instances
- Up to eight GPUs per instance with 256GB total GPU memory
- 700 Gbps EFA-enabled networking
- Up to 7.6TB local NVMe SSD storage
NVIDIA cuVS in Amazon OpenSearch Serverless defaults GPU-accelerated vector indexing for all new vector collections:
- Up to 10× faster vector indexing
- Cost approximately one-quarter of CPU-only approaches
NVIDIA Exemplar Cloud Status certifies that AWS meets NVIDIA's benchmarks for GB300 training workloads, ensuring consistent high-performance infrastructure for large-scale training runs.
As NVIDIA framed the partnership's aim: "Building AI systems at scale is demanding, requiring low-latency inference, fast vector search, strong GPU price-performance and infrastructure that can grow without multiplying operational complexity."
Context
The G7 instances mark the first Blackwell-architecture GPU entering the AWS general availability EC2 catalog. Previous Blackwell availability on AWS was limited to higher-end training configurations. The Exemplar Cloud certification for GB300 training separately validates AWS as a tier-1 platform for the largest NVIDIA training workloads.
NVIDIA's cuVS vector library was previously an optional configuration in AWS managed services. Making it the default for all new OpenSearch Serverless vector collections removes a manual setup step that previously deterred teams from enabling GPU-accelerated retrieval in production.
This follows NVIDIA's NemoClaw and Agent Toolkit announcements earlier this week, signaling a broader push to move GPU acceleration from training-only into the full AI stack: inference, retrieval, and agent orchestration.
Why it matters
For enterprise teams building retrieval-augmented generation pipelines, these changes reduce two distinct bottlenecks simultaneously. G7 instances lower the cost-per-inference for GPU-class compute at AWS, while GPU-accelerated vector indexing by default in OpenSearch Serverless removes a friction point that previously required separate configuration or accepting CPU-bound search latency.
The combination means teams can now deploy inference and retrieval on the same GPU-accelerated stack without separate opt-in steps — the practical path to production-scale RAG workloads on AWS just became shorter. The Exemplar certification adds a third layer: organizations planning large training runs now have AWS-certified GB300 infrastructure in the same environment where their inference and retrieval already run.
Corroborating sources
- Blogs.nvidia
https://blogs.nvidia.com/blog/nvidia-aws-ai-production-scale/
“Building AI systems at scale is demanding, requiring low-latency inference, fast vector search, strong GPU price-performance and infrastructure that can grow without multiplying operational complexity.”