NVIDIA Vera Rubin NVL72 ramps into production, delivering 10x tokens per megawatt across CoreWeave, Google Cloud, and Nebius
NVIDIA's Vera Rubin NVL72 rack-scale AI system is now ramping into production, with early cloud partners publishing the first live-hardware benchmarks and NVIDIA detailing a wave of global deployments spanning CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, Nebius, and Mistral.
What's new
NVIDIA says Vera Rubin NVL72 production racks are running at partners "CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius," across "350+ factory sites in 30 countries." CoreWeave, the first cloud to bring up and validate the system, published the first measured numbers from live hardware: a DeepSeek-R1 benchmark showing "10x more throughput per megawatt than Grace Blackwell NVL72 — landing directly on the metric that matters most for power-constrained AI factories." Google Cloud's new A5X instance, built on Vera Rubin NVL72, is now running for London AI startup Ineffable Intelligence and is reported to deliver "up to 10x lower inference cost per token and 10x higher token throughput per megawatt than the prior generation." Ineffable Intelligence cofounder Lasse Espeholt said: "The next era of research requires the next era of hardware. We feel privileged to work with the teams at NVIDIA and Google Cloud, who were able to grant us early access to Vera Rubin."
The platform's gains come from what NVIDIA calls "extreme codesign" across seven chips and five rack trays — the Vera Rubin NVL72 GPU, Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX switch, and Vera BlueField-4 STX — engineered as a single system. The Vera CPU's custom Olympus core delivers 2x single-threaded performance and 3x core-to-core bandwidth versus competing chiplet designs, and independent benchmarks from cloud provider DeepInfra found the Vera CPU supports "up to 1.6x more concurrent AI agents at the same quality of service and up to 2.2x faster orchestration than alternative CPUs." On networking, sixth-generation NVLink delivers more than 2x throughput and 10x higher packet rates than off-the-shelf Ethernet, while Spectrum-X Ethernet's 102.4T Spectrum-6 switches enable 1.6x higher RDMA bandwidth; NVIDIA says CoreWeave, Microsoft, SpaceX AI, and Tesla are among the first to adopt the new switches.
Separately, NVIDIA detailed a new use for the platform in Europe: an expanded Microsoft-Mistral partnership, backed by a multibillion-dollar infrastructure agreement, will run on Vera Rubin GPUs to give European governments and regulated industries open models on cloud, cloud-connected, and fully disconnected private infrastructure. NVIDIA says Vera Rubin NVL72 delivers "up to 10x more tokens per megawatt and one-tenth the cost per million tokens" compared with its prior-generation GB200 NVL72. On July 24, NVIDIA added that Nebius will bring Vera Rubin NVL72 and Spectrum-6 to its cloud in Europe and the US, having received its first system at its Finland AI Factory.
Context
Vera Rubin succeeds NVIDIA's Grace Blackwell (GB200/GB300) generation, which the company has spent the past year deploying with partners including Bristol Myers Squibb and the Naval Postgraduate School. This release marks the point where Vera Rubin moves from platform announcement to validated, live-hardware production numbers from multiple independent cloud operators, rather than NVIDIA's own projections.
Why it matters
Tokens-per-megawatt has emerged as the metric AI infrastructure buyers now optimize for as power, not chip supply, becomes the binding constraint on scaling AI factories. Independent, order-of-magnitude gains reported by CoreWeave, Google Cloud, and DeepInfra — rather than NVIDIA's own marketing claims alone — give enterprises and sovereign cloud operators concrete grounds to plan multi-year capacity commitments around the platform. The parallel Microsoft-Mistral sovereign AI deal also signals that rack-scale efficiency gains are becoming a lever in the broader contest over where regulated industries and governments choose to run frontier AI workloads.
Corroborating sources
- Blogs.nvidia
https://blogs.nvidia.com/blog/vera-rubin/
“CoreWeave ran a DeepSeek-R1 benchmark on Vera Rubin NVL72 and saw 10x improvement in tokens per second per megawatt compared with Grace Blackwell NVL72.”