NVIDIA's Vera Rubin NVL72 delivers up to 3.7x the throughput of GB300 in its first MLPerf Inference submission
NVIDIA published its first MLPerf Inference results for Vera Rubin NVL72 on September 16, 2026, reporting up to 3.7x higher throughput than the current-generation GB300 NVL72 system on the v6.1 benchmark round, alongside near-linear scaling as the system grows from one rack to four.
What's new
- "In its first MLPerf Inference preview submission, NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than GB300 NVL72," NVIDIA wrote, citing results on the Qwen3-VL workload; the company also reported up to 2.5x higher throughput than GB300 on DeepSeek-R1.
- On agentic workloads specifically, NVIDIA cites a 30x performance advantage for Vera Rubin over GB300 on SemiAnalysis's AgentX benchmark suite.
- Separately, NVIDIA highlighted scaling efficiency on the existing GB300 NVL72 platform: "NVIDIA's DeepSeek-R1 submission scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency," meaning near-zero performance loss as the same workload is spread across four times as many GPUs.
- NVIDIA also reported pure software gains on the existing generation, with Qwen3-VL throughput improving up to 1.6x between the v6.0 and v6.1 rounds of MLPerf on the same GB300 hardware.
Context
MLPerf Inference is the industry's standard, third-party-audited benchmark suite for comparing AI hardware on realistic inference workloads, run twice a year by the MLCommons consortium. A "preview" submission, as NVIDIA is making here for Vera Rubin, lets a vendor post results on hardware that has not yet shipped in volume, which is standard practice ahead of a new GPU generation's general availability. GB300 NVL72 is NVIDIA's current flagship rack-scale system; Vera Rubin is the next-generation platform succeeding it. The scaling and software-only results reported for GB300 in the same submission show NVIDIA continuing to extract more performance from already-deployed hardware even as it prepares to ship the next generation.
Why it matters
For NVIDIA's customers — cloud providers and AI labs planning multi-billion-dollar compute buildouts — these are the first credible, standardized numbers on what Vera Rubin actually delivers over the hardware they're buying today, and they arrive at a moment when every major AI lab is capacity-constrained. A 3.7x throughput gain on a workload like Qwen3-VL, if it holds up once the hardware ships broadly, would materially change the compute-per-dollar math for anyone deciding whether to wait for Vera Rubin or buy GB300 now. The 30x agentic-workload figure is the more pointed number for NVIDIA's pitch: agentic and reasoning workloads are widely seen as the fastest-growing and most compute-hungry category of AI inference, and framing Vera Rubin's advantage specifically around that benchmark signals where NVIDIA expects the next wave of demand — and spending — to concentrate. The 99% scaling efficiency result on GB300 is a smaller but practically important data point for buyers already running or planning large multi-rack deployments: it says that scaling a demanding model like DeepSeek-R1 across four racks does not meaningfully sacrifice per-GPU efficiency, a common failure point in large distributed inference setups.
Corroborating sources
- Blogs.nvidia
https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/
“In its first MLPerf Inference preview submission, NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than GB300 NVL72.”