NVIDIA says Vera Rubin NVL72 delivers 30x more agentic AI throughput per watt than GB300
NVIDIA published new efficiency benchmarks on August 24, 2026 for its Vera Rubin NVL72 rack-scale system, claiming a dramatic jump in performance-per-watt and per-token cost over its current-generation GB300 NVL72 platform when running real agentic AI workloads.
What's new
According to NVIDIA's own blog post, "New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30x higher throughput per megawatt and 35x lower token costs than NVIDIA GB300 NVL72." NVIDIA says the figures were generated using the SemiAnalysis AgentX workload, a benchmark built on actual agentic coding sessions rather than synthetic throughput tests — an important distinction, since agentic workloads (multi-step tool calls, long context, iterative reasoning) tend to stress systems differently than single-shot inference benchmarks.
The two headline numbers — 30x higher throughput per megawatt and 35x lower token cost — are both framed as agentic-workload-specific gains rather than general-purpose throughput improvements, suggesting NVIDIA is tuning its messaging (and likely aspects of the Vera Rubin architecture itself) specifically toward the agent-and-coding-assistant use case that has become the dominant driver of enterprise AI infrastructure demand in 2026.
Context
Vera Rubin is NVIDIA's next-generation rack-scale AI system, the successor to the Blackwell-based GB300 NVL72 that data center operators have been deploying through 2026. NVIDIA has been rolling out Vera Rubin details and partner deployments in stages this month, including a deal with SpaceXAI to adopt NVIDIA Vera CPUs at scale and an expansion of NVLink Fusion to support custom XPUs in large AI factory buildouts. Per-watt and per-token efficiency have become NVIDIA's primary competitive argument as power availability, not raw chip supply, increasingly caps how much AI compute large operators can deploy.
Why it matters
Power and cost per token are now the binding constraints on scaling agentic AI systems for most large operators, more so than raw model capability. A claimed 30x throughput-per-watt improvement — even measured on NVIDIA's own chosen benchmark — would materially change the economics of running large fleets of coding agents and other long-running agentic workloads, since power draw and token cost compound directly into what an operator can afford to run at scale. As with any vendor-published benchmark, the numbers should be read as NVIDIA's own measurement rather than independently verified, and real-world gains will depend on how closely a given deployment's workload matches the AgentX-style agentic coding trajectories NVIDIA used for testing. Independent benchmarking once Vera Rubin systems ship broadly will be the real test of whether the efficiency claims hold up outside NVIDIA's own lab.
Corroborating sources
- Blogs.nvidia
https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/
“that translates directly into 30x more agentic work for the same energy footprint”