NVIDIA says GB300 NVL72 delivers up to 25x performance per watt over Hopper on open models
NVIDIA published new performance-per-watt figures on July 14, 2026, arguing the metric — not raw throughput — is now the deciding factor in AI infrastructure economics, and backing that claim with rack-scale efficiency numbers for its GB300 NVL72 platform.
What's new
According to NVIDIA's blog post, GB300 NVL72 delivers up to 25x the performance per watt of the Hopper generation on DeepSeek V4 Pro, one of the newest open-weight models. The company cited additional gains across other current open models:
- Up to 20x performance-per-watt improvement on GLM 5.1
- Up to 10x performance-per-watt improvement on Kimi K2.6
- Software-only optimizations lifted DeepSeek V4 performance by up to 5x in a single month, without new hardware
- NVIDIA's DSX MaxLPS power-management software lets operators run up to 40% more GPUs within the same facility power budget
NVIDIA named production deployments already running on the platform, citing Anthropic, OpenAI, CoreWeave, Perplexity, and Fireworks AI among the companies using GB300-class infrastructure in live workloads.
Context
The post continues a shift in how NVIDIA pitches its hardware generations: rather than leading with raw FLOPs, the company is increasingly framing GB300 NVL72 around watts-per-token economics — the metric that determines how many GPUs a data center operator can actually power and cool at scale, not just how fast an individual chip runs. It follows a string of recent NVIDIA infrastructure announcements, including revenue-share compute financing deals and multibillion-dollar U.S. manufacturing commitments, all aimed at the same underlying bottleneck: power availability, not chip supply, increasingly caps how much AI infrastructure can be deployed.
The specific benchmarks against DeepSeek V4 Pro, GLM 5.1, and Kimi K2.6 — all open-weight models from Chinese labs — also signal how central open-model inference has become to NVIDIA's efficiency pitch, alongside its more traditional frontier-lab customers.
Why it matters
Power, not silicon, is the binding constraint on data center buildouts industry-wide, and a 25x performance-per-watt gain (if it holds up under independent verification) would materially change the math for operators sizing new facilities. Software-only gains of up to 5x in a month also suggest a meaningful share of near-term efficiency improvements will come from optimization work rather than new hardware generations — good news for operators who can't get new racks installed fast enough to keep up with demand.
Corroborating sources
- Blogs.nvidia
https://blogs.nvidia.com/blog/performance-per-watt-ai-infrastructure-efficiency/
“Across the newest generation of leading open models, NVIDIA GB300 NVL72 delivers up to 25x performance per watt compared with the NVIDIA Hopper generation.”