NVIDIA Rubin platform: six new chips deliver 10x lower inference cost and 4x fewer GPUs for MoE training
NVIDIA unveiled the Rubin platform, a next-generation AI computing architecture built from six co-designed chips that the company says delivers 10x lower inference token cost and requires 4x fewer GPUs to train mixture-of-experts models compared to the Blackwell generation. Products are set to be available from major cloud providers in the second half of 2026.
What's new
The Rubin platform comprises six chips designed as an integrated system:
- NVIDIA Vera CPU: 88 custom Olympus cores with Armv9.2 compatibility
- NVIDIA Rubin GPU: Third-generation Transformer Engine with 50 petaflops of NVFP4 compute
- NVIDIA NVLink 6 Switch: 3.6TB/s bandwidth per GPU
- NVIDIA ConnectX-9 SuperNIC: Advanced networking for AI workloads
- NVIDIA BlueField-4 DPU: Data processing unit for storage and security
- NVIDIA Spectrum-6 Ethernet Switch: Next-generation network fabric
Two rack-scale configurations ship with the platform:
- Vera Rubin NVL72: 72 GPUs and 36 CPUs per rack, 260TB/s aggregate bandwidth — the configuration for large-scale training and inference
- HGX Rubin NVL8: 8-GPU server board for x86-based platforms — the integration path for existing infrastructure
Key performance claims versus the Blackwell generation:
- 10x lower inference token cost
- 4x fewer GPUs needed for MoE model training
- 5x better power efficiency in Spectrum-X photonics-based systems
- 18x faster rack assembly and servicing
Availability: AWS, Google Cloud, Microsoft, OCI, CoreWeave, Lambda, Nebius, and Nscale in H2 2026.
Jensen Huang stated: "Rubin takes a giant leap toward the next frontier of AI" through "extreme codesign across six new chips."
Context
Rubin follows Blackwell, which launched in 2025 and became the dominant training and inference platform for frontier AI models. The Vera CPU is notable: it replaces third-party x86 CPUs in NVIDIA's own rack-scale system with NVIDIA-designed Arm-based cores, giving the company more control over the full compute stack in NVL72 deployments. The NVLink 6 bandwidth increase (3.6TB/s per GPU) is the connective tissue that allows dozens of GPUs to act as one large memory pool — critical for models with trillion-parameter scales.
Why it matters
The 10x inference cost reduction is the number that will get enterprise attention. For organizations already spending significantly on GPU inference, a generational leap of that magnitude can change the economics of deploying frontier models in production. The 4x fewer GPUs for MoE training also matters: as MoE architectures dominate the frontier (from Mixtral to GPT-4 to Grok), training cost is the main barrier to developing competitive models. A platform that cuts that cost by 75% expands who can realistically compete in frontier model development. The six-chip co-design approach also signals that NVIDIA's moat is moving further up the system stack — not just the GPU, but the full rack, networking, and CPU.
Corroborating sources
- Investor.nvidia
https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Kicks-Off-the-Next-Generation-of-AI-With-Rubin--Six-New-Chips-One-Incredible-AI-Supercomputer/default.aspx
“Rubin takes a giant leap toward the next frontier of AI”