NVIDIA debuts Nemotron 3 family: 30B Nano with 1M context available now, 100B Super and 550B Ultra on the way
NVIDIA announced the Nemotron 3 family of open models on June 4, 2026 — a three-tier lineup built on a hybrid mixture-of-experts (MoE) architecture that uses NVIDIA's 4-bit NVFP4 training format on Blackwell hardware. The family introduces a new generation of open reasoning models designed for agentic AI workflows, with the Nano tier immediately available and the larger models targeted for the first half of 2026.
What's new
The Nemotron 3 family spans three sizes, each with a distinct active-parameter count and use case:
| Model | Total Params | Active per Token | Context | Primary Use |
|---|---|---|---|---|
| Nemotron 3 Nano | 30B | ~3B | 1M tokens | Edge / high-throughput |
| Nemotron 3 Super | ~100B | ~10B | — | Multi-agent reasoning |
| Nemotron 3 Ultra | ~550B | ~50B | — | Complex AI workflows |
Nemotron 3 Nano is the lead model available at launch. NVIDIA says it delivers "4x higher token throughput compared with Nemotron 2 Nano" and reduces reasoning-token generation by up to 60% — a meaningful improvement for inference cost. Artificial Analysis ranked it "the most open and efficient among models of the same size, with leading accuracy" among open models in its class.
Nemotron 3 Super is designed for multi-agent applications that require collaborative reasoning across multiple model instances. Nemotron 3 Ultra positions as an advanced reasoning engine for demanding AI workflows, operating closer to frontier-model capability.
All three models use NVFP4 training on NVIDIA's Blackwell architecture, reducing memory requirements and accelerating training while maintaining accuracy at lower numerical precision.
Availability:
- Nemotron 3 Nano: available immediately on Hugging Face, with inference through Baseten, DeepInfra, Fireworks, and Together AI
- Nemotron 3 Super and Ultra: expected H1 2026
- License: not specified in the announcement
Context
Nemotron 3 is the third generation of NVIDIA's own open-model series, which started as a benchmark-training exercise and has expanded into a production-grade offering. The series runs on NVIDIA infrastructure and serves as a reference implementation for what's achievable on Blackwell hardware.
The 1M-token context window on the Nano model is notable at that parameter count. Most 30B-class open models have not offered long context alongside competitive accuracy — the usual trade-off is context versus efficiency. Nemotron 3 Nano's 4x throughput improvement suggests NVIDIA is prioritizing inference economics alongside capability.
The MoE architecture (sparse activation) means the 30B total parameters translate to roughly 3B of active compute per token — closer in inference cost to a dense 3–4B model while retaining the knowledge capacity of a larger network. This is the same architectural pattern used in DeepSeek-V4, Mixtral, and Qwen MoE models.
Why it matters
NVIDIA's decision to release open models — not just chips and software — reflects how competitive the inference and model layer has become. Nemotron models run on NVIDIA hardware and are distributed through NVIDIA's partner inference network, which keeps the company relevant in the model layer even as open-weight alternatives from Meta, DeepSeek, and others commoditize base capabilities.
For builders, Nemotron 3 Nano offers a compelling combination: high throughput, 1M context, and open weights, distributed through inference providers that already run at scale. The Super and Ultra tiers, once available, will expand the family into multi-agent and complex-reasoning territory where proprietary frontier models currently hold the advantage.
Corroborating sources
- Nvidianews.nvidia
https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models
“4x higher token throughput compared with Nemotron 2 Nano”