Moonshot AI releases Kimi K3, a 2.8-trillion-parameter open-weight model
Moonshot AI unveiled Kimi K3 on July 16, 2026, a 2.8-trillion-parameter model the company calls the world's first open 3-trillion-class model. The API is live today on Kimi.com, Kimi Work, Kimi Code, and the Kimi API, with full model weights scheduled for public release by July 27.
What's new
Kimi K3 is built on two new architectural components: Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), which Moonshot says improve how information flows across sequence length and model depth. The model uses a Stable LatentMoE framework with 896 total experts, of which only 16 are active per inference — a sparsity ratio the company says delivers roughly a 2.5x improvement in scaling efficiency over the prior Kimi K2 generation.
Specs at a glance:
- 2.8 trillion total parameters
- 1-million-token context window
- Native visual understanding (not a bolted-on vision module)
- Weights quantized to MXFP4, with MXFP8 activations
- "Thinking mode" reasoning enabled by default
Pricing on the Kimi API is tiered by cache status: $0.30 per million tokens for cache-hit input, $3.00 per million for cache-miss input, and $15.00 per million for output tokens.
In its own release materials, Moonshot describes the model directly: "Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window." The company frames K3's ambitions plainly: "It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning."
Context
Moonshot has moved fast since Kimi K2 — this is the second major model generation from the Beijing-based lab in under a year, and it lands just ahead of the World Artificial Intelligence Conference in Shanghai. K3's long-horizon coding focus, including sustaining extended engineering sessions, navigating large repositories, and orchestrating terminal tools, puts it squarely in competition with the agentic-coding push from Anthropic, OpenAI, and DeepSeek. Coverage from VentureBeat and TechCrunch both frame K3 as Moonshot's most direct attempt yet to close the gap with frontier closed models rather than simply undercutting them on price, as DeepSeek and Z.ai have generally done with their open releases.
The staggered rollout — API access now, full weights in 11 days — mirrors a pattern several open-weight labs have adopted this year: ship the hosted product first to capture usage and feedback, then follow with the weights once serving infrastructure and safety review are settled.
Why it matters
At 2.8 trillion parameters, Kimi K3 is by a wide margin the largest open-weight model yet released, and Moonshot's own benchmark claims put it ahead of Anthropic's Opus 4.8 and OpenAI's GPT-5.6 Sol and GPT-5.5, while still trailing Anthropic's newest Fable 5 tier. Those are the lab's own numbers pending independent verification, but even a partial confirmation would mark a meaningful shift: an open-weight model within striking distance of the frontier, not just a cheaper alternative several rungs below it. For developers and enterprises weighing build-vs-buy decisions on model infrastructure, a model with genuine frontier-level coding and reasoning performance — and weights due for self-hosting in under two weeks — raises the stakes for how much of the frontier gap closed-model providers can still charge a premium for.
Corroborating sources
- Techcrunch
https://techcrunch.com/2026/07/16/moonshots-upcoming-kimi-3-is-expected-to-close-the-gap-with-anthropics-opus-4-8/
- Venturebeat
https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems
- Kimi
https://www.kimi.com/blog/kimi-k3
“Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window.”