DeepSeek ships V4.1-Flash, a smaller, cheaper model that beats its own flagship on benchmarks
DeepSeek released V4.1-Flash on September 10, 2026, a new smaller model in its V4 line that the company says outperforms its own larger V4-Pro flagship on speed, cost, and total runtime.
What's new
According to DeepSeek's own release notes, V4.1-Flash is "the smallest model in our new architecture family, with native visual understanding." The model uses a 552-billion-parameter Mixture-of-Experts design built on what DeepSeek calls a Causal Encoder-Decoder architecture, activating just 8 billion parameters for input processing and 16 billion for output generation.
DeepSeek says the smaller active-parameter footprint translates directly into cheaper inference: the model's KV cache needs only one-quarter the HBM and one-eighth the SSD storage of the prior generation, which the company says cuts the cache-hit charges that make up a large share of typical agent costs.
On performance, the company states plainly: "Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime." As a result, DeepSeek is retiring V4-Pro outright. Starting September 14, 2026, all deepseek-v4-pro API requests will route to V4.1-Flash at V4.1-Flash's rates, an arrangement that continues until a future V4.1-Pro model ships. The older V4-Flash and V4-Flash-Vision-Exp models are also retired, with their API endpoints temporarily redirected to V4.1-Flash for compatibility.
Pricing changes took effect September 10, 2026 at 04:00 UTC, with DeepSeek continuing its peak/off-peak pricing model where off-peak rates run at half the peak rate. Model weights are published on Hugging Face, alongside a technical report, and DeepSeek says it will work with the open-source community on inference support and deployment options. Coding-agent partners WorkBuddy (including CodeBuddy) and OpenCode have already added full support for the new model.
Context
V4.1-Flash is the latest entry in a rapid release cadence for DeepSeek's V4 line this year: V4 Preview shipped in April 2026, V4-Pro reached general availability in August, and a vision-focused V4-Flash variant followed shortly after in late August. Retiring both V4-Pro and the prior Flash variants within weeks of releasing V4.1-Flash is an unusually fast turnover, even by DeepSeek's own pace.
Why it matters
DeepSeek's pitch here is unusual: a smaller, cheaper model claiming to beat the company's own more expensive flagship on real-world metrics, not just cost. If that holds up under independent testing, it reinforces DeepSeek's pattern of using architecture efficiency, rather than raw parameter scale, to compete with better-funded labs on both capability and price. The aggressive retirement schedule for V4-Pro and the older Flash models also signals confidence that V4.1-Flash is a strict upgrade rather than a cheaper, weaker alternative, since customers on the old models get force-migrated rather than given an indefinite choice.
Corroborating sources
- Api-docs.deepseek
https://api-docs.deepseek.com/news/news260910/
“Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime.”