DeepSeek releases V4-Pro in general availability with major agent-capability upgrades
DeepSeek shipped the general-availability release of DeepSeek-V4-Pro on August 13, 2026, moving its flagship model out of preview with a significant upgrade to agent capabilities and native support for OpenAI's Responses API format.
What's new
DeepSeek's own changelog states plainly: "The GA release of DeepSeek-V4-Pro has been rolled out on the APP, Web, and API." The GA build, designated DeepSeek-V4-Pro-0813, is built around agentic performance — the tasks where a model has to call tools, execute code, and carry out multi-step workflows on its own rather than answer a single prompt.
The V4-Pro API now works with OpenAI's Responses API format out of the box and ships with built-in support for Codex integration, letting existing Codex-based tooling point at DeepSeek's model with minimal changes. Thinking effort is now configurable across three levels — low, high, and max — for both V4-Pro and the lighter V4-Flash, so developers can trade latency for reasoning depth per task. The model supports a context window of up to 1 million tokens and can generate outputs as long as 384,000 tokens, running in either thinking or non-thinking mode.
Pricing also changed alongside GA: DeepSeek introduced a peak/off-peak structure, with off-peak pricing set at 50% of the peak rate, effective August 16, 2026 at 16:00 UTC. The move follows V4-Flash's public beta release on July 31 and preceded the experimental V4-Flash-Vision model released August 21, which added visual understanding for agent benchmarks — together forming a rapid, month-long V4 rollout across DeepSeek's model line.
Context
DeepSeek has spent the past several months iterating quickly through the V4 family — Flash in public beta, Pro moving to general availability, then a vision-capable experimental variant — while pushing hard on agent-specific capability and API compatibility with the tooling ecosystem that has grown up around OpenAI's formats. Native Responses API and Codex support in particular lowers the switching cost for developers already building on OpenAI-shaped tooling who want to route work to DeepSeek's models instead.
Why it matters
V4-Pro's GA release is DeepSeek confirming its flagship model is ready for production agent workloads, not just benchmark demonstrations — a distinction that matters given how much of the current AI competition has shifted from raw chat quality to reliability in multi-step, tool-using tasks. Shipping compatibility with the Responses API and Codex specifically targets developers who've already built around OpenAI's ecosystem, making DeepSeek a lower-friction alternative rather than requiring a separate integration. Combined with the peak/off-peak pricing model, it's a pitch aimed squarely at cost-sensitive agent deployments running at scale, where DeepSeek has consistently positioned itself against higher-priced frontier models from U.S. labs.
Corroborating sources
- Api-docs.deepseek
https://api-docs.deepseek.com/updates/
“The GA release of DeepSeek-V4-Pro has been rolled out on the APP, Web, and API.”