DeepSeek introduces peak/off-peak pricing for DeepSeek-V4-Pro, cutting off-peak rates 50%
DeepSeek will move DeepSeek-V4-Pro and DeepSeek-V4-Flash to a peak/off-peak pricing structure starting at 16:00 UTC on August 16, 2026, giving developers a way to cut inference costs by shifting workloads to lower-demand hours.
What's new
According to DeepSeek's own API announcement, the new schedule splits the day into peak and off-peak windows, with off-peak API calls billed well below standard rates. "Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling," the company said in the announcement. The change applies to both DeepSeek-V4-Pro and DeepSeek-V4-Flash, and model names and endpoints are staying the same — this is a rate-card change, not a new model or API version.
The announcement frames the move as part of a broader V4-Pro production push that also includes flexible reasoning-effort levels (low, high, and max, aimed respectively at simple tasks, daily agent workflows, and complex tasks) and native support for OpenAI's Responses API with what DeepSeek describes as "one-click setup" for Codex integration. DeepSeek did not publish exact per-token dollar figures in the announcement text itself, pointing instead to an accompanying pricing graphic; the 50% off-peak discount is the concrete, on-the-record number.
Context
DeepSeek-V4-Pro exited preview and went generally available earlier this month, positioned as the company's flagship model with heavier agent capabilities than V4-Flash. Time-of-day pricing is a increasingly common move among high-volume inference providers: it lets a lab smooth out GPU utilization across the day by pulling price-sensitive, latency-tolerant traffic — batch jobs, evals, synthetic data generation — into cheaper off-peak windows, while peak-hour interactive traffic keeps paying full price. DeepSeek has leaned hard on aggressive pricing as a competitive lever since its earliest releases, and this restructuring continues that pattern rather than reversing it.
Why it matters
For developers running large, schedulable workloads against DeepSeek's API — bulk data labeling, offline evaluation suites, agent training loops — a 50% off-peak discount is a meaningful lever on unit economics, and it rewards teams that can decouple job timing from real-time user demand. It also signals DeepSeek is optimizing for utilization and margin as V4-Pro scales in production, rather than only competing on flat sticker price. The August 16 effective date gives API users a short window to adjust request scheduling, batching, or rate-limit logic before the new tiers take effect.
Corroborating sources
- Api-docs.deepseek
https://api-docs.deepseek.com/news/news260813/
“Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling.”