DeepSeek’s official API docs now post a clock. From 16:00 UTC on August 16, 2026, V4 Flash and V4 Pro move to peak and off-peak billing. Off-peak is half of peak. Peak windows are 01:00–04:00 and 06:00–10:00 UTC. Every other hour is off-peak.
Today’s posted rates are still the flat ones: Flash is $0.0028 per million input tokens on a cache hit, $0.14 on a miss, and $0.28 on output. Pro is $0.003625 / $0.435 / $0.87. After the cutover, Flash off-peak becomes $0.007 / $0.22 / $0.66, and peak is $0.014 / $0.44 / $1.32. Pro off-peak is $0.022 / $0.66 / $1.98; peak is $0.044 / $1.32 / $3.96. Cache-hit input is the sharpest jump, especially on Pro.
The models themselves do not change in this notice. Flash is DeepSeek-V4-Flash-0731. Pro is DeepSeek-V4-Pro-0813. Both still advertise a 1 million token context, up to 384K output, thinking and non-thinking modes, JSON, tools, and OpenAI- plus Anthropic-format endpoints. Concurrency stays 2,500 on Flash and 500 on Pro.
If you run batch jobs, the practical move is to shift work out of those two UTC peak bands. If you live in long chats that hit the cache, budget for a higher bill even off-peak. DeepSeek says it can change prices again and tells customers to check the pricing page. This writeup uses that official page, not a third-party recap.
Source: https://api-docs.deepseek.com/quick_start/pricing
