DeepSeek is replacing its flat API rates with time-of-day billing for V4-Flash and V4-Pro on August 16, 2026. Depending on the model, token type, and hour, prices rise from about 50% to as much as 1,100% — though they stay cheap in absolute terms.
2026-08-13 · by Moe Ameen
DeepSeek is raising the prices developers pay to call its models. After posting a warning on its pricing documentation on August 6, 2026 that a "significant" increase was coming, the Hangzhou-based lab published the actual rate card in mid-August. The new prices take effect at 16:00 UTC on August 16, 2026, and they replace the single flat rate on V4-Flash and V4-Pro with time-of-day billing.
The new structure splits every rate into peak and off-peak, with off-peak set at half the peak price. DeepSeek defines peak hours as 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak. The increases are steep and uneven. V4-Flash output tokens go from a flat $0.28 per million to $0.66 off-peak and $1.32 at peak; V4-Pro output goes from $0.87 to $1.98 off-peak and $3.96 at peak. On the input side, V4-Flash cache-miss input rises from $0.14 to $0.22/$0.44 (off-peak/peak) and V4-Pro cache-miss input from about $0.435 to $0.66/$1.32. Cached-input rates climb the most in percentage terms — the V4-Pro cache-hit rate, currently about $0.003625 per million, takes the steepest percentage jump of all, landing near the top of the increase range. Overall, reporting put the range of increases at roughly 50% to as much as 1,100% depending on model, token type, and time of day.
The context matters as much as the numbers. DeepSeek built its reputation partly on price: V4-Flash was recently flagged by Artificial Analysis as the cheapest well-known model to run globally, and the lab had run promotional discounts as deep as 75% on V4-Pro earlier in 2026. This update reads as the end of that race to the bottom — a normalization of a business that had been pricing near cost. Even after the increase, the rates remain low next to comparable Western frontier models; the shock is the size and the direction of the change, not the absolute figure.
DeepSeek says it is repricing "to allocate resources more reasonably," with the peak/off-peak tiers meant to push workloads toward less-congested off-peak hours, and it notes on its pricing page that rates may change again. As always with a fast-moving model line, treat the figures as a snapshot and confirm current rates on DeepSeek's own page before committing budget.
The practical question this raises for a creator is not "which model is cheapest this week" — it is "why is my content cost tied to a token meter I do not control at all?" DeepSeek just doubled-to-12x its rates with ten days' notice and added a peak/off-peak wrinkle that penalizes you for generating during normal working hours in Asia. If your content pipeline's economics live and die on one lab's token price, this is the second time in 2026 that assumption moved under you.
[Kompozy](/) is metered differently on purpose. You buy credits and spend them per finished asset — a captioned [Persona Short](/glossary/persona-shorts), a brand-exact carousel via [HyperFrames](/glossary/hyperframes), a blog, a newsletter, a batch of platform-native posts — not per token, and never with a clock watching whether you generated at 02:00 or 14:00 UTC. The generation inside runs on Claude and OpenAI, governed by your [Persona Brief](/glossary/persona-brief), and on the Founding tier you can bring your own API keys to control model cost directly. So the move today is not to chase DeepSeek's new rate card across a spreadsheet — it is to price your content in finished, published posts across the eight social platforms plus blog and email, and let token-price swings be someone else's problem.
At 16:00 UTC on August 16, 2026. DeepSeek warned of a "significant" increase on August 6, then published the rate card in mid-August. The current flat rates apply through August 15; the new peak/off-peak structure begins August 16.
Reporting put the range at roughly 50% to as much as 1,100%, depending on the model (V4-Flash or V4-Pro), the token type (cached input, cache-miss input, or output), and the time of day. The largest percentage jumps are on cached-input tokens; output tokens roughly double to quadruple.
Peak hours are 01:00–04:00 and 06:00–10:00 UTC; every other hour is off-peak. Off-peak rates are set at half the peak rates, so the same request can cost twice as much depending on when you send it.
Relatively, yes. Even at the new peak rates, DeepSeek remains inexpensive next to comparable frontier models from OpenAI, Anthropic, and Google. The change is significant because of its size and direction — it ends DeepSeek's flat, near-cost pricing — not because the API is now expensive in absolute terms.