The first-party, pay-as-you-go gateway to DeepSeek's V4 models — and as of August 16, 2026 it prices tokens by the clock, with peak/off-peak billing that raised rates roughly 50% to as much as 1,100%.
Last verified · 2026-08-14 · by Moe Ameen
The DeepSeek API is the first-party, pay-as-you-go way to call DeepSeek's open-weight V4 model family without running the weights yourself. It exposes two tiers — V4-Pro, the flagship for the hardest reasoning and coding, and V4-Flash, the fast, cheap volume tier — both text-in, text-out, with a "thinking" reasoning mode, a faster non-thinking mode, and a 1-million-token context window. Because the weights are MIT-licensed on Hugging Face, the API is one of two ways to use the models; the other is to self-host and pay only for compute. This page is about the API and its billing, not a fresh review of the models themselves.
The reason the DeepSeek API is worth a page of its own in August 2026 is that its pricing changed. After warning developers on August 6 that a "significant" increase was coming, DeepSeek published a new rate card that took effect at 16:00 UTC on August 16, 2026. The single flat rate on V4-Flash and V4-Pro was replaced with time-of-day billing: peak hours are 01:00–04:00 and 06:00–10:00 UTC, every other hour is off-peak, and off-peak is set at half the peak rate. The company said it repriced "to allocate resources more reasonably," with the tiers meant to push heavy workloads toward less-congested hours.
The increases are steep and uneven. V4-Flash output moved from a flat $0.28 per million tokens to $0.66 off-peak and $1.32 at peak; V4-Pro output from $0.87 to $1.98 off-peak and $3.96 at peak. On the input side, V4-Flash cache-miss input rose from $0.14 to $0.22/$0.44 (off-peak/peak) and V4-Pro cache-miss input from about $0.435 to $0.66/$1.32. Cached-input rates climbed the most in percentage terms. Overall, reporting put the range at roughly 50% to as much as 1,100% depending on the model, token type, and hour. DeepSeek notes rates can change again, so confirm the live numbers on its pricing page before you budget.
The honest framing for a creator: even at the new peak rates the DeepSeek API is still cheap next to comparable frontier APIs from OpenAI, Anthropic, and Google — the shock is the size and direction of the change, not the absolute figure. What did not change is what the API structurally is. It drafts text and reasons over documents; it generates no image, cuts no clip, records no avatar, designs no carousel, and posts to nothing. It is the upstream drafting layer, now metered by the clock.
The August 16 change makes one workflow decision concrete: draft on DeepSeek's clock, finish on Kompozy's credit. The DeepSeek API is now genuinely good at exactly one job in a content pipeline — producing cheap words, cheapest of all if you batch the heavy drafting into the off-peak UTC windows where the rate is half of peak. What the new rate card does not touch is the part that was always the real work: turning those words into a captioned [Persona Shorts](/glossary/persona-shorts) avatar video, a brand-exact Carousel and Quote Graphics through [HyperFrames](/glossary/hyperframes), Photo Posts, a Blog Article, and an Email Newsletter, then reframing each to 9:16, 1:1, and 16:9 and publishing the set across platforms. A token API does none of that, at any hour.
That finishing-and-distribution layer is [Kompozy](/), and it is metered on purpose differently from DeepSeek: you buy credits and spend them per finished, published asset, with no clock watching whether you generated at 02:00 or 14:00 UTC. So the clean split is to let the DeepSeek API be your batch-drafting endpoint — pour a script or a week of hooks out of V4-Flash off-peak — and hand the raw text to Kompozy, whose own copy engine (Claude and OpenAI, governed by your [Persona Brief](/glossary/persona-brief) and banned-word filters) rewrites it in your voice and renders every format across the eight primary social platforms plus blog and email, with a per-post review pipeline and Autopilot. DeepSeek's price table stays a line item you can batch around; your content cost stays a fixed number of credits per post.
From 16:00 UTC on August 16, 2026, DeepSeek bills V4-Flash and V4-Pro on peak/off-peak rates. V4-Flash output is $0.66 per million tokens off-peak and $1.32 at peak (up from a flat $0.28); V4-Pro output is $1.98 off-peak and $3.96 at peak (up from $0.87). Cache-miss input also rose, and cached input climbed the most in percentage terms. Confirm current rates on DeepSeek's pricing page, as they can change again.
Peak hours are 01:00–04:00 and 06:00–10:00 UTC; every other hour is off-peak. Off-peak rates are half the peak rates, so the same request can cost twice as much depending on when you send it. Batching heavy jobs into off-peak windows is the main way to soften the increase.
Relatively, yes. Even at the new peak rates, the DeepSeek API remains inexpensive next to comparable frontier APIs from OpenAI, Anthropic, and Google. The change is significant because of its size — roughly 50% to as much as 1,100% depending on model, token type, and hour — and because it ends DeepSeek's flat, near-cost pricing, not because the API is now expensive in absolute terms.
No. The DeepSeek API is text-in, text-out — it drafts, summarizes, reasons, and codes, but generates no images, video, or audio and publishes nothing. To turn its scripts and drafts into finished visual posts and distribute them, pair it with a content engine like Kompozy, which renders the media and schedules it across platforms.
Use the DeepSeek API for the drafting step only, batched off-peak, then finish and publish in Kompozy. Kompozy is metered per finished asset in credits rather than per token, so a fixed credit cost per post replaces a variable, time-of-day token bill for everything downstream of the draft.