// AI TOOLS · DEEPSEEK API

DeepSeek API

The first-party, pay-as-you-go gateway to DeepSeek's V4 models — and as of August 16, 2026 it prices tokens by the clock, with peak/off-peak billing that raised rates roughly 50% to as much as 1,100%.

Last verified · 2026-08-14 · by Moe Ameen

What DeepSeek API is

The DeepSeek API is the first-party, pay-as-you-go way to call DeepSeek's open-weight V4 model family without running the weights yourself. It exposes two tiers — V4-Pro, the flagship for the hardest reasoning and coding, and V4-Flash, the fast, cheap volume tier — both text-in, text-out, with a "thinking" reasoning mode, a faster non-thinking mode, and a 1-million-token context window. Because the weights are MIT-licensed on Hugging Face, the API is one of two ways to use the models; the other is to self-host and pay only for compute. This page is about the API and its billing, not a fresh review of the models themselves.

The reason the DeepSeek API is worth a page of its own in August 2026 is that its pricing changed. After warning developers on August 6 that a "significant" increase was coming, DeepSeek published a new rate card that took effect at 16:00 UTC on August 16, 2026. The single flat rate on V4-Flash and V4-Pro was replaced with time-of-day billing: peak hours are 01:00–04:00 and 06:00–10:00 UTC, every other hour is off-peak, and off-peak is set at half the peak rate. The company said it repriced "to allocate resources more reasonably," with the tiers meant to push heavy workloads toward less-congested hours.

The increases are steep and uneven. V4-Flash output moved from a flat $0.28 per million tokens to $0.66 off-peak and $1.32 at peak; V4-Pro output from $0.87 to $1.98 off-peak and $3.96 at peak. On the input side, V4-Flash cache-miss input rose from $0.14 to $0.22/$0.44 (off-peak/peak) and V4-Pro cache-miss input from about $0.435 to $0.66/$1.32. Cached-input rates climbed the most in percentage terms. Overall, reporting put the range at roughly 50% to as much as 1,100% depending on the model, token type, and hour. DeepSeek notes rates can change again, so confirm the live numbers on its pricing page before you budget.

The honest framing for a creator: even at the new peak rates the DeepSeek API is still cheap next to comparable frontier APIs from OpenAI, Anthropic, and Google — the shock is the size and direction of the change, not the absolute figure. What did not change is what the API structurally is. It drafts text and reasons over documents; it generates no image, cuts no clip, records no avatar, designs no carousel, and posts to nothing. It is the upstream drafting layer, now metered by the clock.

What you can make with it

  • Cheap, high-volume text drafting — captions, hooks, post copy, and short scripts (V4-Flash)
  • Long-form drafts and outlines — blog posts, article skeletons, video scripts — from either tier
  • Reasoning over a messy source pile with the 1M-token context: pull angles, structure a content calendar, extract a brief (V4-Pro thinking mode)
  • Batch generation — dozens of headline, subject-line, or CTA variations in a single call
  • Summaries and repurposing drafts from a full transcript, PDF, or course
  • Off-peak batching: schedule heavy jobs into the half-price UTC windows to blunt the increase

How Kompozy turns DeepSeek API output into content

The August 16 change makes one workflow decision concrete: draft on DeepSeek's clock, finish on Kompozy's credit. The DeepSeek API is now genuinely good at exactly one job in a content pipeline — producing cheap words, cheapest of all if you batch the heavy drafting into the off-peak UTC windows where the rate is half of peak. What the new rate card does not touch is the part that was always the real work: turning those words into a captioned [Persona Shorts](/glossary/persona-shorts) avatar video, a brand-exact Carousel and Quote Graphics through [HyperFrames](/glossary/hyperframes), Photo Posts, a Blog Article, and an Email Newsletter, then reframing each to 9:16, 1:1, and 16:9 and publishing the set across platforms. A token API does none of that, at any hour.

That finishing-and-distribution layer is [Kompozy](/), and it is metered on purpose differently from DeepSeek: you buy credits and spend them per finished, published asset, with no clock watching whether you generated at 02:00 or 14:00 UTC. So the clean split is to let the DeepSeek API be your batch-drafting endpoint — pour a script or a week of hooks out of V4-Flash off-peak — and hand the raw text to Kompozy, whose own copy engine (Claude and OpenAI, governed by your [Persona Brief](/glossary/persona-brief) and banned-word filters) rewrites it in your voice and renders every format across the eight primary social platforms plus blog and email, with a per-post review pipeline and Autopilot. DeepSeek's price table stays a line item you can batch around; your content cost stays a fixed number of credits per post.

  1. Batch your DeepSeek API drafting — scripts, hooks, outlines — into the off-peak UTC windows (any hour outside 01:00–04:00 and 06:00–10:00) to pay half the peak rate.
  2. Bring the raw text into Kompozy as the source for a format.
  3. Let Kompozy rewrite it in your voice through the Persona Brief and render the finished formats — persona video, carousels, images, blog, newsletter, text posts.
  4. Review the batch in one pipeline; approve or edit inline.
  5. Schedule and publish across the eight social platforms plus blog and email on Autopilot — priced per asset in credits, not per token by the hour.

Frequently asked questions

How much does the DeepSeek API cost after the August 16, 2026 price increase?

From 16:00 UTC on August 16, 2026, DeepSeek bills V4-Flash and V4-Pro on peak/off-peak rates. V4-Flash output is $0.66 per million tokens off-peak and $1.32 at peak (up from a flat $0.28); V4-Pro output is $1.98 off-peak and $3.96 at peak (up from $0.87). Cache-miss input also rose, and cached input climbed the most in percentage terms. Confirm current rates on DeepSeek's pricing page, as they can change again.

What are DeepSeek's peak and off-peak hours?

Peak hours are 01:00–04:00 and 06:00–10:00 UTC; every other hour is off-peak. Off-peak rates are half the peak rates, so the same request can cost twice as much depending on when you send it. Batching heavy jobs into off-peak windows is the main way to soften the increase.

Is the DeepSeek API still cheap after the hike?

Relatively, yes. Even at the new peak rates, the DeepSeek API remains inexpensive next to comparable frontier APIs from OpenAI, Anthropic, and Google. The change is significant because of its size — roughly 50% to as much as 1,100% depending on model, token type, and hour — and because it ends DeepSeek's flat, near-cost pricing, not because the API is now expensive in absolute terms.

Can the DeepSeek API generate images, video, or social posts?

No. The DeepSeek API is text-in, text-out — it drafts, summarizes, reasons, and codes, but generates no images, video, or audio and publishes nothing. To turn its scripts and drafts into finished visual posts and distribute them, pair it with a content engine like Kompozy, which renders the media and schedules it across platforms.

How do I keep my content cost predictable if DeepSeek prices by the hour?

Use the DeepSeek API for the drafting step only, batched off-peak, then finish and publish in Kompozy. Kompozy is metered per finished asset in credits rather than per token, so a fixed credit cost per post replaces a variable, time-of-day token bill for everything downstream of the draft.

Related tools

  • DeepSeek V4The latest generation of DeepSeek's open-weight model family — a two-tier lineup (V4-Pro and V4-Flash) of mixture-of-experts language models with a 1M-token context, MIT-licensed weights, and API pricing well below Western frontier models.
  • DeepSeek V4 Pro 0813The general-availability build of DeepSeek's flagship model — a 1.6-trillion-parameter mixture-of-experts LLM with a 1M-token context, MIT-licensed weights, and API pricing well under Western frontier models.
  • DeepSeek-V4-FlashDeepSeek's fast, low-cost frontier language model — a 284B-parameter mixture-of-experts LLM (13B active) with a 1M-token context, open weights under the MIT license, and API pricing near the bottom of the market.
  • Claude Opus 5Anthropic's frontier Claude model — deep reasoning, agentic coding, vision, and a 1M-token context, positioned near Fable 5 intelligence at half the cost. A text-output model, not an image, audio, or video generator.
  • GPT-5.6OpenAI's three-tier frontier model family — Sol, Terra, and Luna — with sharper image reading and stronger text-and-interface generation.

← All AI tools · Get started →