The July 31 public-beta release keeps the same 284B-parameter architecture, size, and rock-bottom price as the preview but re-post-trains the model for stronger agent and coding performance — and adds native Responses-API support.
2026-07-31 · by Moe Ameen
DeepSeek marked the official release of DeepSeek-V4-Flash in its API changelog on July 31, 2026, moving the model out of preview and into public beta under the build name DeepSeek-V4-Flash-0731. The company says the release keeps the same architecture and model size as the earlier V4-Flash preview and was re-post-trained rather than rebuilt — same name, same price, notably different behavior on agent tasks. DeepSeek reports the new build's agent-benchmark results now exceed those of the earlier V4-Pro preview, citing scores on Terminal Bench, an NL2Repo coding benchmark, and a cybersecurity agent benchmark.
The economics are the story. DeepSeek-V4-Flash is a mixture-of-experts model with about 284 billion total parameters and roughly 13 billion active per token, paired with a 1-million-token context window. On DeepSeek's first-party API it runs near the bottom of the market — about $0.14 per million input tokens on a cache miss and $0.28 per million output tokens, with cached input far cheaper — roughly a third of the output price of the larger DeepSeek-V4-Pro tier. The weights are open under the MIT license and published on Hugging Face, so teams can self-host instead of paying per token.
The release also leans toward agent and coding workflows. DeepSeek says the model natively supports the Responses API format and is adapted for coding-agent tooling such as Codex, and it runs in both a thinking (reasoning) and a non-thinking mode. DeepSeek's legacy `deepseek-chat` and `deepseek-reasoner` endpoints now route into V4-Flash's non-thinking and thinking modes. The V4 family first appeared as a preview on April 24, 2026; this is the point at which the Flash tier became an official, if still beta, release.
When drafting gets this cheap, the constraint in a content operation stops being "can I write it" and becomes "can I produce and ship it everywhere." DeepSeek-V4-Flash makes the words nearly free; it still renders no image, cuts no clip, records no avatar, and publishes to nothing. That is exactly the gap Kompozy fills. Use DeepSeek to draft a script or batch out fifty caption variations for pennies, then bring that text into Kompozy — a full AI content generation and multi-platform publishing engine — and it produces the actual assets: a Persona Shorts avatar video reading your script, a brand-exact Carousel, Quote Graphics, a Blog Article, and an Email Newsletter, all held to your voice by the Persona Brief.
From there the volume you can now afford to write becomes volume you can actually publish. Kompozy's Autopilot and per-post review pipeline reframe each asset to 9:16, 1:1, and 16:9 and schedule and publish across nine destinations — the eight primary social platforms plus blog and email. Cheap text without a production-and-distribution engine just gives you more drafts sitting in a doc; Kompozy is where a $0.14-per-million-token draft turns into a week of finished posts on every platform your audience uses. (Kompozy's own copy generation runs on Claude and OpenAI models, so DeepSeek is your upstream drafting choice and Kompozy is the downstream engine.)
DeepSeek moved V4-Flash from preview to an official public beta (build DeepSeek-V4-Flash-0731). It kept the same architecture, model size, and price but was re-post-trained for stronger agent and coding performance — DeepSeek says its agent-benchmark scores now exceed the earlier V4-Pro preview — and added native Responses-API support and Codex adaptation.
On DeepSeek's first-party API it is roughly $0.14 per million input tokens (cache miss) and $0.28 per million output tokens, with cached input far cheaper — about a third of the larger V4-Pro tier's output rate. The weights are also open under the MIT license, so teams can self-host to avoid per-token fees.
The weights are open under the MIT license and published on Hugging Face, so you can download and self-host the model. It is a mixture-of-experts model with about 284 billion total parameters, roughly 13 billion active per token, and a 1-million-token context window.
No. It is a text-and-reasoning model — it writes and analyzes text but produces no images, video, or audio and publishes nothing. To turn its drafts into finished posts, pair it with a content engine like Kompozy, which generates video, carousels, and images and publishes across the eight social platforms plus blog and email.