In late July and early August 2026, three of China's biggest AI labs pushed major launches within days of each other — open-weights video from MiniMax, a longer-clip video model from ByteDance, and a rock-bottom-priced text model from DeepSeek — a concentrated debut domestic press flagged as a milestone launch stretch for Chinese models.
2026-08-04 · by Moe Ameen
Across late July and early August 2026, three of China's largest AI labs shipped major model launches inside the same short window, and the domestic press framed it as a concentrated debut of China's big models — a milestone launch stretch for the country's labs. Two of the launches — MiniMax's H3 and ByteDance's Seedance 2.5 — were video models pointed at the same market; the third, DeepSeek's V4 line, was a text-and-reasoning model competing on price. The common thread is direction: open weights and aggressive pricing aimed squarely at the proprietary Western leaders.
MiniMax released H3 through its API and the Hailuo platform on July 31 and published the weights to Hugging Face on August 3. H3 is a multimodal video model of roughly 33 billion parameters that reads text, images, video, and audio in one context and generates clips of about 4 to 15 seconds at up to 2K resolution with native stereo sound, conditioning on up to 9 image, 3 video, and 3 audio references. On the independent Artificial Analysis leaderboard it placed first in video editing — the first time an openly-released model has led a mainstream video category outright — and MiniMax says it costs under a third of proprietary rivals at 2K. One caveat matters for self-hosters: the open weights cap around 768p because the 2K upscaler and a context layer stay proprietary.
ByteDance's Seedance 2.5, unveiled at its Volcano Engine FORCE conference in June and rolling out publicly through the summer, took the opposite bet: it stayed a closed, hosted model. Its headline is a continuous 30-second clip generated in a single pass without stitching, with native 4K output and up to 50 reference inputs per request. On the text side, DeepSeek moved its V4-Flash tier to an official public beta on July 31 (build DeepSeek-V4-Flash-0731), re-post-trained for stronger agent and coding performance. V4-Flash is a mixture-of-experts model with about 284 billion total and 13 billion active parameters, a 1-million-token context window, and MIT-licensed open weights — part of the two-tier V4 family that first previewed on April 24, 2026 alongside the larger V4-Pro. On DeepSeek's API it runs near the bottom of the market, roughly $0.14 per million input tokens and $0.28 per million output.
The takeaway is not any single spec but the cadence and the strategy behind it. Leaderboards for video and text are changing hands week to week, and the models doing it increasingly ship free-to-download weights and sub-market pricing rather than a closed API. Treat individual resolution, length, and price figures as fast-moving snapshots — the durable fact is that capable video, image, and text generation is getting cheaper and more open from several vendors at once.
Start with the move you can make this week, because the wave itself is the story your audience is watching: publish a credible take before the recap crest passes. Drop your point of view — "China's labs just shipped open, cheap video and text in one week, here's what it changes" — into [Kompozy](/) as a source, and it spins that one angle into a Blog Article explaining the open-vs-closed split, a brand-exact Carousel breaking down H3 versus Seedance 2.5, Quote Graphics pulling the sharpest line, captioned short-form, and native Text Posts in your voice through the Persona Brief, then schedules and publishes the set across the eight primary social platforms plus blog and email from one queue. Being early on a launch wave like this is how a single opinion becomes a week of content.
The deeper fit is what a multimodal wave does to your inputs. You can now draft with DeepSeek V4 for pennies, generate clips with a self-hosted H3 or a hosted Seedance 2.5, and shoot images from any of a dozen cheap models — and you're left with a folder of parts from four vendors, no shared brand, and nothing published. Kompozy is the assembly layer that unifies them: it holds one voice across every asset through the Persona Brief, renders the formats no video or text model makes — [Persona Shorts](/glossary/persona-shorts) and avatar video, brand-exact carousels and quote cards through HyperFrames — and fans one idea into [18 output formats](/glossary/output-buckets) that [Autopilot](/glossary/autopilot) publishes across nine destinations behind a per-post review gate. Because Kompozy is source-agnostic, it doesn't care whether this week's best model is [MiniMax H3](/ai-tools/minimax-h3-video-model), [DeepSeek V4](/ai-tools/deepseek-v4), or whatever tops the board next week — you swap the cheap commodity input, and the finished, on-brand, everywhere output stays exactly the same.
The three headline launches were MiniMax H3, an open-weights multimodal video model that makes ~2K clips with native audio (API on July 31, weights on August 3, 2026); ByteDance's Seedance 2.5, a closed video model that generates a continuous 30-second clip in one pass with native 4K; and DeepSeek's V4 line, whose low-cost, MIT-licensed V4-Flash tier moved to official public beta on July 31, 2026. All three came from major Chinese labs within days of each other.
They launched into the same market on opposite strategies. MiniMax open-sourced H3's weights so anyone under roughly US$20M in revenue can self-host and use it commercially, and it topped the video-editing category on Artificial Analysis — but the open weights cap around 768p. Seedance 2.5 stayed closed and hosted, and its edge is length and resolution: a continuous 30-second single-pass clip in native 4K with up to 50 reference inputs.
The V4 weights are open under the MIT license and published on Hugging Face, so teams can self-host. V4 is a two-tier family: the larger V4-Pro and the smaller, cheaper V4-Flash (~284B total, ~13B active, 1M-token context), which moved to an official public beta on July 31, 2026. On DeepSeek's API, V4-Flash runs near the bottom of the market at roughly $0.14 per million input and $0.28 per million output tokens.
Treat the model as a swappable input, not the product. Draft text with a cheap model like DeepSeek V4, generate clips with H3 or Seedance, then run everything through a source-agnostic engine like Kompozy that captions, reframes, and brands the output, fans one idea into carousels, blogs, avatar video, and newsletters, and publishes across the eight social platforms plus blog and email — so when the best model changes next week, only the input swaps.