With its weights on Hugging Face as of August 3, 2026, H3 took the No. 1 spot in video editing on Artificial Analysis — the first time an openly-released model has led a category previously owned by closed systems.
2026-08-04 · by Moe Ameen
On August 3, 2026, MiniMax published the weights for its H3 video model to Hugging Face under the MiniMax H3 Community License, days after the July 31 API launch. That release did something no open model had managed before: it put an openly-downloadable model at the top of a mainstream AI video ranking. On Artificial Analysis — the independent benchmark that scores video generators through blind, head-to-head comparisons — H3 placed first in Video Editing, second in Text-to-Video, and third in Image-to-Video. Video-model leaderboards have been the domain of closed, hosted systems from Google, ByteDance, Kuaishou, and Alibaba; H3 is the first with published weights to lead one outright. MiniMax called it the state-of-the-art open video model on both the Artificial Analysis and Arena boards.
H3 is a roughly 33-billion-parameter multimodal model. It reads text, images, video, and audio in one context and generates clips of about 4 to 15 seconds at 24 fps with native stereo sound, and its standout is editing and revising existing footage rather than only generating from a prompt — which is the exact category it now tops. The Community License permits free non-commercial use and free commercial use for organizations under roughly US$20 million in annual revenue, with attribution.
One honest caveat matters for anyone planning to self-host: the open weights are not the full hosted model. The 2K-resolution upscaling module and H3's context-translation layer stay proprietary, so a locally-run H3 maxes out around 768p — the 2K clips and some of the reference-conditioning polish remain on MiniMax's API. The benchmark crown is real, but "open" here means a very capable base model, not a byte-for-byte copy of the paid product.
The milestone lands in a crowded week: ByteDance shipped Seedance 2.5 the same day with longer 30-second, audio-native clips, and the top of the video board keeps changing hands. What's new is the license under the leader's name. For the first time, the model sitting at No. 1 in a category is one any small studio can download and run.
The headline everyone will take from this is "the best video model is now free." The more useful read: when the leader open-sources its weights, clip quality becomes table stakes, and the only thing left to compete on is everything a leaderboard doesn't measure — brand consistency, the formats a video model can't make, and getting the thing published. That downstream layer is precisely what [Kompozy](/) is. A benchmark measures one 8-second clip in isolation; Kompozy takes that clip and renders brand-exact carousels, quote cards, and hook text through HyperFrames, generates the [17 other formats](/glossary/output-buckets) H3 will never produce — a Blog Article, an Email Newsletter, [Persona Shorts](/glossary/persona-shorts) and avatar video — and pushes the whole set live across the eight social platforms plus blog and email behind a per-post review gate. Open weights hand every creator the same footage; what a benchmark can't rank is the operation that turns it into finished, on-brand content everywhere.
H3's crown is specifically for editing existing footage, and that plays directly into how Kompozy already works. Feed it a recut H3 clip — or a raw one — and Clipped Shorts and Marketing Shorts trim it to vertical cuts, burn in captions styled to your Persona Brief, and reframe to each destination's ratio, then [Autopilot](/glossary/autopilot) schedules and publishes the batch from one queue. Because the open weights cap around 768p locally while the sharpest 2K output stays on MiniMax's API, staying source-agnostic matters more than usual: Kompozy doesn't care whether your clip came from a self-hosted H3, the paid API, or the next model that tops the board next week — it finishes and ships it the same way. When clip quality is commoditized, the workflow is the product. (See the [MiniMax H3 tool breakdown](/ai-tools/minimax-h3-video-model) for how it fits a content stack.)
Yes. With its weights published to Hugging Face on August 3, 2026, H3 took the No. 1 spot in the Video Editing category on Artificial Analysis's independent leaderboard — the first time an openly-released model has led a mainstream AI video ranking outright, in a space previously dominated by closed systems from Google, ByteDance, Kuaishou, and Alibaba. It also ranks 2nd in text-to-video and 3rd in image-to-video.
Not exactly. The open weights are a very capable base model, but H3's 2K-resolution upscaling module and its context-translation layer remain proprietary, so a self-hosted H3 maxes out around 768p. The full 2K output and some reference-conditioning polish stay on MiniMax's hosted API. Treat "open and free" and "the exact clip on the leaderboard" as related but not identical.
Its top-ranked category is video editing — revising, re-cutting, and transforming existing footage rather than only generating from a text prompt. It is a roughly 33-billion-parameter multimodal model that reads text, images, video, and audio together and outputs about 4–15 second clips at 24 fps with native stereo audio.
A leaderboard scores the raw clip; it does not caption, reframe, brand, or publish it. An engine like Kompozy takes an H3 export, burns in on-brand captions, reframes to 9:16, 1:1, and 16:9, and fans the idea into carousels, quote graphics, a blog, and a newsletter, then schedules and publishes across the eight social platforms plus blog and email — the work the model itself never touches.