// AI NEWS · MODEL RELEASE

MiniMax Releases H3, an Open-Weights Video Model It Says Costs a Third of the Proprietary Leaders

The Shanghai firm's new multimodal model makes 2K clips with native stereo sound, conditions on up to 9 image, 3 video, and 3 audio references, and ships its weights under a license that's free for smaller companies.

2026-07-31 · by Moe Ameen

What happened

MiniMax, the Shanghai AI company founded in 2022 and listed in Hong Kong in January 2026, teased H3 on July 30, 2026 under the tag #MiniMaxH3 and released it on July 31 through its API and the Hailuo platform. The company said it would publish the model's weights "within days," making H3 the open-weights entry in its Hailuo line rather than a closed API-only product.

H3 is a general-purpose multimodal video model: it reads text, images, video, and audio in one unified context and generates a coherent audiovisual result from any mix of them. It produces clips of roughly 5 to 15 seconds at up to 2K (2560×1440) resolution with native stereo sound, reported at 24 fps. Its most distinctive feature is reference conditioning — you can feed up to 9 reference images, 3 video clips, and 3 audio clips at once to lock a style, a character's identity, a motion, or a voice — plus edit existing footage and transfer motion between videos.

The pitch is price and openness. MiniMax says generating 2K video costs less than one-third of mainstream rivals, and less than half at 768p versus mainstream models at 720p, and is targeting commercial work in advertising, e-commerce, product design, games, and film. The weights ship under the MiniMax Community License: free for non-commercial use, and free for commercial use by organizations under roughly US$20 million in annual revenue, with attribution. MiniMax also said H3 is built to run across a broad range of AI hardware, including Chinese-made chips.

On independent benchmarks the picture is mixed. Artificial Analysis rated H3 the strongest model for video editing at launch while placing it behind Google's Gemini Omni Flash in text-to-video and behind both Gemini Omni Flash and ByteDance's Seedance 2.0 in image-to-video — so H3's edge is open weights, multi-reference control, and cost rather than a top raw-quality ranking.

Why it matters for creators

  • Open weights plus a sub-$20M-revenue free commercial tier put a capable 2K, audio-native video model in reach of small creators and studios without an API bill.
  • A claimed one-third the cost of proprietary rivals makes high-volume generation viable — the constraint shifts from "can I afford the clips" to "can I finish and publish them."
  • Conditioning on up to 9 image, 3 video, and 3 audio references means consistent characters, products, and voices across a whole batch, not just within one render.
  • Native footage editing and motion transfer let you revise a clip with instructions instead of regenerating from scratch.
  • It is not a clean benchmark win — Gemini Omni Flash and Seedance 2.0 still lead on parts of text- and image-to-video, so H3 is a value and control play, not a quality king.

How to act on this with Kompozy

You don't have to wait for the weights to act on this. The moment H3 lowers the cost of a controllable 2K clip, the bottleneck moves downstream — to captioning, reframing, staying on-brand, and posting — and that's exactly the tail Kompozy removes. Generate a batch of reference-locked H3 clips, drop each export into Kompozy, and it burns in captions styled to your brand, reframes cleanly to 9:16, 1:1, and 16:9 per destination, and stacks hook text and lower-thirds through HyperFrames so the muted first second reads on the feed.

Then it multiplies and ships. The reference-consistency H3 gives you inside one render, Kompozy extends across your operation: your Persona Brief holds one voice across every caption, and a single H3 clip seeds a full unit — the video plus a brand-exact Carousel, a Quote Graphic, native Text Posts, a Blog Article, and an Email Newsletter. Autopilot then schedules and publishes the whole package across nine social platforms plus blog and email from one queue, each piece clearing a per-post review pipeline first. Cheap, controllable motion from H3; finished, on-brand, everywhere from Kompozy.

Quick takeaways

  • H3 released July 31, 2026; weights to follow "within days."
  • Up to 2K (2560×1440) clips, ~5–15s, native stereo audio, multi-reference conditioning.
  • Free commercial use under the MiniMax Community License for orgs under ~US$20M revenue, with attribution.
  • Claimed at under a third the 2K cost of mainstream rivals; strongest at video editing, behind on parts of text/image-to-video.

Frequently asked questions

What is MiniMax H3?

H3 (marketed as Hailuo 3.0) is an open-weights, multimodal AI video model from MiniMax, released July 31, 2026. It reads text, images, video, and audio in one context and generates clips of roughly 5–15 seconds at up to 2K resolution with native stereo audio, and it can edit existing footage and condition on up to 9 image, 3 video, and 3 audio references.

Is MiniMax H3 free to use commercially?

The weights ship under the MiniMax Community License, which allows free non-commercial use and free commercial use for organizations under roughly US$20 million in annual revenue, with attribution required. Larger companies fall outside that free tier. It is more permissive than closed rivals but not a standard open-source license.

How does H3 compare to Seedance 2.0 and Gemini Omni Flash?

At launch, Artificial Analysis rated H3 the strongest model for video editing but placed it behind Google's Gemini Omni Flash in text-to-video and behind both Gemini Omni Flash and ByteDance's Seedance 2.0 in image-to-video. H3 competes on open weights, multi-reference control, and a claimed price of under a third of mainstream rivals.

Related news

← All AI news · Get started →