MiniMax's open-weights, multimodal video model — 2K clips with native stereo audio, conditioned on up to 9 image, 3 video, and 3 audio references, priced to undercut the proprietary leaders.
Last verified · 2026-07-31 · by Moe Ameen
MiniMax H3 (marketed as Hailuo 3.0) is a general-purpose multimodal video generation model from MiniMax, the Shanghai AI company founded in 2022 and listed in Hong Kong in January 2026. MiniMax teased it on July 30, 2026 under the tag #MiniMaxH3 and released it on July 31, 2026 through its API and the Hailuo platform, with the model weights slated to follow "within days." What sets it apart from the rest of MiniMax's Hailuo line is that H3 ships as an open-weights model, not a closed API-only product.
The technical pitch is one unified context that reads text, images, video, and audio together and generates a coherent audiovisual result from any mix of them. H3 outputs clips of roughly 5 to 15 seconds at up to 2K (2560×1440) resolution with native stereo sound, reported at 24 fps. Its most distinctive control is reference conditioning: you can feed up to 9 reference images, 3 video clips, and 3 audio clips at once to lock a style, a character's identity, a motion, or a voice — plus edit existing footage and transfer motion from one video onto another.
MiniMax is positioning H3 as a low-cost challenger to the proprietary leaders. The company says generating 2K video costs less than one-third of mainstream rivals, and less than half at 768p versus mainstream models at 720p. The weights are released under the MiniMax Community License: free for non-commercial use, and free for commercial use by organizations under roughly US$20 million in annual revenue, with attribution required. MiniMax frames H3 for commercial work — advertising, e-commerce, product design, games, and film — and notes it is built to run across a broad range of AI hardware, including Chinese-made chips.
The honest boundary: on independent benchmarks H3 is not a clean sweep. Artificial Analysis rated it the strongest model for video editing while placing it behind Google's Gemini Omni Flash in text-to-video and behind both Gemini Omni Flash and ByteDance's Seedance 2.0 in image-to-video. And like every generator, H3 stops at the clip — it writes no caption in your voice, sizes nothing for a specific feed, builds no carousel or blog, and posts to no platform. Treat resolution ceilings, the fps figure, and pricing as an early snapshot and confirm the current spec on MiniMax's own site before quoting it.
H3's standout is control at low cost: because you can bind a character's face, a product's look, or a specific voice with up to nine image, three video, and three audio references, you can crank out clip after clip that all feel like the same campaign — and do it for a fraction of what the closed models charge. That makes H3 a volume engine. But a locked, consistent 2K clip is still a raw, silent-until-you-caption asset, not a post — and generating ten of them just multiplies the finishing work you still owe each one. Kompozy is the layer that absorbs that work. Drop an H3 export in and it burns in captions styled to your brand, reframes the clip cleanly to 9:16, 1:1, and 16:9 for each destination, and stacks hook text and lower-thirds through HyperFrames so the muted first second actually reads on the feed.
The reference-consistency H3 gives you inside a single render, Kompozy extends across your whole content operation. Your Persona Brief and banned-word filters hold one voice across every caption; the AI Influencer persona pool with Gemini face-lock keeps a recurring on-brand identity across posts, not just across frames of one clip. And a single H3 clip seeds a full unit: the video for short-form feeds plus a brand-exact Carousel, a Quote Graphic, native Text Posts, a Blog Article, and an Email Newsletter — then Kompozy schedules and publishes the whole package across nine social platforms plus blog and email from one queue, on Autopilot with a per-post review pipeline. Generate cheap, controllable motion in H3; make it finished, on-brand, multiplied, and everywhere in Kompozy.
H3 (marketed as Hailuo 3.0) is an open-weights, multimodal AI video model from MiniMax, released July 31, 2026. It reads text, images, video, and audio in one context and generates clips of roughly 5–15 seconds at up to 2K resolution with native stereo audio. It supports editing existing footage and conditioning on up to 9 image, 3 video, and 3 audio references.
The weights are open under the MiniMax Community License, which is more permissive than closed rivals but not a standard open-source license. Non-commercial use is free, and commercial use is free for organizations under roughly US$20 million in annual revenue with attribution required. Larger companies fall outside that free tier.
MiniMax positions H3 as a low-cost model, claiming 2K generation costs less than one-third of mainstream rivals and less than half at 768p versus mainstream 720p output. It is available via the Hailuo platform and API. Because it launched days ago, confirm current per-video or credit pricing on MiniMax's own page before budgeting.
On Artificial Analysis at launch, H3 rated strongest for video editing but trailed Google's Gemini Omni Flash in text-to-video and sat behind both Gemini Omni Flash and ByteDance's Seedance 2.0 in image-to-video. Its edge is open weights, multi-reference control, and price rather than a top raw-quality ranking.
No. H3 generates the clip but does not caption in your voice, brand it, size it per platform, schedule, or publish. To turn an H3 clip into finished, on-brand posts across nine platforms plus blog and email, use a content engine like Kompozy.