On August 3, 2026, ComfyUI shipped native support for MiniMax's open-weights H3 video model the same day the weights landed — with optimizations that cut the smallest variant's memory footprint by 66% so a 2K model can run locally on a card like an RTX 3060.
2026-08-03 · by Moe Ameen
On August 3, 2026, ComfyUI — the open-source, node-based interface for running AI models — announced day-0 native support for MiniMax H3, MiniMax's open-weights multimodal video model, timed to the same day the model's weights were released. "Day-0" is the point: rather than waiting weeks for community integration, the workflow templates, repackaged weights, and node support shipped alongside the model itself, so ComfyUI users could generate with H3 immediately.
Inside a ComfyUI graph, the integration exposes H3's full mode set — text-to-video from a prompt, image-to-video, first-and-last-frame control (set the opening frame, closing frame, or both), and reference-to-video (carry a subject, motion, or voice through the clip). Every clip is generated with native stereo audio in the same pass rather than dubbed afterward, at up to 2K resolution and durations of roughly 5–15 seconds. The same five tasks that usually need five different tools collapse into one model driven from one workflow.
The engineering story is running a very large model on hardware people actually own. Comfy Org reports that H3's modulation weights — about 40% of its parameters — could be pruned and replaced with a functionally equivalent lookup table, and combined with int8 quantization and custom kernels this cut the smallest variants' memory footprint from 123.6 GB in full precision to 42.5 GB, a 66% reduction. With dynamic VRAM offloading, that puts a 2K, native-audio video model within reach of a consumer GPU as modest as an RTX 3060. Repackaged weights in bf16, INT8, pruned, and NVFP4 formats, plus official workflow templates, are published on Hugging Face. Treat the exact resolution, duration, and VRAM figures as an early snapshot and confirm current numbers in ComfyUI's documentation.
The useful way to read this launch: the thing that was scarce yesterday — generating a 2K, native-audio clip — is now free and unlimited on a card you might already own, and the thing that's still scarce is turning a folder of raw renders into a published week of on-brand content across every feed. When generation stops being the ceiling, distribution becomes it, and that's exactly the work [Kompozy](/) is built to absorb. Point it at your ComfyUI exports and it does the part a node graph can't: it burns in captions styled to your brand so H3's clips read in muted autoplay feeds, reframes each render cleanly to 9:16, 1:1, and 16:9 per destination, and layers hook text through HyperFrames so the first second stops the scroll.
Then one H3 render becomes a full unit instead of a single post. Kompozy fans it into [18 output formats](/glossary/output-buckets) — [Clipped Shorts](/glossary/clipped-short) for the video feeds, a brand-exact Carousel, Quote Graphics, native Text Posts, a Blog Article, and an Email Newsletter — every piece held to one voice by your Persona Brief and banned-word filters. [Autopilot](/glossary/autopilot) and a per-post review pipeline then schedule and publish the whole set across the eight primary social platforms plus blog and email from one queue. Because Kompozy is source-agnostic, it doesn't care that the footage came from a local open-weights model rather than a paid API — generate at zero marginal cost in ComfyUI, and let Kompozy remove the finishing-and-posting work that now stands between cheap video and an actual audience.
It means ComfyUI added native support for MiniMax H3 on the same day the model's open weights were released — August 3, 2026 — so users could generate with H3 immediately instead of waiting for community integration. The launch included workflow templates, repackaged weights on Hugging Face, and node support for all of H3's modes inside a ComfyUI graph.
Within limits, yes. Comfy Org reports pruning H3's modulation weights into a lookup table plus int8 quantization and custom kernels, cutting the smallest variant's memory footprint from 123.6 GB to 42.5 GB — a 66% reduction — and using dynamic VRAM offloading so a 2K model can run on a GPU as modest as an RTX 3060. Higher-quality variants want more VRAM.
H3 supports text-to-video from a prompt, image-to-video, first-and-last-frame control, and reference-to-video that carries a subject, motion, or voice through the clip. Every clip is generated with native stereo audio in the same pass, at up to 2K resolution and roughly 5–15 second durations.
ComfyUI generates the file but does not caption in your brand voice, reframe per platform, build carousels or blogs, or publish. Bring the render into a content engine like Kompozy, which burns in branded captions, reframes to 9:16, 1:1, and 16:9, fans it into a carousel, quote graphics, a blog, and a newsletter, and schedules and publishes across the eight social platforms plus blog and email.