The FLUX image lab's first video model makes 20-second clips with native audio in one pass — and even drives factory robots — but ships in a gated early-access program, not general availability.
2026-07-28 · by Moe Ameen
Black Forest Labs, the Freiburg, Germany lab behind the FLUX image models, announced FLUX 3 on July 23, 2026 — its first multimodal frontier model and its first move into video. Where FLUX.1 and FLUX.2 were image generators, FLUX 3 is trained jointly on image, video, audio, and robot-action prediction inside one architecture, built on a method the lab calls Self-Flow for aligning multimodal generation and understanding. Black Forest Labs frames it as a step toward "real-world visual intelligence" — models that perceive, predict, and act rather than just draw.
The headline capability is video with sound. FLUX 3 Video generates clips up to 20 seconds long in a single generation with optional native audio, produced from the same flow-matching framework that renders the frames rather than dubbed on afterward. It supports text-to-video, image-to-video, and video-to-video, plus keyframe transitions, multilingual dialogue, on-screen typography, and chaining clips into longer sequences. The lab says it is especially strong at human facial expressions and at matching sound to on-screen physical events. On the physical-AI side, an action model the lab calls FLUX-mimic drives robots — Black Forest Labs said it is already in production testing at Audi — from the same shared architecture.
FLUX 3 is a family that ships in stages, and most of it is gated at launch. FLUX 3 Video entered an "Early Access" program on announcement day: anyone can apply, but the lab must approve each applicant. The action component is held tighter — available only through selected research and commercial partners, beginning with mimic robotics, not a public application. FLUX 3 Image was slated to follow in the weeks after, and the lab said API access, private model weights, and an open-weight version called FLUX 3 Dev would come later in the year. Black Forest Labs did not publish parameter counts or pricing, so any circulating figure is unverified. It did share its own human-preference numbers on 10-second 720p clips — evaluators reportedly preferred FLUX 3 Video over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%, with narrower margins against Kling v3 Pro and Seedance 2.0 — which are the lab's evaluations, so read them as directional rather than independent.
The honest first reaction to a launch like this is that you probably cannot use it yet. FLUX 3 Video is gated behind an approval queue, FLUX 3 Image lands later, and there is no public pricing — so the practical question isn't "how do I build my pipeline on FLUX 3," it's "how do I stay ready to use the best video model I can actually get into, without rebuilding my whole workflow each time a new one wins the leaderboard." That is exactly the layer Kompozy occupies: it is model-agnostic on the render side and owns everything after the clip.
Whichever frontier model you get access to — FLUX 3 when your invite lands, or Runway, Kling, or Seedance while you wait — the export is the same shape: an MP4 that no feed can publish as-is. Bring it into Kompozy and it burns in branded captions for muted autoplay, reframes to 9:16, 1:1, and 16:9 per destination, stacks a hook overlay through HyperFrames so the opening second lands, and fans the concept into a Carousel, a Quote Graphic, a Blog Article, and platform-native captions in your Persona Brief voice — then schedules and publishes the set across the eight social platforms plus blog and email from one review pipeline. And on the days no invite has come through, Kompozy still generates net-new video FLUX 3 can't touch: HeyGen-powered Persona Shorts and avatar video that hold one face and voice across every post. The model you render in will keep changing; the generation-and-publishing engine around it doesn't have to.
FLUX 3 is Black Forest Labs' first multimodal frontier model, announced July 23, 2026. Unlike the earlier image-only FLUX.1 and FLUX.2, it is trained jointly on image, video, audio, and robot-action prediction in one architecture. Its headline feature is generating video up to 20 seconds long with optional native audio, and it includes an action model, FLUX-mimic, for robotics.
Only partly. At launch, FLUX 3 Video was in an approval-based early-access program — you can apply, but Black Forest Labs must accept you. The action component was limited to selected partners (starting with mimic robotics), not open to public application. FLUX 3 Image was expected in the following weeks, and API access, private weights, and an open-weight FLUX 3 Dev were planned for later in the year. Check the lab's site for current availability.
Black Forest Labs published its own human-preference results on 10-second 720p clips, reporting evaluators preferred FLUX 3 Video over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%, with narrower margins against Kling v3 Pro and Seedance 2.0. These are the lab's own tests, so treat them as directional rather than independent benchmarks.
FLUX 3 renders the video and audio but doesn't publish anything. Bring the export into a content engine like Kompozy to add branded captions, reframe per platform, stack a hook overlay, fan it into a carousel, quote card, and blog with captions in your voice, then schedule and publish across TikTok, Reels, YouTube Shorts, X, LinkedIn, and more from one queue.