Black Forest Labs' multimodal foundation model — one system trained jointly on image, video, and audio, plus a robotics action head.
Last verified · 2026-07-23 · by Moe Ameen
Flux 3 (styled FLUX 3) is Black Forest Labs' multimodal foundation model, announced on July 23, 2026. Where the company's earlier FLUX line was known for image generation, Flux 3 is trained jointly on images, video, and audio inside a single unified architecture — Black Forest Labs frames it as a step toward "real-world visual intelligence": models that perceive, predict, and act. It is built on a method the lab calls Self-Flow, its approach for aligning multimodal generation and understanding within a single underlying architecture.
The headline capability is video. Flux 3 Video generates clips up to 20 seconds long in a single generation, with optional native audio — a first for Black Forest Labs, which had not shipped a video model before. It supports text-to-video, image-to-video, and video-to-video, plus keyframe-to-video transitions, multilingual dialogue, on-screen typography, and chaining clips into longer sequences. The lab says Flux 3 Video is especially strong at capturing human facial expressions and matching sound to on-screen physical events.
Flux 3 is a family, and not all of it is out at once. At announcement, Flux 3 Video (with optional audio) and an action component were available in early access; Flux 3 Image was slated to follow "in the coming weeks"; and the lab said it plans to release API access, private model weights, and an open-weight version called Flux 3 Dev later in the year. The action side is Flux-mimic, a video-action model for robotics that Black Forest Labs said was already in production testing at Audi, initially available through select partners.
Black Forest Labs did not publish parameter counts or pricing at announcement, so treat any circulating figure as unverified. What it did share is a set of human-preference comparisons: on 10-second 720p clips, evaluators are reported to have preferred Flux 3 Video over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%, with narrower margins against models like Kling v3 Pro and Seedance 2.0. Those are the lab's own evaluations, so read them as a directional signal rather than an independent benchmark.
Flux 3's real unlock for creators is native audio inside a 20-second clip — that's long enough, and complete enough, to stand as a finished short rather than a silent 5-second fragment you have to pad and score by hand. But a rendered MP4 with baked-in audio is still not a post. It has no captions for the large share of feeds that autoplay muted, no per-platform aspect ratio, no hook text on the opening frame, and no surrounding copy. Kompozy is the layer that closes that gap: drop a Flux 3 export in and it burns in branded, on-style captions, reframes the clip to 9:16, 1:1, and 16:9 per destination, and stacks a hook overlay through HyperFrames so the first second lands even before the audio kicks in — then it schedules and publishes the same short across Instagram, TikTok, YouTube, LinkedIn, Facebook, X, Pinterest, and Threads from one queue.
Because a Flux 3 clip carries dialogue, typography, and multiple beats, it also seeds a whole content unit rather than a single upload. Feed the concept into Kompozy and it fans the idea into a Carousel breaking down each beat, a Quote Graphic pulled from the dialogue, a Blog Article, and platform-native captions written in your voice through the Persona Brief — plus net-new formats Flux 3 doesn't touch, like HeyGen-powered Persona Shorts that hold one face and voice across every post. One Flux 3 generation becomes a week of cross-platform content. Flux 3 owns the render and the sound; Kompozy owns the captions, the formats, the schedule, and the publish.
Flux 3 is Black Forest Labs' multimodal foundation model, announced on July 23, 2026. Unlike the earlier image-focused FLUX line, it is trained jointly on images, video, and audio in one architecture, and adds a robotics action model called Flux-mimic. Its headline feature is generating video up to 20 seconds long with optional native audio.
Yes. Flux 3 Video generates clips up to 20 seconds in a single generation with optional native audio — the first video model from Black Forest Labs. The lab describes Flux 3 Video as especially strong at capturing human facial expressions and matching sound to on-screen physical events.
Partly. At announcement, Flux 3 Video (with optional audio) and an action component were in early access, while Flux 3 Image was expected in the following weeks. Black Forest Labs said it plans to release API access, private weights, and an open-weight Flux 3 Dev version later in the year, so confirm current availability on its site.
Black Forest Labs published human-preference results on 10-second 720p clips, reporting that evaluators preferred Flux 3 Video over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%, with narrower margins against Kling v3 Pro and Seedance 2.0. These are the lab's own evaluations, so treat them as directional rather than independent.
Flux 3 renders the clip but doesn't publish it. Bring the export into Kompozy to add branded captions, reframe per platform, stack a hook overlay, fan it into a carousel and captions in your voice, and schedule and publish across TikTok, Reels, YouTube Shorts, X, LinkedIn, and more from one queue.