// AI NEWS · MODEL RELEASE

Black Forest Labs Unveils FLUX 3, a Multimodal Model That Generates Video, Images, and Audio Together

The FLUX image lab's first video model makes 20-second clips with native audio in one pass — and even drives factory robots — but ships in a gated early-access program, not general availability.

KompozyTurn one idea into a week of content — across every platform, published for you.
Get Started →

2026-07-28 · by Moe Ameen

What happened

Black Forest Labs, the Freiburg, Germany lab behind the FLUX image models, announced FLUX 3 on July 23, 2026 — its first multimodal frontier model and its first move into video. Where FLUX.1 and FLUX.2 were image generators, FLUX 3 is trained jointly on image, video, audio, and robot-action prediction inside one architecture, built on a method the lab calls Self-Flow for aligning multimodal generation and understanding. Black Forest Labs frames it as a step toward "real-world visual intelligence" — models that perceive, predict, and act rather than just draw.

The headline capability is video with sound. FLUX 3 Video generates clips up to 20 seconds long in a single generation with optional native audio, produced from the same flow-matching framework that renders the frames rather than dubbed on afterward. It supports text-to-video, image-to-video, and video-to-video, plus keyframe transitions, multilingual dialogue, on-screen typography, and chaining clips into longer sequences. The lab says it is especially strong at human facial expressions and at matching sound to on-screen physical events. On the physical-AI side, an action model the lab calls FLUX-mimic drives robots — Black Forest Labs said it is already in production testing at Audi — from the same shared architecture.

FLUX 3 is a family that ships in stages, and most of it is gated at launch. FLUX 3 Video entered an "Early Access" program on announcement day: anyone can apply, but the lab must approve each applicant. The action component is held tighter — available only through selected research and commercial partners, beginning with mimic robotics, not a public application. FLUX 3 Image was slated to follow in the weeks after, and the lab said API access, private model weights, and an open-weight version called FLUX 3 Dev would come later in the year. Black Forest Labs did not publish parameter counts or pricing, so any circulating figure is unverified. It did share its own human-preference numbers on 10-second 720p clips — evaluators reportedly preferred FLUX 3 Video over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%, with narrower margins against Kling v3 Pro and Seedance 2.0 — which are the lab's evaluations, so read them as directional rather than independent.

Why it matters for creators

  • The best-known open-image lab now makes video with native audio. A 20-second clip with synced sound is long and complete enough to stand as a finished short, not a silent five-second fragment you have to score by hand.
  • It ships gated, not open. Early access requires approval, image generation and open weights come later, and there is no public price — so for most creators this is a preview to plan around, not a tool to use this week.
  • Frontier video is consolidating into a crowded field — Runway, Kling, Seedance, Hailuo, and now FLUX 3 — where no single model wins every shot. Betting your workflow on one model is riskier than staying able to use whichever one you can access.
  • A raw clip is still not a post. FLUX 3 renders video and sound; it writes no captions, picks no aspect ratio, adds no hook, and cannot schedule or publish anything.
  • The Audi robotics angle is a signal, not a creator feature. It shows where the lab is aiming, but the content-relevant piece is FLUX 3 Video, and that is the part behind the access gate.

How to act on this with Kompozy

The honest first reaction to a launch like this is that you probably cannot use it yet. FLUX 3 Video is gated behind an approval queue, FLUX 3 Image lands later, and there is no public pricing — so the practical question isn't "how do I build my pipeline on FLUX 3," it's "how do I stay ready to use the best video model I can actually get into, without rebuilding my whole workflow each time a new one wins the leaderboard." That is exactly the layer Kompozy occupies: it is model-agnostic on the render side and owns everything after the clip.

Whichever frontier model you get access to — FLUX 3 when your invite lands, or Runway, Kling, or Seedance while you wait — the export is the same shape: an MP4 that no feed can publish as-is. Bring it into Kompozy and it burns in branded captions for muted autoplay, reframes to 9:16, 1:1, and 16:9 per destination, stacks a hook overlay through HyperFrames so the opening second lands, and fans the concept into a Carousel, a Quote Graphic, a Blog Article, and platform-native captions in your Persona Brief voice — then schedules and publishes the set across the eight social platforms plus blog and email from one review pipeline. And on the days no invite has come through, Kompozy still generates net-new video FLUX 3 can't touch: HeyGen-powered Persona Shorts and avatar video that hold one face and voice across every post. The model you render in will keep changing; the generation-and-publishing engine around it doesn't have to.

Quick takeaways

  • Black Forest Labs announced FLUX 3 on July 23, 2026 — its first multimodal model and its first video model, trained jointly on image, video, audio, and robot-action prediction.
  • FLUX 3 Video makes clips up to 20 seconds with optional native audio in one generation; an action model called FLUX-mimic drives robots and is in production testing at Audi.
  • It launched gated: FLUX 3 Video in approval-based early access, the action component limited to selected partners (starting with mimic robotics), FLUX 3 Image following later, with API, private weights, and an open-weight FLUX 3 Dev planned for later in the year.
  • No parameter counts or pricing were published; the lab's own preference tests put FLUX 3 Video ahead of Luma Ray 3.2 and Runway Gen-4.5, so treat them as directional.

Frequently asked questions

What is FLUX 3?

FLUX 3 is Black Forest Labs' first multimodal frontier model, announced July 23, 2026. Unlike the earlier image-only FLUX.1 and FLUX.2, it is trained jointly on image, video, audio, and robot-action prediction in one architecture. Its headline feature is generating video up to 20 seconds long with optional native audio, and it includes an action model, FLUX-mimic, for robotics.

Can I use FLUX 3 right now?

Only partly. At launch, FLUX 3 Video was in an approval-based early-access program — you can apply, but Black Forest Labs must accept you. The action component was limited to selected partners (starting with mimic robotics), not open to public application. FLUX 3 Image was expected in the following weeks, and API access, private weights, and an open-weight FLUX 3 Dev were planned for later in the year. Check the lab's site for current availability.

How does FLUX 3 compare to Runway, Kling, and Seedance?

Black Forest Labs published its own human-preference results on 10-second 720p clips, reporting evaluators preferred FLUX 3 Video over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%, with narrower margins against Kling v3 Pro and Seedance 2.0. These are the lab's own tests, so treat them as directional rather than independent benchmarks.

How do I turn a FLUX 3 clip into posts across platforms?

FLUX 3 renders the video and audio but doesn't publish anything. Bring the export into a content engine like Kompozy to add branded captions, reframe per platform, stack a hook overlay, fan it into a carousel, quote card, and blog with captions in your voice, then schedule and publish across TikTok, Reels, YouTube Shorts, X, LinkedIn, and more from one queue.

Related news

← All AI news · Get started →