Black Forest Labs unveils FLUX 3, a single model spanning image, video, audio, and robot action
Black Forest Labs introduced FLUX 3, a multimodal foundation model the company says jointly learns from images, video, and audio inside one architecture rather than stitching together separate generators for each medium. The German image-generation startup, founded by former Stability AI researchers, is positioning the release as a step toward a single system that understands and generates across modalities, including physical action.
What's new
In its announcement, Black Forest Labs states: "FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture." The model builds on the company's Self-Flow approach to aligning multimodal generation and understanding, with training scaled across video, image, and audio data simultaneously rather than treating each modality as a separate problem.
On the video side, FLUX 3 generates text-to-video, image-to-video, and video-to-video clips up to 20 seconds long with native audio — dialogue, sound effects, and ambient noise generated alongside the visuals — plus keyframe-controlled transitions, multilingual dialogue, and agentic chaining for multi-shot sequences. For images, the company says "FLUX 3 can synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions," with improved handling of complex prompts and multilingual text rendering.
The more unusual piece is FLUX 3 Action, which integrates action prediction directly into the model as a "dynamics-aware foundation" for robotics. Partner company mimic has built FLUX-mimic on top of it for robot learning and is testing the system on production tasks at Audi. Black Forest Labs also published preliminary video-quality comparisons claiming FLUX 3 is preferred over Runway's Gen-4.5 in 77% of head-to-head evaluations and over Luma's Ray 3.2 in 93% — figures the company generated using its own test set and should be read as self-reported until independently reproduced.
Availability is staggered: FLUX 3 Video is in early access now, FLUX 3 Image follows in early access "in the following weeks," FLUX-mimic and FLUX 3 Action remain limited to selected partners, and an open-weight FLUX 3 Dev variant is planned for later release.
Context
Black Forest Labs built its reputation on the FLUX image-generation family, which became a widely used open-weight alternative to closed image models after the company's 2024 founding. FLUX 3 marks its first move beyond still images into video, audio, and now physical action prediction, putting it in more direct competition with video-generation specialists like Runway and Luma as well as broader multimodal players building toward embodied AI.
Why it matters
A single model trained jointly across image, video, audio, and action — rather than separate models bolted together — is a meaningful architectural bet: if it holds up, it should generalize better across tasks and make consistent multi-shot, multi-modal outputs easier to produce than pipelines that hand off between disconnected generators. The action-prediction angle is the more speculative wager, treating video generation and robot manipulation as versions of the same underlying prediction problem; the Audi pilot through mimic will be an early real-world test of whether that bet pays off outside the lab.
Corroborating sources
- Bfl
https://bfl.ai/blog/flux-3
“FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture.”