Black Forest Labs has opened early access to FLUX 3, a new multimodal foundation model designed to work across images, video, audio and actions within one architecture. The first accessible component is FLUX 3 Video, which generates video with native audio, while image generation, action prediction and an open-weight backbone are scheduled for later stages.
The announcement broadens Black Forest Labs beyond the image-focused FLUX 1 and FLUX 2 families. It also places the company in the growing market for models that combine media generation with representations intended for simulation, robotics and other forms of physical AI. For developers and creative teams, however, the practical offer is still an early-access release rather than a finished, generally available product suite.
What FLUX 3 can do in early access
Black Forest Labs says FLUX 3 was trained jointly on images, video and audio, rather than adding separate media systems around a single-modality model. The company describes this as a way to learn relationships between appearance, motion, sound and instructions within a shared representation.
FLUX 3 Video is the most developed part of the launch. According to the official announcement, it can create clips of up to 20 seconds with native audio and supports several generation and editing patterns:
- text-to-video and image-to-video generation;
- video-to-video transformation using a reference clip;
- video and audio continuation from existing material;
- keyframe-controlled transitions;
- multilingual dialogue and synchronised sound; and
- agentic chaining of clips into longer, multi-shot sequences.
Early access is important context. Black Forest Labs says its evaluations are preliminary and that further improvements are expected. It reports preference-test results against several competing video models, but those figures were produced for the company’s own launch material and have not been independently validated. Buyers should assess output quality, consistency, latency and cost against their own production workloads.
A staged plan for image, action and open weights
FLUX 3 Image is not part of the initial public early-access offer. The company says image synthesis and editing will enter early access in the following weeks. It also plans API and private-weight access for video and audio, selected-partner access for action prediction, and an open-weight multimodal backbone called FLUX 3 Dev.
That distinction matters when comparing the announcement with an immediately available multimodal API. Some capabilities demonstrated in the launch are planned products, not features that every customer can use today. Black Forest Labs has not published final general-availability dates, production service levels or complete pricing for the new family.
The planned action component connects FLUX 3 to robotics. A companion technical post describes FLUX-mimic, a video-action model developed with mimic robotics using the FLUX 3 backbone. Black Forest Labs says the system has been tested on factory tasks at Audi, including manipulation of flexible parts that are difficult for conventional automation.
The company reports that FLUX-mimic can run its backbone in under 80 milliseconds on a single NVIDIA RTX 5090, with a complete robot system reacting in 101 milliseconds. Those numbers describe a partner implementation and should not be read as universal performance guarantees. Hardware, sensors, deployment software and task complexity will all affect real-world results.
Why the unified architecture matters
The central technical claim is that media generation and action prediction can share the same learned representation. Video training requires a model to predict motion and cause-and-effect relationships; Black Forest Labs argues that these representations can also support an action decoder for robotic control.
This approach could reduce the need to build separate foundations for creative generation and physical-world tasks. It may also help transfer knowledge from large media datasets into smaller, specialised robotics datasets. The companion research material says action prediction initially reduced video quality during training, before the model recovered its earlier performance while retaining the new capability.
For organisations evaluating FLUX 3, the most immediate questions are operational. They should confirm which early-access tier is available, what data-handling terms apply, whether outputs include usable provenance controls, and when each promised component will receive stable APIs and commercial terms. Teams using the model for robotics will also need substantially different safety, testing and failure-handling processes from teams using it for creative production.
What comes next
Black Forest Labs plans to release the family progressively over the coming weeks and months. The sequencing gives the company time to gather feedback and conduct safety testing, but it also leaves important details unresolved. The launch does not yet specify final model sizes, rate limits, geographic availability, general-availability dates or open-weight licence terms for FLUX 3 Dev.
Even with those gaps, FLUX 3 is a material expansion of the FLUX product line. Its initial video-and-audio offer puts a unified multimodal model into customer testing now, while the image, action and open-weight roadmap signals a broader attempt to connect generative media with physical AI.