Black Forest Labs has moved beyond still-image generation with FLUX.3 Video, a new model family that creates video and sound together. The first release is available through the company’s API and selected partners, with support for text-to-video, image-to-video and video continuation workflows.
The headline capability is duration: FLUX.3 Video can generate clips of up to 20 seconds. Output is available at 720p, with 1080p delivery through an upscaling step. Native audio generation covers dialogue, effects and ambient sound, allowing a prompt to describe both what happens on screen and what should be heard.
One prompt can now describe a complete shot
FLUX.3 Video is designed to keep visual and audio generation inside the same workflow. A user can specify a scene, camera movement, spoken dialogue and background sound without first generating silent footage and then building a separate audio track. Black Forest Labs says the model supports multilingual dialogue and can ground generated audio in the action taking place on screen.
The release also supports multiple shots within a single clip. That matters for sequences that need a change of framing or scene rather than one continuous camera move. The company presents this as a way to produce more structured short-form material while retaining coherence between shots.
For image-led work, creators can provide keyframes to guide the result. Video continuation accepts up to four seconds of existing video and audio, then generates what follows. That gives production teams a route to extend an approved opening, develop variations from a selected take or bridge an existing asset into generated footage.
Draft mode separates exploration from final rendering
A draft mode is included for faster iteration. Teams can use lower-cost or faster previews to test composition, timing and prompt wording before committing to a final result. This is especially useful for video, where an unsuccessful full render costs more time and compute than an image variation.
The model supports conventional text instructions as well as image and video inputs, but the release does not remove the need for creative review. A 20-second ceiling still places FLUX.3 Video firmly in the short-clip category. Longer stories will require multiple generations, editing and attention to continuity across segments.
Black Forest Labs has published its own benchmark comparisons and says FLUX.3 Video performs strongly on visual quality, prompt adherence and audio-video alignment. Those results are vendor-reported rather than an independent evaluation. Buyers comparing systems should test the kinds of motion, dialogue, characters and brand assets they actually expect to use.
The first part of a broader FLUX.3 rollout
The “Part 1” label is deliberate. Black Forest Labs says more reference combinations are planned, alongside FLUX.3 Image and a FLUX.3 Dev release. The current announcement therefore establishes the generation layer, while later releases are expected to broaden control and development options.
Availability through the BFL API gives developers a direct integration path, while selected partner access may suit teams already using a creative platform. The announcement does not provide a single universal price table for every route, so costs and supported controls should be checked with the chosen provider.
FLUX.3 Video is a substantial expansion of the FLUX line: it brings moving images, synchronised sound and continuation into the same product family. Its practical value will depend less on headline benchmark scores than on consistency across repeated characters, usable takes and the amount of editing required after generation.