What FLUX 3 Video is
FLUX 3 Video is Black Forest Labs' video generation model. The first version, covering generation from text and from images, became generally available on August 4, 2026.
The underlying FLUX 3 is built as a model that generates and predicts video, audio, images and actions together, and this release opens up the video generation part of it for the first time. Audio is created together with the video rather than added to finished footage. That is the premise of the model.
What FLUX 3 Video can do today
The one that earns its keep in production work is video continuation. Hand over up to four seconds of existing footage and audio, say what should happen next, and movement, camera behaviour and dialogue all carry across the seam. Material you have already shot does not have to be thrown away.
Starting today, an initial version of FLUX 3 Video for generation from text and images is generally available via the BFL API and select partners. / The model generates clips up to 20 seconds long in HD resolution, with Full HD output via upscaling and native audio created alongside the video. / FLUX 3 is our frontier multimodal model for generating and predicting video, audio, images, and actions. / In its initial form, FLUX 3 Video can create video clips of up to 20 seconds length with native audio. / We are releasing our model at HD (720p) and Full HD (1080p) resolutions … / Text-to-Video: Describe a scene in simple language or using a detailed prompt. / FLUX 3 follows complex instructions while generating natural movements, scene logic, and audio. / Image-to-Video and Keyframes: Start with an image, specify an end frame, or set multiple keyframes in a clip. / FLUX 3 Video connects these in sequence while following the intended visual language. / Video Continuation: Provide FLUX 3 Video with up to four seconds of existing video and audio and tell it what should happen next. / The model takes both components into account to continue movement, camera behavior, dialogue, and audio across the video seam. / Creating Multiple Shots: Create multiple scenes and camera angles within a single video, while keeping the sequence coherent. / Audio and Dialogue: Generate dialogue, sound effects, and ambient sounds along with the individual frames. / Multi-Linguality: FLUX 3 Video is built to be a powerful tool for people of many different ethnicities and languages. — From the passages on how it became available (page dated August 4, 2026), clip length and resolution, the position of the model, and each of the six capabilities in the table above
Pricing and output: the rate changes with the mode
Billing is per second of output, and the rate changes with both mode and resolution. Getting that wrong will put an estimate off by nearly double.
| Mode | Length | Full generation rate | Draft rate |
|---|---|---|---|
| Text to video | 5–20 s | $0.17/s HD, $0.29/s FHD | $0.06/s |
| Image to video | 5–20 s | $0.17/s HD, $0.29/s FHD | $0.06/s |
| Video continuation | 5–15 s | $0.43/s HD, $0.54/s FHD | $0.12/s |
Every figure is per second of output. Continuation is the expensive mode, and it also caps out at 15 seconds instead of 20. The multiple is roughly 2.5x at HD, 1.9x at Full HD and 2x for Draft. Twenty seconds of Full HD from text works out at $0.29 × 20 = $5.80. The same 20 seconds cannot be produced through continuation at all.
Every mode outputs 24fps, with aspect ratios chosen from 21:9, 2:1, 16:9, 4:3, 1:1, 3:4 and 9:16. Drafts render at HD.
| <strong>Text to Video</strong> | a prompt | 5–20 s | \$0.17/s hd, \$0.29/s fhd | \$0.06/s | / | <strong>Image to Video</strong> | a prompt + 1–10 images | 5–20 s | \$0.17/s hd, \$0.29/s fhd | \$0.06/s | / | <strong>Video Continuation</strong> | a prompt + your clip | 5–15 s | \$0.43/s hd, \$0.54/s fhd | \$0.12/s | / Every mode outputs 24 fps at `hd` or `fhd`, in aspect ratios 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, and 9:16. / Drafts render at `hd`. — From the three table rows giving the generatable length and per-second rate for each mode, and the description of the output specifications shared by all modes (frame rate, aspect ratios, Draft resolution)
Draft mode exists so that missing is cheap
The reason costs balloon in video generation is that you cannot tell whether the direction is right until you see the finished render. Draft mode aims at exactly that.
A draft returns a fast preview at a fraction of the cost, and once a draft is satisfactory the same content is rendered at full quality. Black Forest Labs states explicitly that the subjects, composition and motion carry over unchanged into the final output. If what you approved is what you get, then settling the composition at the draft stage is worth doing properly.
In per-second terms, a full HD-resolution generation is $0.17 against $0.06 for a draft. Three misses still cost less than one full render.
Draft Mode: Enables the exploration of creative directions easily. / A draft generation returns a fast preview of your prompt at a fraction of the cost, so you can iterate on ideas instead of waiting for a full high-quality generation every time. / When a draft is satisfactory, FLUX 3 renders the video at full quality. / It includes the same subjects, same composition and same motion so the final output matches the version you approved. — From the passages on the purpose of Draft mode, the cost and speed of a draft generation, and what carries over into the full render
How it differs from Seedance 2.5 and MiniMax H3
Set beside Seedance 2.5 and MiniMax H3, released around the same time, FLUX 3 Video's position becomes clear.
| Aspect | FLUX 3 Video | Seedance 2.5 | MiniMax H3 |
|---|---|---|---|
| How you get it | BFL API and partners | Through services (Jimeng AI, Doubao Pro and others) | Weights published; runs locally |
| Length per generation | 5–20 s (5–15 s for continuation) | Up to 30 s (minutes via multi-stage extension) | 4–15 s |
| Resolution | HD (720p) / Full HD (1080p), 24fps | Not stated in the official blog post | 768px on the short side by default, 24fps |
| Audio | Generated alongside the video | Generated alongside the video | 32kHz stereo generated alongside |
| Reference material | 1–10 images / up to 4 s of video | 30 images, 10 videos, 10 audio files | 9 images, 3 videos, 3 audio files (12 files total) |
| Pricing visibility | Per-second rates published ($0.06–$0.54) | Not stated in the official blog post | Your own compute, since weights are published |
Figures for Seedance 2.5 and MiniMax H3 come from each company's own materials (sources are cited in the respective articles).
Read the table down its columns and the distinctions are clean. Seedance 2.5 has the longest output. MiniMax H3 is the only one you can run yourself. And FLUX 3 Video is the only one with published per-second rates you can budget against in advance.
The reason to choose FLUX 3 Video is neither length nor freedom. It is that the cost is knowable up front, and that the draft-to-final transition is guaranteed as a specification. If you need a single clip longer than 30 seconds, that is Seedance 2.5; if data cannot leave your premises, MiniMax H3. The split is straightforward.
The quality position BFL claims
In its own evaluation, Black Forest Labs says FLUX 3 clearly outperforms existing state-of-the-art models on text-to-video, and on image-to-video ties Seedance 2.0 while beating the rest. Human raters, it says, preferred the model on both.
This is an internal evaluation, not a third-party benchmark. And the model it ties with is Seedance 2.0, not the newer Seedance 2.5. Those two qualifications need to be kept in view.
Human raters found it to be the preferred model for both text-to-video and image-to-video generation. / In our internal evaluation, FLUX 3 outperforms existing SOTA models in text to video generation by a solid margin. / It ties Seedance 2.0 and beats all other existing SOTA models in image-to-video. / Our next releases will expand FLUX 3 Video for enhanced controllability and ship capabilities for new modalities. / Our roadmap further includes FLUX 3 Image for image generation and editing, and FLUX 3 Dev as an open-weight variant. — From the passages on the human and internal evaluation results, the models compared against, and the planned future releases
What to know before using the FLUX 3 Video API
The API is asynchronous. You submit a request, poll the URL that comes back, and once the status is Ready you fetch the video from the result URL.
Result URLs are signed and they expire. The documentation is inconsistent here: the body text says they expire about two hours after the job finishes, while a diagram label on the same page says about 10 minutes. Build against the shorter figure to be safe. Either way, unless you save the file immediately on detecting completion, you will lose results.
Requests and responses are both JSON, and the parameter shape changes by mode, so it is quicker to format and inspect them locally while you build.
Result URLs are signed and expire about 2 hours after the job finishes. / Download the video promptly once the status is `Ready`. / The signed result URL is valid for about 10 minutes — download it promptly. — From the statement about result URL expiry in the body text and the statement inside a diagram label on the same page (the two give different values)
The safety work went through a third party
On safety, Black Forest Labs says it worked with the third-party partner Cinder to evaluate FLUX 3 Video for a range of risks before release, validating its mitigations across the full range of supported modalities, including non-consensual intimate imagery (NCII) and child sexual abuse material (CSAM).
We are committed to responsible development and deployment of AI models, and apply multiple layers of mitigation before, during, and after release to combat the risk of misuse. / Working with a trusted third-party partners, Cinder / we evaluated FLUX 3 Video for a range of risks prior to release to validate our mitigations across the full range of supported modalities, including non-consensual intimate imagery (NCII) and child sexual abuse material (CSAM). — From the responsible development and deployment section, on when mitigations are applied, the evaluation carried out with a named third-party partner, and the range of risks covered
How to use FLUX 3 Video
What sits at the centre of FLUX 3 Video is cost predictability. Per-second rates are published, and the path of checking direction cheaply in Draft before committing to a full render is provided as a specification rather than a workaround.
The 20-second ceiling and the unreleased weights are fixed in this version. If you need length, Seedance 2.5; if you need local execution, MiniMax H3. These three are less competitors than tools divided by use case, which is closer to how they actually behave.
To try it, run a handful of drafts to find the composition, then move that same content into a full render. That order costs the least.



