sakutto
Generative AI

FLUX 3 Video Pricing and Limits: 20-Second Clips With Native Audio

Video Generation AIFLUXBlack Forest LabsGenerative AI
FLUX 3 Video Pricing and Limits: 20-Second Clips With Native Audio

What FLUX 3 Video is

FLUX 3 Video is Black Forest Labs' video generation model. The first version, covering generation from text and from images, became generally available on August 4, 2026.

The underlying FLUX 3 is built as a model that generates and predicts video, audio, images and actions together, and this release opens up the video generation part of it for the first time. Audio is created together with the video rather than added to finished footage. That is the premise of the model.

What FLUX 3 Video can do today

Text to video
Takes a single sentence or a detailed prompt, generating movement, scene logic and audio
Image to video / keyframes
Start from an image, set an end frame, or set multiple keyframes and let it connect them
Video continuation
Hand it up to four seconds of your own video and audio and have it carry on
Multiple shots
Switch scenes and camera angles within one clip while keeping the sequence coherent
Audio and dialogue
Dialogue, sound effects and ambient sound generated along with the frames
Multiple languages
Dialogue in many languages, with lip-syncing — matching mouth movement to the speech

The one that earns its keep in production work is video continuation. Hand over up to four seconds of existing footage and audio, say what should happen next, and movement, camera behaviour and dialogue all carry across the seam. Material you have already shot does not have to be thrown away.

View official source →
Starting today, an initial version of FLUX 3 Video for generation from text and images is generally available via the BFL API and select partners. / The model generates clips up to 20 seconds long in HD resolution, with Full HD output via upscaling and native audio created alongside the video. / FLUX 3 is our frontier multimodal model for generating and predicting video, audio, images, and actions. / In its initial form, FLUX 3 Video can create video clips of up to 20 seconds length with native audio. / We are releasing our model at HD (720p) and Full HD (1080p) resolutions … / Text-to-Video: Describe a scene in simple language or using a detailed prompt. / FLUX 3 follows complex instructions while generating natural movements, scene logic, and audio. / Image-to-Video and Keyframes: Start with an image, specify an end frame, or set multiple keyframes in a clip. / FLUX 3 Video connects these in sequence while following the intended visual language. / Video Continuation: Provide FLUX 3 Video with up to four seconds of existing video and audio and tell it what should happen next. / The model takes both components into account to continue movement, camera behavior, dialogue, and audio across the video seam. / Creating Multiple Shots: Create multiple scenes and camera angles within a single video, while keeping the sequence coherent. / Audio and Dialogue: Generate dialogue, sound effects, and ambient sounds along with the individual frames. / Multi-Linguality: FLUX 3 Video is built to be a powerful tool for people of many different ethnicities and languages. — From the passages on how it became available (page dated August 4, 2026), clip length and resolution, the position of the model, and each of the six capabilities in the table above

Pricing and output: the rate changes with the mode

Billing is per second of output, and the rate changes with both mode and resolution. Getting that wrong will put an estimate off by nearly double.

ModeLengthFull generation rateDraft rate
Text to video5–20 s$0.17/s HD, $0.29/s FHD$0.06/s
Image to video5–20 s$0.17/s HD, $0.29/s FHD$0.06/s
Video continuation5–15 s$0.43/s HD, $0.54/s FHD$0.12/s

Every figure is per second of output. Continuation is the expensive mode, and it also caps out at 15 seconds instead of 20. The multiple is roughly 2.5x at HD, 1.9x at Full HD and 2x for Draft. Twenty seconds of Full HD from text works out at $0.29 × 20 = $5.80. The same 20 seconds cannot be produced through continuation at all.

Every mode outputs 24fps, with aspect ratios chosen from 21:9, 2:1, 16:9, 4:3, 1:1, 3:4 and 9:16. Drafts render at HD.

View official source →
| <strong>Text to Video</strong> | a prompt | 5–20 s | \$0.17/s hd, \$0.29/s fhd | \$0.06/s | / | <strong>Image to Video</strong> | a prompt + 1–10 images | 5–20 s | \$0.17/s hd, \$0.29/s fhd | \$0.06/s | / | <strong>Video Continuation</strong> | a prompt + your clip | 5–15 s | \$0.43/s hd, \$0.54/s fhd | \$0.12/s | / Every mode outputs 24 fps at `hd` or `fhd`, in aspect ratios 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, and 9:16. / Drafts render at `hd`. — From the three table rows giving the generatable length and per-second rate for each mode, and the description of the output specifications shared by all modes (frame rate, aspect ratios, Draft resolution)

Draft mode exists so that missing is cheap

The reason costs balloon in video generation is that you cannot tell whether the direction is right until you see the finished render. Draft mode aims at exactly that.

A draft returns a fast preview at a fraction of the cost, and once a draft is satisfactory the same content is rendered at full quality. Black Forest Labs states explicitly that the subjects, composition and motion carry over unchanged into the final output. If what you approved is what you get, then settling the composition at the draft stage is worth doing properly.

In per-second terms, a full HD-resolution generation is $0.17 against $0.06 for a draft. Three misses still cost less than one full render.

View official source →
Draft Mode: Enables the exploration of creative directions easily. / A draft generation returns a fast preview of your prompt at a fraction of the cost, so you can iterate on ideas instead of waiting for a full high-quality generation every time. / When a draft is satisfactory, FLUX 3 renders the video at full quality. / It includes the same subjects, same composition and same motion so the final output matches the version you approved. — From the passages on the purpose of Draft mode, the cost and speed of a draft generation, and what carries over into the full render

How it differs from Seedance 2.5 and MiniMax H3

Set beside Seedance 2.5 and MiniMax H3, released around the same time, FLUX 3 Video's position becomes clear.

AspectFLUX 3 VideoSeedance 2.5MiniMax H3
How you get itBFL API and partnersThrough services (Jimeng AI, Doubao Pro and others)Weights published; runs locally
Length per generation5–20 s (5–15 s for continuation)Up to 30 s (minutes via multi-stage extension)4–15 s
ResolutionHD (720p) / Full HD (1080p), 24fpsNot stated in the official blog post768px on the short side by default, 24fps
AudioGenerated alongside the videoGenerated alongside the video32kHz stereo generated alongside
Reference material1–10 images / up to 4 s of video30 images, 10 videos, 10 audio files9 images, 3 videos, 3 audio files (12 files total)
Pricing visibilityPer-second rates published ($0.06–$0.54)Not stated in the official blog postYour own compute, since weights are published

Figures for Seedance 2.5 and MiniMax H3 come from each company's own materials (sources are cited in the respective articles).

Read the table down its columns and the distinctions are clean. Seedance 2.5 has the longest output. MiniMax H3 is the only one you can run yourself. And FLUX 3 Video is the only one with published per-second rates you can budget against in advance.

The reason to choose FLUX 3 Video is neither length nor freedom. It is that the cost is knowable up front, and that the draft-to-final transition is guaranteed as a specification. If you need a single clip longer than 30 seconds, that is Seedance 2.5; if data cannot leave your premises, MiniMax H3. The split is straightforward.

The quality position BFL claims

In its own evaluation, Black Forest Labs says FLUX 3 clearly outperforms existing state-of-the-art models on text-to-video, and on image-to-video ties Seedance 2.0 while beating the rest. Human raters, it says, preferred the model on both.

This is an internal evaluation, not a third-party benchmark. And the model it ties with is Seedance 2.0, not the newer Seedance 2.5. Those two qualifications need to be kept in view.

View official source →
Human raters found it to be the preferred model for both text-to-video and image-to-video generation. / In our internal evaluation, FLUX 3 outperforms existing SOTA models in text to video generation by a solid margin. / It ties Seedance 2.0 and beats all other existing SOTA models in image-to-video. / Our next releases will expand FLUX 3 Video for enhanced controllability and ship capabilities for new modalities. / Our roadmap further includes FLUX 3 Image for image generation and editing, and FLUX 3 Dev as an open-weight variant. — From the passages on the human and internal evaluation results, the models compared against, and the planned future releases

What to know before using the FLUX 3 Video API

The API is asynchronous. You submit a request, poll the URL that comes back, and once the status is Ready you fetch the video from the result URL.

Result URLs are signed and they expire. The documentation is inconsistent here: the body text says they expire about two hours after the job finishes, while a diagram label on the same page says about 10 minutes. Build against the shorter figure to be safe. Either way, unless you save the file immediately on detecting completion, you will lose results.

Requests and responses are both JSON, and the parameter shape changes by mode, so it is quicker to format and inspect them locally while you build.

Free ToolJSON Formatter & ValidatorPretty-print or minify JSON data. Catch syntax errors instantly with line numbers and tree view.Try it now →

View official source →
Result URLs are signed and expire about 2 hours after the job finishes. / Download the video promptly once the status is `Ready`. / The signed result URL is valid for about 10 minutes — download it promptly. — From the statement about result URL expiry in the body text and the statement inside a diagram label on the same page (the two give different values)

The safety work went through a third party

On safety, Black Forest Labs says it worked with the third-party partner Cinder to evaluate FLUX 3 Video for a range of risks before release, validating its mitigations across the full range of supported modalities, including non-consensual intimate imagery (NCII) and child sexual abuse material (CSAM).

View official source →
We are committed to responsible development and deployment of AI models, and apply multiple layers of mitigation before, during, and after release to combat the risk of misuse. / Working with a trusted third-party partners, Cinder / we evaluated FLUX 3 Video for a range of risks prior to release to validate our mitigations across the full range of supported modalities, including non-consensual intimate imagery (NCII) and child sexual abuse material (CSAM). — From the responsible development and deployment section, on when mitigations are applied, the evaluation carried out with a named third-party partner, and the range of risks covered

How to use FLUX 3 Video

What sits at the centre of FLUX 3 Video is cost predictability. Per-second rates are published, and the path of checking direction cheaply in Draft before committing to a full render is provided as a specification rather than a workaround.

The 20-second ceiling and the unreleased weights are fixed in this version. If you need length, Seedance 2.5; if you need local execution, MiniMax H3. These three are less competitors than tools divided by use case, which is closer to how they actually behave.

To try it, run a handful of drafts to find the composition, then move that same content into a full render. That order costs the least.

FAQ

Q. How long can a FLUX 3 Video clip be?
Five to 20 seconds for text-to-video and image-to-video. Video Continuation, which extends an existing clip, is limited to five to 15 seconds. Two resolutions are offered, HD (720p) and Full HD (1080p), both at 24fps.
Black Forest Labs Blog — FLUX 3 Video - Generation Capabilities
In its initial form, FLUX 3 Video can create video clips of up to 20 seconds length with native audio. / We are releasing our model at HD (720p) and Full HD (1080p) resolutions … Black Forest Labs Blog — FLUX 3 Video - Generation Capabilities
Q. What does it cost?
Billing is per second of output. Generating from text or images costs $0.17 per second at HD, $0.29 at Full HD, and $0.06 in Draft mode. Video Continuation is more expensive: $0.43 per second at HD, $0.54 at Full HD, and $0.12 for Draft.
Black Forest Labs Documentation — FLUX 3 (the relevant rows of the table of lengths and rates by mode)
| **Text to Video** | a prompt | 5–20 s | \$0.17/s hd, \$0.29/s fhd | \$0.06/s | / | **Video Continuation** | a prompt + your clip | 5–15 s | \$0.43/s hd, \$0.54/s fhd | \$0.12/s | Black Forest Labs Documentation — FLUX 3 (the relevant rows of the table of lengths and rates by mode)
Q. Are the weights available? Can I run it locally?
Not at the moment. FLUX 3 Video is offered only through the BFL API and select partners. Black Forest Labs does list FLUX 3 Image, for image generation and editing, and FLUX 3 Dev, an open-weight variant, on its roadmap.
Black Forest Labs Blog — What comes next
Our roadmap further includes FLUX 3 Image for image generation and editing, and FLUX 3 Dev as an open-weight variant. Black Forest Labs Blog — What comes next
Q. Can it produce dialogue in languages other than English?
Yes. The languages listed by Black Forest Labs include English in various dialects, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi and Punjabi, and the company describes the lip-syncing as precise.
Black Forest Labs Blog — Multi-Linguality
Supported languages include English (various dialects), Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi and more, with precise lip-syncing. Black Forest Labs Blog — Multi-Linguality

Related Tools

Related Tool Categories

Articles