FLUX 3 Videoとは何か
FLUX 3 Videoとは、Black Forest Labsの動画生成モデルです。2026年8月4日にテキストと画像からの生成に対応した最初の版が一般提供になりました。
土台のFLUX 3は動画・音声・画像・行動をまとめて生成・予測するモデルとして作られていて、今回はそのうち動画生成の部分が初めて広く開放されました。映像を作ってから音声を足すのではなく、音声を映像と一緒に作る。これがこのモデルの前提です。
FLUX 3 Video で今日できること
とくに実務で効くのが動画の続き生成です。手持ちの映像と音声を最大4秒渡して「この先どうなるか」を指示すると、動きもカメラワークもセリフもつなぎ目をまたいで続きます。撮ってある素材を捨てずに伸ばせる。
Starting today, an initial version of FLUX 3 Video for generation from text and images is generally available via the BFL API and select partners. / The model generates clips up to 20 seconds long in HD resolution, with Full HD output via upscaling and native audio created alongside the video. / FLUX 3 is our frontier multimodal model for generating and predicting video, audio, images, and actions. / In its initial form, FLUX 3 Video can create video clips of up to 20 seconds length with native audio. / We are releasing our model at HD (720p) and Full HD (1080p) resolutions … / Text-to-Video: Describe a scene in simple language or using a detailed prompt. / FLUX 3 follows complex instructions while generating natural movements, scene logic, and audio. / Image-to-Video and Keyframes: Start with an image, specify an end frame, or set multiple keyframes in a clip. / FLUX 3 Video connects these in sequence while following the intended visual language. / Video Continuation: Provide FLUX 3 Video with up to four seconds of existing video and audio and tell it what should happen next. / The model takes both components into account to continue movement, camera behavior, dialogue, and audio across the video seam. / Creating Multiple Shots: Create multiple scenes and camera angles within a single video, while keeping the sequence coherent. / Audio and Dialogue: Generate dialogue, sound effects, and ambient sounds along with the individual frames. / Multi-Linguality: FLUX 3 Video is built to be a powerful tool for people of many different ethnicities and languages. — 一般提供の開始形態(ページ掲載日 August 4, 2026)、クリップ長と解像度、モデルの位置づけ、および図表に挙げた6機能それぞれの説明より
料金と出力仕様:モードで単価が変わる
料金は秒単位の従量課金で、モードと解像度で単価が変わります。ここを取り違えると見積もりが倍近くずれます。
| モード | 生成できる長さ | 本生成の単価 | Draftの単価 |
|---|---|---|---|
| テキスト→動画 | 5〜20秒 | HD 0.17ドル/FHD 0.29ドル | 0.06ドル |
| 画像→動画 | 5〜20秒 | HD 0.17ドル/FHD 0.29ドル | 0.06ドル |
| 動画の続き生成 | 5〜15秒 | HD 0.43ドル/FHD 0.54ドル | 0.12ドル |
いずれも1秒あたりの金額です。続き生成だけが割高で、生成できる長さも15秒までに縮みます。倍率はHDで約2.5倍、Full HDで約1.9倍、Draftで2倍。20秒のFull HDをテキストから作れば0.29×20で5.8ドル。同じ20秒を続き生成で作ることはできません。
出力はどのモードも24fps(1秒あたり24コマ)で、縦横比は21:9・2:1・16:9・4:3・1:1・3:4・9:16から選びます。Draftで作った下書きはHDで描かれます。
| <strong>Text to Video</strong> | a prompt | 5–20 s | \$0.17/s hd, \$0.29/s fhd | \$0.06/s | / | <strong>Image to Video</strong> | a prompt + 1–10 images | 5–20 s | \$0.17/s hd, \$0.29/s fhd | \$0.06/s | / | <strong>Video Continuation</strong> | a prompt + your clip | 5–15 s | \$0.43/s hd, \$0.54/s fhd | \$0.12/s | / Every mode outputs 24 fps at `hd` or `fhd`, in aspect ratios 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, and 9:16. / Drafts render at `hd`. — モード別の生成可能秒数と秒単価をまとめた表の3行、および全モード共通の出力仕様(フレームレート・縦横比・Draft の解像度)の記述より
Draftモードは「安く外す」ための仕組み
動画生成でコストが膨らむ最大の理由は仕上がりを見るまで方向が合っているか分からないことです。FLUX 3 VideoのDraftモードはここを狙っています。
Draftは本生成の一部の費用で速いプレビューを返し、満足いく下書きができたら同じ内容を高品質で描き直します。公式は被写体・構図・動きが下書きと同じまま最終出力になると明記しています。承認した見た目がそのまま出るなら、下書きの段階で構図を詰めておく手が生きます。
秒単価で見ると、HDの本生成0.17ドルに対してDraftは0.06ドル。3回外しても本生成1回より安い計算です。
Draft Mode: Enables the exploration of creative directions easily. / A draft generation returns a fast preview of your prompt at a fraction of the cost, so you can iterate on ideas instead of waiting for a full high-quality generation every time. / When a draft is satisfactory, FLUX 3 renders the video at full quality. / It includes the same subjects, same composition and same motion so the final output matches the version you approved. — Draft モードの目的、下書き生成の費用と速度、および本生成へ引き継がれる要素に関する記述より
Seedance 2.5・MiniMax H3との違い
同じ時期に公開された Seedance 2.5(シードダンス2.5)・MiniMax H3(ミニマックスH3)と並べると、FLUX 3 Videoの立ち位置がはっきりします。
| 観点 | FLUX 3 Video | Seedance 2.5 | MiniMax H3 |
|---|---|---|---|
| 提供の形 | BFLのAPI・提携先 | サービス上で利用(Jimeng AI・Doubao Pro など) | 重みを公開。手元で動かせる |
| 1回の長さ | 5〜20秒(続き生成は5〜15秒) | 最長30秒(多段延長で数分) | 4〜15秒 |
| 解像度 | HD(720p)/Full HD(1080p)・24fps | 公式ブログに解像度の記載なし | 既定は短辺768px・24fps |
| 音声 | 映像と同時に生成 | 映像と同時に生成 | 32kHzのステレオ音声を同時に生成 |
| 参照素材 | 画像1〜10枚/動画は最大4秒 | 画像30枚・動画10本・音声10本 | 画像9枚・動画3本・音声3本(合計12ファイルまで) |
| 料金の出方 | 秒単価が公開(0.06〜0.54ドル) | 公式ブログに料金の記載なし | 重み公開のため自前の計算資源 |
※ Seedance 2.5・MiniMax H3 の数値は各社の公式資料にもとづきます(出典は各記事内に明記)。
表を縦に見ると差がはっきりします。長さは Seedance 2.5 が最長。手元で動かせるのは MiniMax H3 だけ。そして秒単価が公開されていて事前に見積もれるのは FLUX 3 Video だけです。
FLUX 3 Videoを選ぶ理由は長さでも自由度でもありません。費用が事前に読めること、そして下書きから本生成への移行が仕様として保証されていることです。逆に30秒以上を1本で作りたいならSeedance 2.5、社外にデータを出せないならMiniMax H3。ここは素直に分かれます。
公式が主張する品質の位置
公式は自社評価として、テキストから動画では既存のSOTA(最高水準)モデルを明確に上回り、画像から動画ではSeedance 2.0と同等で他は上回るとしています。人間の評価者による比較でも両方で最も好まれたという説明です。
ただしこれは内部評価であり第三者のベンチマークではありません。同等とされた比較対象がSeedance 2.0であって、より新しいSeedance 2.5ではない。ここは分けて読む必要があります。
Human raters found it to be the preferred model for both text-to-video and image-to-video generation. / In our internal evaluation, FLUX 3 outperforms existing SOTA models in text to video generation by a solid margin. / It ties Seedance 2.0 and beats all other existing SOTA models in image-to-video. / Our next releases will expand FLUX 3 Video for enhanced controllability and ship capabilities for new modalities. / Our roadmap further includes FLUX 3 Image for image generation and editing, and FLUX 3 Dev as an open-weight variant. — 人手評価と内部評価の結果、比較対象モデル、および今後のリリース計画に関する記述より
FLUX 3 VideoのAPIを使う前に押さえる点
API(外部プログラムから機能を呼び出す窓口)は非同期です。リクエストを投げ、返ってきたポーリング用URL(完了したかを繰り返し問い合わせるための宛先)を叩き、状態がReadyになったら結果のURLから動画を取得します。
結果のURLは署名つきで期限があります。ここは公式ドキュメント内で記述が割れていて、本文には「ジョブ完了から約2時間で失効する」とあり、同じページの図解ラベルには「約10分」とあります。短いほうを前提に組むのが安全です。いずれにせよ、完了検知の直後に保存する作りにしないと取りこぼします。
リクエストもレスポンスもJSON形式です。モードごとにパラメータの形が変わるので、手元で整形して中身を見ながら組むほうが早いです。
無料ツールJSON整形・検証JSONデータを見やすく整形&構文エラーを検証。開発やAPI連携に必須。今すぐ使ってみる →
Result URLs are signed and expire about 2 hours after the job finishes. / Download the video promptly once the status is `Ready`. / The signed result URL is valid for about 10 minutes — download it promptly. — 本文中の結果URL失効に関する記述、および同一ページの図解ラベル内の記述より(両者で期限の値が異なる)
安全対策は第三者評価を経ている
安全対策は、公式が第三者のCinder(シンダー)と組んで公開前にリスク評価を実施したと説明しています。非同意の性的画像(NCII)や児童性的虐待表現(CSAM)を含む範囲で、対応するすべての種類の出力について緩和策を検証したという書き方です。
We are committed to responsible development and deployment of AI models, and apply multiple layers of mitigation before, during, and after release to combat the risk of misuse. / Working with a trusted third-party partners, Cinder / we evaluated FLUX 3 Video for a range of risks prior to release to validate our mitigations across the full range of supported modalities, including non-consensual intimate imagery (NCII) and child sexual abuse material (CSAM). — 責任ある開発と展開の節における、緩和策の適用時期、第三者パートナーの名称を含む評価実施、および評価対象リスクの範囲に関する記述より
まとめ:FLUX 3 Videoをどう使うか
FLUX 3 Videoの中心は費用の見通しが立つことです。秒単価が公開され、Draftで安く方向を確かめてから本生成に進む道筋が仕様として用意されています。
20秒という上限と重み非公開という制約は、今の版では動きません。長さが要るならSeedance 2.5、手元実行が要るならMiniMax H3。この3本は競合というより用途で分かれていると見たほうが実態に合います。
まず試すならDraftを数本回して構図の当たりを取り、そこから本生成へ移す。この順番がいちばん安く済みます。



