Seedance 2.5とは?ByteDanceの動画生成AIの位置づけ
Seedance 2.5とは、ByteDance の研究チーム「Seed」が公開した動画生成モデルです。同社は動画生成AIを Seedance という名前で世代を重ねており、2.5 はその最新版にあたります。
公式ブログは、前世代の Seedance 2.0 で作った土台の上に載せたものだと説明しています。土台とは、映像と音声をひとつのモデルの中でまとめて作る仕組みのことです。映像を作ってから音を後付けするのではなく、両方を同時に生成します。
Seedance 2.5 の概要(公式ブログによる)
「ワンテイク生成」と「柔軟な参照」という2本柱
記事タイトルにも置かれている one-take creation(ワンテイク生成) は、映画の撮影用語から来ています。カメラを止めずに一続きで撮ることです。動画生成AIでこれを言うのは、短い断片を何本も作って後でつなぐのではなく、始まりから終わりまでを1回の生成で出すという意味になります。
もう一方の flexible referencing(柔軟な参照) は、手元の素材をモデルに見せて「これを使って作れ」と指示できる範囲の広さを指します。公式は、この2つを軸にして、長い話の組み立て・参照・編集の3方向で大きな前進があったとしています。
Today, we are officially launching Seedance 2.5, the new-generation video creation model. / Building on the unified multimodal audio-video joint-generation architecture of Seedance 2.0, Seedance 2.5 centers on foundational generation and reference-based generation, delivering major breakthroughs in long-form storytelling, multimodal reference, and editing. — 公開の告知および前世代からの位置づけに関する記述より
30秒のワンテイク生成と多段延長でできること
Seedance 2.5 の変化のうち、いちばん数字で分かりやすいのが生成できる長さです。
単発生成が15秒から30秒に伸びた
1回の生成で作れる長さが、15秒から30秒になりました。倍になったこと自体より、公式が「その30秒に何を入れられるか」を強調している点が重要です。単に1つの瞬間を引き伸ばすのではなく、導入・展開・転換・結末という順序で話が進むように、複数のカットを組み立てられるとしています。
公式が挙げている例は、歌手のステージ映像です。舞台に上がる瞬間だけを写すのではなく、楽屋でスタッフとやりとりし、舞台裏の通路を歩き、ダンサーと合流し、そのままステージへ出るまでを一続きで描く、という組み立てです。
Seedance 2.5 can generate high-quality, 30-second audio-video clips in a single pass and supports multiple rounds of extension. / Seedance 2.5 extends single-pass video generation from 15 to 30 seconds and further strengthens its storytelling in longer videos. / Within 30 seconds, the model can organize multiple logically connected shots so that a story unfolds through setup, development, turning points, and resolution, rather than simply extending a single moment. — 単発生成の長さと、30秒の中で複数のカットを組み立てる点に関する記述より
多段延長で数分の動画までつなげる
さらに、生成済みの動画に続きを足していく multi-round extensions(多段延長) が使えます。延長の途中でも、主要な登場人物・環境・話の進み方の一貫性を保つとされています。
これが効くのは編集の手間です。公式は、断片に分けて何度もつなぎ直し、つなぎ目を直す作業を減らせると書いています。動画生成AIを実務で使うときにいちばん時間を食うのがこの後工程なので、そこを削るという主張になります。
Throughout the extension process, it maintains the consistency of main characters, environments, and narrative pacing. / This allows users to output videos lasting several minutes at once, reducing the effort required to split clips, repeatedly splice footage, and fix transitions. — 多段延長中に保たれるものと、後工程の削減に関する記述より
参照素材は画像30枚・動画10本・音声10本まで
もうひとつの柱が参照です。作りたいものを文章だけで説明するのではなく、実物の素材を渡して指定できます。
参照素材の上限は画像30枚・動画10本・音声10本
1回の生成で渡せる参照素材の上限
| 種類 | 上限 | 使われ方の例(公式より) |
|---|---|---|
| 画像 | 30枚 | 会場・演奏者・楽器・観客席をそれぞれ別画像で指定 |
| 動画 | 10本 | 既存映像を渡して続きを生成、または背景だけ差し替え |
| 音声 | 10本 | 複数の登場人物の声をそれぞれ保ったまま生成 |
公式が挙げている例では、コンサート映像を作るのに、会場・ピアニスト・チェロ・バイオリン・歌手・オーケストラ・合唱団・観客席を、すべて別々の画像として指定しています。素材の数と種類が増えるほど、作り手の意図を細かく反映できるという主張です。
Users can now input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. — 1回の生成で渡せる参照素材の上限に関する記述より
クレイレンダー参照で構図とカメラの動きを先に決める
参照の種類として公式が新たに強調しているのが clay render(クレイレンダー)参照 です。クレイレンダーとは、色や質感を貼らない灰色の3Dモデルだけの状態を指します。粘土(clay)で作った模型のように見えるためこの名前です。
これを渡すと、空間の構造・人物の姿勢・動く経路・カメラの角度を先に決めておけます。モデルはその骨組みに沿って映像を作るので、複雑なカットでも構図が意図から外れにくくなるという説明です。加えて、骨組みから得た空間情報をもとに、光源の向き・色温度・強さ・影の落ち方を物理法則に沿って作るとしています。
参照に使う画像を何十枚も用意するときは、サイズをそろえておくと扱いが楽になります。
For instance, with clay render referencing, users can build a scene's spatial structure, character poses, motion paths, and camera angles using textureless 3D models. / Additionally, Seedance 2.5 improves lighting control. By leveraging the spatial information from the clay render, it generates realistic lighting effects that follow physical laws, such as light source direction, color temperature, intensity, and shadow projection. — クレイレンダー参照で指定できる要素と、そこから作られる光と影に関する記述より
無料ツール画像リサイズ画像のサイズ(幅×高さ)を変更。ピクセル指定やパーセント指定に対応。今すぐ使ってみる →
秒数を指定する編集機能と、産業用途への広がり
3つめの柱が編集です。生成しっぱなしではなく、出来上がった映像の一部を狙って直せます。
タイムスタンプ指定でカットの秒数まで制御する
Seedance 2.5 は timestamp-level control(タイムスタンプ単位の制御) に対応しています。タイムスタンプとは、動画の中の「何秒の時点」を指す目印のことです。
生成の段階では、この秒からこの秒まではこう見せる、という形で話の進み方・カメラの視点・動き・全体のテンポを指示できます。生成した後も、特定の区間の人物・動作・筋書きだけを直せます。直した前後で映像がつながったままになるように保つ、という点が公式の主張です。
グリーンスクリーン編集も強化されています。グリーンスクリーンとは、緑一色の背景で撮影して後から別の背景に差し替える手法です。公式は、主役をそのまま残して背景を入れ替え、まったく別の物語に仕立てられるとしています。しかも入れ替えた環境の物理法則に主役が反応する——服のなびく向き、髪の状態、歩き方のリズム、光の当たり方——という点まで挙げています。
Seedance 2.5 offers timestamp-level control for targeted editing of audio and video content, notably improving efficiency and controllability. / During the generation phase, users can use prompts to control the narrative, camera perspective, movement, and overall rhythm for a specific time frame, aligning the output more closely with their creative intent. / In green screen editing, for example, the model can replace backgrounds and tell entirely different stories while keeping the main subject intact. / This includes the fluttering direction of clothes, the state of hair, gait rhythm, and lighting interaction, ensuring the subject blends harmoniously with the scene. — タイムスタンプ単位の制御、秒数を区切った指示、およびグリーンスクリーン編集に関する記述より
教育・製造・自動運転で使われはじめている
公式ブログの後半は、実際に使われている業種の話に移ります。ここは動画生成AIの用途が「作品づくり」から外へ出ている部分です。
公式が挙げた産業用途
自動運転の例に出てくる long-tail scenarios(ロングテールの状況) とは、めったに起きないが起きたときの影響が大きい状況を指します。実走行では滅多に集まらないため、生成して補うという発想です。
For example, Seedance 2.5 can turn the historical context, characters, and storylines behind a lesson into more vivid and immersive visuals. / It also helps teachers produce instructional videos more efficiently, turning abstract content — scientific principles, historical events, experimental procedures — into dynamic demonstrations. / The model can generate high-quality synthetic video data that helps train robots' perception and manipulation skills. / For autonomous driving, the model can simulate long-tail scenarios, such as extreme weather and complex traffic conditions, providing more diverse samples for system testing and training. — 教育・ロボット・自動運転それぞれでの用途に関する記述より
MiniMax H3との違い|重み公開モデルと製品として使うモデル
同じ週に、中国の別のAI企業からも動画と音声を同時に生成するモデルが出ています。MiniMax H3(ミニマックスH3)です。こちらは 重み——学習の結果として得られたデータ本体——が公開されており、手元の機材で動かせます。並べると、性格の違いがはっきりします。
Seedance 2.5 と MiniMax H3 の比較(各公式資料による)
| 観点 | Seedance 2.5 | MiniMax H3 |
|---|---|---|
| 提供の形 | サービス上で利用(Jimeng AI・Doubao Pro など) | 重みを公開。手元で動かせる |
| 1回の長さ | 最長30秒(多段延長で数分) | 4〜15秒 |
| 音声 | 映像と同時に生成 | 32kHzのステレオ音声を同時に生成 |
| 参照素材 | 画像30枚・動画10本・音声10本 | 画像9枚・動画3本・音声3本(合計12ファイルまで) |
| 使う条件 | 公式ブログに利用条件の記載なし | ライセンスの適用地域から米国・EU・英国・韓国を除外 |
数字だけ見れば Seedance 2.5 が上回りますが、そもそも置き場所が違います。片方はサービスの向こう側にあり、もう片方は自分の機材の上に置けます。開発者が自社の仕組みへ組み込みたいなら重みが公開されている側、完成度と尺を取りたいならサービスとして使う側、という分かれ方です。どこで動かすかが先に決まり、性能の比較はその後になります。
なお、生成された映像が後から本物として出回る問題は別途あり、判定する側の技術も動いています。NVIDIA の合成動画検出ツールがその一例です。
Today, Seedance 2.5 is rolling out on Jimeng AI, Doubao Pro, and other platforms, with API access coming soon via BytePlus ModelArk. / Seedance 2.5 extends single-pass video generation from 15 to 30 seconds and further strengthens its storytelling in longer videos. / Users can now input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. — 比較表の Seedance 2.5 側(提供の形・1回の長さ・参照素材)の根拠となる記述より
Output duration | 4–15 seconds / Output audio | 32 kHz stereo / <strong>Images:</strong> ≤ 9 images / <strong>Videos:</strong> ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds / <strong>Audio:</strong> ≤ 3 clips; audio must be accompanied by image or video input and cannot be used as the sole input; each clip must be 2–15 seconds long; total duration ≤ 15 seconds / <strong>Mixed inputs:</strong> Maximum number of files across all input types is 12 / We release the complete model weights to support further development, including fine-tuning. — 比較表の MiniMax H3 側(1回の長さ・音声・参照素材の上限・重みの公開)の根拠となる記述より
“Applicable Territory” means worldwide, excluding the Excluded Territories. / “Excluded Territories” means the European Union, the United Kingdom, the Republic of Korea and the United States of America. — 比較表の「使う条件」欄の根拠となる、適用地域と除外地域の定義より
まとめ:Seedance 2.5は動画生成AIをどこへ動かしたか
Seedance 2.5 の変化を一言でまとめると、断片を作る道具から、作品を仕上げる工程まで含む道具への移動です。公式自身が、クリップ単位の出力から創作の作業全体へ引き上げた、という書き方をしています。
30秒という長さも、参照素材30枚という数も、単体では単なる仕様の拡大です。ただし、秒単位で直せる編集と組み合わさると、意味が変わります。作り直すのではなく直せるようになると、AIで動画を作る作業が撮影後の編集作業に近づきます。
一方で、公式は限界も自分から書いています。複雑な動きの物理的なもっともらしさと、複数の被写体が絡む場面の安定性です。人が入り乱れる場面や、物が複雑に動く場面はまだ壊れうるということなので、使いどころを見きわめる材料になります。
Seedance 2.5 marks a significant step forward in understanding and rendering the real world, elevating video generation from clip-level outputs to comprehensive creative workflows. At the same time, we recognize there is still room for improvement, particularly regarding the physical plausibility of complex motions and the stability of scenes involving interactions among multiple subjects. — 総括および残る課題に関する記述より



