2週間限定 | Seedream 5.0 Pro が20%OFF!

1つのAPIキー、1本の15秒フィルム:Seedream 5.0 15sワークフローの完全版

深度マップ Seedream 5.0 15秒ワークフロー: ジオメトリとスタイルを分離し、9つのショットをストーリーボード化し、1つのAPIキーで映画的なショートをレンダリングします。完全なプロンプトは内部に。

週末で作った。孤独な保管者が無限のアーカイブに足を踏み入れ、本を開くと、ページから光があふれ出す。15秒。キャラクターデザイン、ストーリーボード、モーション、すべて生成され、すべてが単一のAPIキーから。レンダーファームも、3アプリのパイプラインも、プラグインの動物園も必要ない。

秘訣は、より賢いプロンプトではなかった。作業を2つに分割したのだ。ほとんどの人は、色、光、テクスチャ、ジオメトリをすべて一度に一つのモデルに任せ、結果がフラットになる理由に戸惑う。これがSeedream 5.0 15sワークフローの修正点だ。まず深度マップでジオメトリを固定し、後からスタイルを塗る。以下が全パイプラインで、実際に使用したすべてのプロンプトを載せている。

重要なポイント

  • Seedream 5.0 15sワークフローの核心:ジオメトリ(深度マップ)とスタイル(カラーカード)を分離することで、構図を崩さずにショットを無限にリスタイルできる。
  • 9つのカメラアングルを事前に1つの3x3深度グリッドとしてストーリーボード化し、ビデオモデルが連続性を維持し、無関係な9フレームを生成しないようにする。
  • 2026年7月下旬確認の価格:Seedream 5.0 Proの画像は1.5Kティアで$0.045から、高品質のGPT Image 2の約5分の1。Seedream 5.0 Proで画像、Seedance 2.0でビデオ、1つのキーで。

Seedream 5.0 15sワークフローで深度マップを先に使う理由

ジオメトリは画像モデルが最初に落とすものだ。色、照明、マテリアル、3D空間を一度に解決させようとすると、空間が崩壊する。これがAIアートの多くが、奥行きのないペラッとしたステッカーのように見える理由だ。この修正は、拡散ツールでよく知られたアイデアを借りている。モデルに制御画像を渡し、その固定された構造の上に構築させる(ControlNet、ICCV 2023)。

よく使われる制御画像は3種類あり、その差は見た目以上に大きい。以下はこのワークフロー内での比較だ。

制御画像エンコードするもの役立つ理由失敗する点
線画エッジ、ポーズ、構図レイアウトとキャラクターのポーズを固定深度情報がないため、フレームは依然としてフラットに読める
3Dホワイトモデルボリュームとシルエット質量感のあるリアルな感覚過制御:クレイレンダーの見た目が最終画像ににじみ出て、除去できない
深度マップ距離のみ(白が近く、黒が遠く)正確な空間、スタイルの混入ゼロ別途スタイルパスが必要。それがポイント

線画ストーリーボード:エッジとポーズは固定されているが、すべてのパネルがフラットで奥行きがない

線画のストーリーボード:エッジとポーズは固定されているが、すべてのパネルがフラットで奥行きがない。

同じストーリーボードを3Dホワイトモデルレンダリングで。灰色のクレイ調が画像を支配する

同じストーリーボードを3Dホワイトモデルでレンダリング。灰色のクレイ調が画像を乗っ取る。

3Dホワイトモデルにはボリュームがあるが、過制御になる。モデルはその灰色のクレイ肌を画像そのものとして扱い、後続のすべてのステップでプラスチックのようなレンダリングルックと戦うことになる。

深度マップはただ一つのこと、カメラからの距離だけをエンコードする。白が近く、黒が遠く、その間を連続的なグレーのランプでつなぐ(Depth Anything V2、2024)。これにより線画よりも正確で、実際の空間を運び、ホワイトモデルよりもクリーンで、リークするマテリアルやスタイルがない。どんなルックを上から重ねても、ジオメトリが汚染されることはない。ジオメトリをきれいに固定し、後からスタイルを施す。この単一の習慣が、Seedream 5.0 15sワークフローの残りをまとめている。

1つのAPIキーでSeedream 5.0 15sワークフロー全体を実行

これが週末プロジェクトであって、システム統合の頭痛の種にならない理由は、チェーン内のすべてのモデルが1つのエンドポイントの背後にあるからだ。画像はSeedream 5.0 Pro、ビデオはSeedance 2.0で実行され、両方ともAtlas Cloud上の同じAPIキーに応答する。キャラクターカードと最終レンダーの間でプラットフォームを切り替える必要も、2つの請求ダッシュボードを調整する必要も、2つ目のSDKも必要ない。

これは聞こえる以上に重要だ。このパイプラインでの15秒のショートは、少なくとも5回の画像生成と1回のビデオパスに触れる。それを3つのベンダーに分散すると、時間の半分は配管作業に費やされる。1つのキーにまとめれば、摩擦はほぼゼロになる。

キーの取得は約1分で完了する。コンソールを開き、API Keysに移動し、作成し、コピーする。その後、任意のHTTPクライアントをOpenAI互換の画像エンドポイントに向ければ、生成が始まる。

Seedream 5.0 15sワークフロー、顔からフィルムへの6ステップ

パイプラインは6ステップで、最初の2つだけが再利用可能なアセットを構築する。それ以降は構図とモーションだ。キャラクターとスタイルを前方にフィードし、ロックしておけば、モデルはドリフトを止める。以下は全体の概要をステップごとに説明する前に。

ステップロックするものモデル入力から出力へ
1 キャラクターカードアイデンティティ、衣装Seedream 5.0 Pro複数ビュー参照から5アングルのキャラクターシートへ
2 スタイルカードカラーグレード、ルックSeedream 5.0 Proフィルムスチルからパレットカラーカードへ
3 エスタブリッシングフレーム最初のショットSeedream 5.0 Proキャラクターカード+スタイルカードからフレーム1へ
4 深度マップ純粋なジオメトリSeedream 5.0 Proエスタブリッシングフレームからグレースケール深度パスへ
5 3x3ストーリーボード9つのカメラアングルSeedream 5.0 Proフレーム+深度マップから9パネル深度グリッドへ
6 ビデオモーション、連続性Seedance 2.0顔+フレーム+ストーリーボードから15秒シーケンスへ
  1. キャラクターカード:アイデンティティを固定する。 複数ビューの参照をSeedream 5.0 Proにフィードし、5つのアングルで同じ顔を保持するシートを取得する。後々の大きな痛手を避けるためのヒント:クリーンな正面の顔写真がすでにある場合、全身ビューの顔は削除すること。アングルごとに1つの顔があれば、ビデオモデルがショットの途中で見知らぬ人にドリフトするのを防げる。

女性が長い濃緑のコートを着た複数のアングル

アーキビストの5アングルキャラクターカード:全身、背面、正面顔写真、2つの3/4ビュー。

  1. スタイルカード:グレードを固定する。 モデルにルックを自由に設定させる代わりに、参照を取得する。好きな映画から数フレームを取得し、パレットをカラーカードに抽出し、それをスタイルアンカーとしてフィードバックする。ルックはここから来る、構図は来ない。私はSeedream 5.0 ProにBlade Runner 2049のスチルを数枚渡し、パレットを抽出させた。ロジャー・ディーキンスがその映画のために構築した、暖かな琥珀色が深いインクティールの影に溶け込むパレットだ(American Cinematographer、2017)。

4つの黄色いフィルムスチルの上に一致するカラーパレット

Blade Runner 2049のスチルから構築されたカラーカード、6色の琥珀からティールへのパレット。

6つのアースカラーのブラウンとゴールドのスウォッチ、16進コード付き

抽出された6色パレットストリップ:ダークブラウンから琥珀、ウォームグレーへ。

  1. エスタブリッシングフレーム。 キャラクターカードとスタイルカードを組み合わせて、シーンの最初のフレームを生成する。ここでステップ2のカラーリードが完全に現れる。

人物がランタンを持ち、そびえ立つ日光の差す図書館に立つ

エスタブリッシングショット:アーキビストは小さく、対称的な一点透視図法の通路の中央に配置されている。

Plain
1Cinematic establishing shot, composed in the style of Denis Villeneuve / Roger Deakins: rigorous central one-point-perspective symmetry, monumental sense of scale, strong negative space, quiet and reverent.
2
3SCENE: the interior of a colossal, seemingly infinite library-archive at cathedral scale. Towering vertical bookshelves recede symmetrically on both sides down a single central aisle toward a distant glowing warm vanishing point. Volumetric amber god-ray light shafts angle down through fine floating dust and a few drifting loose pages. Three clear depth layers: sharp foreground shelf edges near camera, the hero in the midground aisle, and an infinitely receding golden-hazed background.
4
5HERO, identity and wardrobe locked to image 1. Use image 1 ONLY as the identity and wardrobe reference; do NOT reproduce its reference-sheet layout or its gray studio background. The same exact woman: her face and features unchanged, floor-length deep ink-teal wool coat over a cream turtleneck, low ponytail. She holds a small lit brass oil lantern that casts a warm pool of light around her. She stands small and centered far down the deep aisle, seen from a low angle, dwarfed by the towering shelves.
6
7GRADE and OPTICS, match image 2: warm burnished amber-gold highlights rolling into deep desaturated ink-teal shadows, low contrast with softly lifted clean blacks; tall oval anamorphic bokeh, a faint horizontal lens flare off the brightest shaft, gentle Black Pro-Mist bloom on the highlights, fine 35mm Kodak film grain, 2.00 anamorphic widescreen framing.
8
9Cinematic 35mm film still, photoreal, ultra-detailed, monumental. Not CGI, not a 3D render, not a game-engine look.
  1. 深度マップ:Seedream 5.0 15sワークフローの核心。 そのフレームを純粋なグレースケール深度パスに変換する。白が近く、黒が遠く、スタイルは除去され、ジオメトリのみ。これで構図はそれ自体のレイヤーとなり、構造に触れることなく何度でもリスタイルできる。

背の高い本棚の間に立つ白いシルエットが暗い虚無に向かっている

エスタブリッシングフレームをグレースケール深度マップに変換:白い前景の本棚から黒い深度へとフェード。

Plain
1Reinterpret this image as a single-channel linear depth pass, a grayscale Z-depth map of the kind a 3D renderer writes from its depth buffer, or a LiDAR range image. Brightness encodes distance from camera only: pure white on the nearest visible surface, pure black at the farthest, one continuous monotonic ramp of mid-grays across everything between. Anchor the scale once to the whole frame so identical distances read as identical gray anywhere in the image; never per-region auto-contrast. Hold geometry exact: rounded volumes get smooth continuous gradients, overlapping objects break at crisp hard-edged occlusion boundaries, thin structures and silhouettes stay legible, connected surfaces keep stable values. Distance is the only variable: flat depth-driven gray with no albedo, no texture, no cast shadows, no directional light, no outlines, no ambient occlusion. Output the clean depth pass and nothing else.
  1. 3x3深度ストーリーボード。 画像が1枚だけでいいなら、これはスキップする。ビデオを作るなら、ストーリーボード全体が必要だ。9つのショットを事前に構築し、1つの深度規則と1つの連続性を共有し、それぞれが勝手な方向に行く9フレームではない。2つの画像をフィードする:カラーのエスタブリッシングフレーム(シーンに何があるか)と深度マップ(距離がどのようにエンコードされているか)。以下のプロンプトは、各パネルに異なるカメラを強制する。これにより、ビデオモデルが同じ中央通路を9回デフォルトするのを防ぐ。

巨大でシュールな図書館を探索する人物の9つのグレースケールパネル

同じアーカイブの異なるカメラアングルによる9つのグレースケール深度パスの3x3グリッド。

Plain
1You are a depth-map storyboard generator. Output ONE image: a clean 3x3 grid of nine sequential shots from one continuous scene, every panel a grayscale linear depth pass (white = nearest, black = farthest) and nothing else.
2
3INPUT, Image 1 is a centered symmetric establishing depth pass. Use it for TWO things: (a) the grayscale linear-depth CONVENTION and tonal range to match across ALL nine panels; (b) the EXACT composition to reproduce in panel 2. Read the character, wardrobe (floor-length coat, brass lantern) and the archive design language from it too.
4
5CRITICAL, COMPOSITIONAL VARIETY (top priority): every panel MUST use a DISTINCT camera, different shot size, height, azimuth, tilt. ONLY panel 2 may use the centered symmetric one-point-perspective aisle; ALL other panels are FORBIDDEN from using a centered symmetric aisle. Embrace cinematic framing: oblique corners, raking diagonal colonnades, worm's-eye verticals, high top-down angles, strongly off-center asymmetric framing, Dutch tilts, deep negative space.
6
7STORY (nine panels = one continuous event; a lone woman archivist in a long coat with a brass lantern discovers one book and reaches it; read left to right, top row, middle, bottom):
81. ESTABLISHING, high aerial angle near the vaulted ceiling looking obliquely down, the aisle running diagonally, hero tiny on the diagonal.
92. ARRIVAL, the centered symmetric establishing composition from Image 1.
103. DISCOVERY, extreme worm's-eye looking almost straight up a towering shelf toward one small target book high above.
114. REACTION, medium close-up, strongly off-center: hero's face on the right third looking up-left, slight Dutch tilt.
125. PREPARATION, an oblique corner composition, hero small at the turn reaching upward, a vast dark negative-space void filling one half.
136. INSERT DETAIL, extreme macro; the shelf runs as a steep diagonal, her hand and one book spine sharp in the near corner.
147. MAIN ACTION, oblique view down a colonnade of tall repeating vertical fins raking diagonally, the opened book held large in the near foreground.
158. CONSEQUENCE, high angle looking down as loose pages scatter and fall through several depth layers.
169. RESOLUTION, elevated oblique extreme-wide from a high corner, revealing the archive as a vast asymmetric structure.
17
18CONTINUITY: keep identical across panels, character identity, coat, lantern, hair, the shelf and architecture design, scene scale. ONLY camera and pose change.
19
20DEPTH RENDER: every panel a single-channel linear depth pass, grayscale Z-depth, brightness = distance only. Pure white nearest, pure black farthest, one continuous monotonic ramp; anchor scale once and apply identically to all nine. No color, no texture, no directional light, no cast shadows, no glow, no ambient occlusion, no depth-of-field blur, no grain.
21
22GRID: exactly nine panels, three equal rows and three equal columns, identical aspect ratio, thin uniform gutters. No captions, numbers, labels, arrows, borders, or color anywhere.
  1. ビデオ。 顔、エスタブリッシングフレーム、3x3ストーリーボードをSeedance 2.0にフィードすると、ストーリーボードからショットごとにシーケンスをレンダリングする。深度グリッドはここで実際に機能している。各パネルが距離をエンコードするため、モデルは視差をシミュレートし、カットごとにシーンをシャッフルする代わりに、オブジェクトの位置を安定して保持できる。

図書館の本棚の間から明るい光線が紙吹雪のように輝く

完成したフィルムのフレーム:アーカイブの真上からのショット、光が爆発しページが散乱する。

Plain
1A cinematic 15-second single continuous piece, 2.00 anamorphic widescreen, photoreal 35mm film look with fine Kodak grain, no AI gloss. SCENE: the interior of a colossal, seemingly infinite library-archive of towering bookshelves receding into warm darkness. VISUAL STYLE: warm burnished amber-gold key light rolling into deep desaturated ink-teal shadows, volumetric god-ray shafts through fine floating dust, a Villeneuve / Deakins epic look. DIRECTOR THESIS: a lone keeper walks the infinite archive of every story ever written, finds one glowing book and opens it, and light and worlds pour out of its pages.
2
3LOCKS:
4- Identity: image 1 is the SOLE source of the woman's face and identity, keep her face identical in every shot, never drift.
5- Wardrobe, first frame and grade: image 2 is the FIRST FRAME and the look, match its warm amber-and-ink-teal grade, anamorphic optics and 35mm grain across the whole film.
6- Composition and camera: image 3 is a 3x3 depth-map storyboard of nine shot compositions; drive the sequence shot by shot from image 3 IN ORDER (panel 1 through panel 9), each shot matching the framing, shot size and camera angle of its panel.
7- Continuity: same woman, same coat, same lantern, same archive architecture and scale throughout.
8- Negative locks: no hand morphing, consistent natural fingers, no face distortion, no duplicated or floating limbs, no text or captions anywhere.
9
10TIMELINE (drive from image 3; all light motivated by the lantern and the glowing book):
110-2s SHOT 1 (panel 1), high oblique aerial looking down, slow drift in.
122-3.5s SHOT 2 (panel 2), centered symmetric wide, slow push-in along the axis.
133.5-5s SHOT 3 (panel 3), extreme worm's-eye craning up toward one glowing book.
145-6.5s SHOT 4 (panel 4), off-center medium close-up, her eyes lifting in quiet awe.
156.5-8s SHOT 5 (panel 5), oblique corner, she reaches upward into deep negative space.
168-9.5s SHOT 6 (panel 6), tight macro, her fingers slide one book free.
179.5-11.5s SHOT 7, first-person POV, her hands hold the book open and warm light erupts toward the lens.
1811.5-13.5s SHOT 8 (panel 8), high angle, loose pages drift through the aisle around her.
1913.5-15s SHOT 9 (panel 9), elevated oblique extreme-wide, she stands tiny in the vast archive now glowing warm.

Seedream 5.0 15sワークフローの実際のコスト

ショート全体の生成にかかるコストは、コーヒー一杯分よりも安い。このパイプラインでは、1フレームも動く前に画像のスタックを生成する。キャラクターカード、スタイルカード、エスタブリッシングフレーム、深度マップ、9パネルのストーリーボード。つまり5回の画像生成と1回のビデオパスだ。ここで画像あたりの価格が丸め誤差でなくなり、自由に反復するか試行を制限するかを決定する。以下の数字はすべて、2026年7月下旬時点のAtlas Cloud モデル価格ページから確認済みだ。

ワークフロー内のアセットモデルティア定価現在(20%オフ)
キャラクターカードSeedream 5.0 Pro1.5K$0.05$0.04
スタイル/カラーカードSeedream 5.0 Pro1.5K$0.05$0.04
エスタブリッシングフレームSeedream 5.0 Pro2K$0.09$0.07
深度マップSeedream 5.0 Pro1.5K$0.05$0.04
3x3深度ストーリーボードSeedream 5.0 Pro2K$0.09$0.07
画像小計$0.32$0.25
15秒ビデオ(720p)Seedance 2.0720p$0.2419 / 秒秒単位

720pで$0.2419/秒の場合、15秒のシーケンス全体で約$3.63追加され、画像とモーションを合わせたショート全体の生成コストは$4以下になる。Seedance 2.0は、ビデオ入力を与えると$0.1486/秒に下がり、その参照-to-ビデオモードは最大9枚の参照画像を受け入れる。これはまさにストーリーボード駆動のシーケンスが求めるものだ。

画像の経済性は、私が二度見した部分だ。Seedream 5.0 Proのフレームは1.5Kティアで$0.045から始まり、高品質のGPT Image 2の約5分の1だ。GPT Image 2は1024ピクセルフレームで$0.21572、1536x1024で$0.16964。2Kティアでも、GPT Image 2が到達しない領域で、Seedream 5.0 Proは画像あたり$0.09で、より大きなフレームを出力しながら半分以下の価格だ。ショートあたり5枚、バッチで50枚生成すると、その差は急速に拡大する。

注意:Seedream 5.0 Proは現在20%の期間限定割引を実施中で、1.5Kフレームは実質$0.036、2Kフレームは$0.072になる。価格とプロモーションは変動するため、大きなバッチの予算を組む前にモデルページを確認すること。

Seedream 5.0ワークフローを単一ショートの先へ

深度とスタイルの分割は、映画的なフィルムだけのためではない。同じ規律(最初に構造を固定し、連続性をそれ自体のレイヤーとして扱う)により、Seedream 5.0 Proは一度に多くのパネルにわたって一貫したデザインシステムを保持できる。これは、1つのヒーローフレームではなく、互いに属する画像のセットが必要なあらゆる場所で役立つ。

建築家ル・コルビュジエの1907年から1965年までの生涯と作品のタイムライン

各カードに一貫したデザインシステムを持つ9パネルの伝記タイムラインレイアウト。

マルチパネルのナレッジカードやカルーセルも、9ショットのストーリーボードと同じ問題だ。セルごとにコンテンツが変わる中で、タイプ、スペーシング、カラー言語を同一に保つ。Seedream 5.0ファミリーは、構造化レイアウトとテキストレンダリングをうまく処理できるため、ここで特に強力だ。これは他のフロンティア画像モデルと互角に戦えるカテゴリの1つだ。グリッドを一度構築し、規則をロックしておけば、すべてのパネルが1つの作品として読める。

深度マップ、ストーリーボード、カラーカード。これらはすべて、実際にコミットする前にショットを考えるための方法に過ぎない。ツールはますますワイルドで安価になっている。それでもできないことは、あなたが実際に伝えたいストーリーを決めることだ。

よくある質問

Seedream 5.0 15sワークフローを一言で言うと?

ジオメトリとスタイルを分離する6ステップのパイプラインです。シーンの構図をグレースケール深度マップとして固定し、9つのカメラアングルを1つの3x3深度グリッドとしてストーリーボード化し、それに固定されたキャラクターとカラーグレードを加えてビデオモデルにフィードし、15秒のショートをレンダリングします。

このワークフローにStable DiffusionやComfyUIは必要ですか?

いいえ。深度マップのアイデアはControlNetツールから来ていますが、このワークフロー全体はホスト型APIモデルで実行されます。画像はすべてSeedream 5.0 Pro、ビデオはSeedance 2.0で生成し、どちらも1つのAPIキーで行うため、ローカルインストール、GPU、ノードグラフのメンテナンスは不要です。

Seedream 5.0 15sワークフローのショートあたりのコストは?

2026年7月下旬の定価で約$4です。5つの画像アセットはSeedream 5.0 Proで合計約$0.315、15秒の720pクリップはSeedance 2.0で$0.2419/秒で約$3.63追加されます。現在の20%割引で画像コストはさらに低くなります。必ずライブのモデルページで確認してください。

Seedream 5.0 15sワークフローはキャラクターを一貫して保てますか?

はい、それが最初のステップの目的です。5アングルのキャラクターカードを構築し、重複する顔を削除して各アングルに1つのクリーンな参照を残し、その画像をストーリーボードとビデオプロンプトの唯一のアイデンティティソースとしてロックします。深度ストーリーボードが構図を処理し、キャラクターカードが顔を処理します。

なぜ深度マップが線画や3Dホワイトモデルよりも優れているのですか?

線画はポーズを固定するが空間を持たないため、フレームはフラットなままです。3Dホワイトモデルはボリュームを追加するが、灰色のクレイ調の外観を最終画像に強制します。深度マップは距離のみ(白が近く、黒が遠く)をエンコードするため、正確なジオメトリをスタイルの混入ゼロで提供し、同じ構図を何度でもリスタイルできます。

最新モデル

ひとつのAPIで、あらゆるメディアAIを。

すべてのモデルを探索
Seedream 5.0 15秒ワークフロー: シネマティック・ショートを1つのAPIキーで作成