僅限兩週 | Seedream 5.0 Pro 立享 8 折!

一個API金鑰,一部15秒影片:完整的Seedream 5.0 15秒工作流程

一個基於深度圖的 Seedream 5.0 15 秒工作流程:將幾何與風格分離,分鏡九張,並使用一個 API 金鑰渲染一部短片。完整提示詞在內。

我在一個週末完成了這個作品:一位孤獨的守護者走進無垠的檔案庫,翻開一本書,光芒從書頁中傾瀉而出。十五秒。角色設計、分鏡、動態,全部由AI生成,全部來自一個API金鑰。沒有渲染農場,沒有三個軟體的管線,沒有外掛動物園。

關鍵不在於更聰明的提示詞,而是將工作拆成兩部分。多數人讓一個模型同時搞定顏色、光線、紋理和幾何,然後困惑為什麼結果看起來很平。這是Seedream 5.0 15秒工作流程的解決方案:先用深度圖鎖定幾何,再將風格疊加上去。以下是完整管線,包含我實際使用的每個提示詞。

關鍵要點

  • Seedream 5.0 15秒工作流程的核心操作:將幾何(深度圖)與風格(色卡)分離,這樣就能無限重新為鏡頭設定風格,而構圖不會崩壞。
  • 九個攝影機角度預先以一個3x3深度網格形式建立分鏡,讓影片模型保持連續性,而不是憑空創造九個無關的畫面。
  • 2026年7月下旬驗證的定價:Seedream 5.0 Pro圖片在1.5K等級約$0.045,約為GPT Image 2高品質價格的五分之一。圖片用Seedream 5.0 Pro,影片用Seedance 2.0,一個金鑰搞定。

為什麼Seedream 5.0 15秒工作流程要先做深度圖

幾何是影像模型最先放棄的東西。要求它一次性解決顏色、光線、材質和3D空間,空間就會崩塌,這也是為什麼許多AI藝術看起來像是一張漂亮的貼紙,內部沒有空間感。解決方法借用了擴散工具中一個常見的概念:交給模型一張控制影像,讓它在這個固定結構上建構(ControlNet,ICCV 2023)。

人們常用的控制影像有三種,它們之間的差距比看起來更大。以下是它們在這個工作流程中的比較。

控制影像編碼的內容為什麼有幫助在哪裡失敗
線稿邊緣、姿勢、構圖鎖定佈局和角色姿勢沒有深度資訊,畫面仍然顯得扁平
3D白色模型體積和輪廓真實的質量感過度控制:黏土渲染的外觀會滲入最終影像,無法去除
深度圖僅距離(白色近,黑色遠)準確的空間,零風格污染需要單獨的風格處理,而這正是重點

故事板:一個男人變身成巨型機器人對抗怪獸

線稿故事板:邊緣和姿勢被鎖定,但每個畫面都顯得扁平,沒有深度

線稿只鎖定姿勢,其他什麼都沒有。模型無法判斷近遠,所以畫面保持扁平。

故事板:一個男人變身成巨型機器人對抗怪獸

同一個故事板以3D白色模型渲染,灰色黏土外觀佔據了畫面

3D白色模型有體積,但過度控制。模型將那層灰色黏土皮膚視為圖片本身,之後你會在每一步都與塑膠渲染外觀搏鬥。

深度圖只編碼一件事:與攝影機的距離。白色是近,黑色是遠,之間是連續的灰色漸層(Depth Anything V2,2024)。這使得它比線稿更準確,因為它承載了真實的空間,並且比白色模型更乾淨,因為它沒有材質或風格會洩漏。在頂層疊加任何外觀,幾何都不會被污染。先鎖定乾淨的幾何,再設定風格。這個單一習慣就是讓Seedream 5.0 15秒工作流程其餘部分保持一致的關鍵。

一個API金鑰運行整個Seedream 5.0 15秒工作流程

這之所以能保持在一個週末專案而不變成系統整合的頭痛,是因為鏈中的每個模型都位於一個端點之後。圖片在Seedream 5.0 Pro上運行,影片在Seedance 2.0上運行,兩者都響應Atlas Cloud上的同一個API金鑰。無需在角色卡和最終渲染之間切換平台,無需調和兩個計費儀表板,無需第二個SDK。

這比聽起來更重要。一個15秒的短片在這個管線中至少會觸及五次圖片生成和一次影片處理。如果分散到三個供應商,一半的時間都會花在管線連接上。保持在一個金鑰下,摩擦幾乎為零。

獲取金鑰大約需要一分鐘:打開控制台,進入API金鑰,創建一個,複製它。然後將任何HTTP客戶端指向與OpenAI相容的圖片端點,你就可以開始生成了。

Seedream 5.0 15秒工作流程:從臉部到影片的六個步驟

這個管線有六個步驟,只有前兩個步驟建立可重複使用的資產。之後的一切都是構圖和動態。將角色和風格向前傳遞,保持鎖定,模型就不會偏移。以下是整個流程的概覽,然後我們再逐步進行。

步驟鎖定的內容模型輸入到輸出
1 角色卡身份、服裝Seedream 5.0 Pro多視角參考到五個角度的角色表
2 風格卡色調、外觀Seedream 5.0 Pro電影劇照到調色盤色卡
3 建立鏡頭第一個鏡頭Seedream 5.0 Pro角色卡加風格卡到第一個畫面
4 深度圖純幾何Seedream 5.0 Pro建立鏡頭到灰階深度通道
5 3x3故事板九個攝影機角度Seedream 5.0 Pro畫面加深度圖到九面板深度網格
6 影片動態、連續性Seedance 2.0臉部加畫面加故事板到15秒序列
  1. 角色卡:鎖定身份。 將多視角參考輸入Seedream 5.0 Pro,獲得一張在五個角度上保持相同臉孔的角色表。一個能省去很多後續麻煩的提示:如果你已經有乾淨的正面頭像,刪除全身視圖上的臉部。每個角度只有一張臉,可以防止影片模型在鏡頭中段偏離到一個陌生人。

一位穿著深綠色長外套的女性多個角度

檔案管理員的五角度角色卡:全身、背面、正面頭像,以及兩個四分之三視角

  1. 風格卡:鎖定色調。 不要讓模型自由發揮外觀,而是提取參考。從你喜歡的電影中抓取幾個畫面,提取調色盤到一張色卡,然後將它作為風格錨點回饋。外觀來自這裡,構圖則不來自這裡。我給了Seedream 5.0 Pro一些《銀翼殺手2049》的劇照,讓它提取調色盤——那種溫暖的琥珀色融入深墨青色陰影,正是Roger Deakins為這部電影打造的風格(American Cinematographer,2017)。

四張黃色電影劇照顯示在一個匹配的調色盤上方

從《銀翼殺手2049》劇照建立的色卡,包含六色琥珀到青色的調色盤

六個土褐色和金色色票,標有十六進位色碼

提取的六色調色盤條:深褐色到琥珀色再到暖灰色

  1. 建立鏡頭。 角色卡加風格卡給出場景的第一個畫面。這是步驟2中的顏色讀取完整呈現的地方。

一個人拿著燈籠站在高大、陽光普照的圖書館中

建立鏡頭:檔案管理員渺小居中,在對稱的單點透視走道中

Plain
1Cinematic establishing shot, composed in the style of Denis Villeneuve / Roger Deakins: rigorous central one-point-perspective symmetry, monumental sense of scale, strong negative space, quiet and reverent.
2
3SCENE: the interior of a colossal, seemingly infinite library-archive at cathedral scale. Towering vertical bookshelves recede symmetrically on both sides down a single central aisle toward a distant glowing warm vanishing point. Volumetric amber god-ray light shafts angle down through fine floating dust and a few drifting loose pages. Three clear depth layers: sharp foreground shelf edges near camera, the hero in the midground aisle, and an infinitely receding golden-hazed background.
4
5HERO, identity and wardrobe locked to image 1. Use image 1 ONLY as the identity and wardrobe reference; do NOT reproduce its reference-sheet layout or its gray studio background. The same exact woman: her face and features unchanged, floor-length deep ink-teal wool coat over a cream turtleneck, low ponytail. She holds a small lit brass oil lantern that casts a warm pool of light around her. She stands small and centered far down the deep aisle, seen from a low angle, dwarfed by the towering shelves.
6
7GRADE and OPTICS, match image 2: warm burnished amber-gold highlights rolling into deep desaturated ink-teal shadows, low contrast with softly lifted clean blacks; tall oval anamorphic bokeh, a faint horizontal lens flare off the brightest shaft, gentle Black Pro-Mist bloom on the highlights, fine 35mm Kodak film grain, 2.00 anamorphic widescreen framing.
8
9Cinematic 35mm film still, photoreal, ultra-detailed, monumental. Not CGI, not a 3D render, not a game-engine look.
  1. 深度圖:Seedream 5.0 15秒工作流程的關鍵。 將該畫面轉換為純灰階深度通道。白色近,黑色遠,風格被剝離,只有幾何。現在構圖是它自己的層,你可以隨意重新設定風格,而不會影響結構。

一個白色人影站在高書架之間,面對黑暗虛空

建立鏡頭轉換為灰階深度圖:白色前景書架漸變為黑色深度

Plain
1Reinterpret this image as a single-channel linear depth pass, a grayscale Z-depth map of the kind a 3D renderer writes from its depth buffer, or a LiDAR range image. Brightness encodes distance from camera only: pure white on the nearest visible surface, pure black at the farthest, one continuous monotonic ramp of mid-grays across everything between. Anchor the scale once to the whole frame so identical distances read as identical gray anywhere in the image; never per-region auto-contrast. Hold geometry exact: rounded volumes get smooth continuous gradients, overlapping objects break at crisp hard-edged occlusion boundaries, thin structures and silhouettes stay legible, connected surfaces keep stable values. Distance is the only variable: flat depth-driven gray with no albedo, no texture, no cast shadows, no directional light, no outlines, no ambient occlusion. Output the clean depth pass and nothing else.
  1. 3x3深度故事板。 只需要一張圖片?跳過這一步。要製作影片?你需要一個完整的 storyboard,預先建立九個鏡頭,共享一個深度規範和連續性,而不是九個各自為政的畫面。輸入兩張圖片:彩色建立鏡頭(場景中有什麼)和深度圖(距離如何編碼)。下面的提示強制每個面板使用不同的攝影機,這可以防止影片模型預設九次都使用相同的居中走道。

九個灰階面板,展示一個人在巨大的超現實圖書館中探索

一個3x3網格,包含九個灰階深度通道,每個都是同一個檔案庫的不同攝影機角度

Plain
1You are a depth-map storyboard generator. Output ONE image: a clean 3x3 grid of nine sequential shots from one continuous scene, every panel a grayscale linear depth pass (white = nearest, black = farthest) and nothing else.
2
3INPUT, Image 1 is a centered symmetric establishing depth pass. Use it for TWO things: (a) the grayscale linear-depth CONVENTION and tonal range to match across ALL nine panels; (b) the EXACT composition to reproduce in panel 2. Read the character, wardrobe (floor-length coat, brass lantern) and the archive design language from it too.
4
5CRITICAL, COMPOSITIONAL VARIETY (top priority): every panel MUST use a DISTINCT camera, different shot size, height, azimuth, tilt. ONLY panel 2 may use the centered symmetric one-point-perspective aisle; ALL other panels are FORBIDDEN from using a centered symmetric aisle. Embrace cinematic framing: oblique corners, raking diagonal colonnades, worm's-eye verticals, high top-down angles, strongly off-center asymmetric framing, Dutch tilts, deep negative space.
6
7STORY (nine panels = one continuous event; a lone woman archivist in a long coat with a brass lantern discovers one book and reaches it; read left to right, top row, middle, bottom):
81. ESTABLISHING, high aerial angle near the vaulted ceiling looking obliquely down, the aisle running diagonally, hero tiny on the diagonal.
92. ARRIVAL, the centered symmetric establishing composition from Image 1.
103. DISCOVERY, extreme worm's-eye looking almost straight up a towering shelf toward one small target book high above.
114. REACTION, medium close-up, strongly off-center: hero's face on the right third looking up-left, slight Dutch tilt.
125. PREPARATION, an oblique corner composition, hero small at the turn reaching upward, a vast dark negative-space void filling one half.
136. INSERT DETAIL, extreme macro; the shelf runs as a steep diagonal, her hand and one book spine sharp in the near corner.
147. MAIN ACTION, oblique view down a colonnade of tall repeating vertical fins raking diagonally, the opened book held large in the near foreground.
158. CONSEQUENCE, high angle looking down as loose pages scatter and fall through several depth layers.
169. RESOLUTION, elevated oblique extreme-wide from a high corner, revealing the archive as a vast asymmetric structure.
17
18CONTINUITY: keep identical across panels, character identity, coat, lantern, hair, the shelf and architecture design, scene scale. ONLY camera and pose change.
19
20DEPTH RENDER: every panel a single-channel linear depth pass, grayscale Z-depth, brightness = distance only. Pure white nearest, pure black farthest, one continuous monotonic ramp; anchor scale once and apply identically to all nine. No color, no texture, no directional light, no cast shadows, no glow, no ambient occlusion, no depth-of-field blur, no grain.
21
22GRID: exactly nine panels, three equal rows and three equal columns, identical aspect ratio, thin uniform gutters. No captions, numbers, labels, arrows, borders, or color anywhere.
  1. 影片。 將臉部、建立鏡頭和3x3故事板輸入Seedance 2.0,它會根據你的故事板逐鏡頭渲染序列。深度網格在這裡發揮了真正的作用:因為每個面板都編碼了距離,模型可以模擬視差,並在鏡頭之間保持物體位置穩定,而不是每次剪輯都重新排列場景。

發光的紙張從圖書館書架之間的一道明亮光束中爆炸開來

完成影片中的一個畫面:從上方俯瞰檔案庫,光芒爆發,紙張散落

Plain
1A cinematic 15-second single continuous piece, 2.00 anamorphic widescreen, photoreal 35mm film look with fine Kodak grain, no AI gloss. SCENE: the interior of a colossal, seemingly infinite library-archive of towering bookshelves receding into warm darkness. VISUAL STYLE: warm burnished amber-gold key light rolling into deep desaturated ink-teal shadows, volumetric god-ray shafts through fine floating dust, a Villeneuve / Deakins epic look. DIRECTOR THESIS: a lone keeper walks the infinite archive of every story ever written, finds one glowing book and opens it, and light and worlds pour out of its pages.
2
3LOCKS:
4- Identity: image 1 is the SOLE source of the woman's face and identity, keep her face identical in every shot, never drift.
5- Wardrobe, first frame and grade: image 2 is the FIRST FRAME and the look, match its warm amber-and-ink-teal grade, anamorphic optics and 35mm grain across the whole film.
6- Composition and camera: image 3 is a 3x3 depth-map storyboard of nine shot compositions; drive the sequence shot by shot from image 3 IN ORDER (panel 1 through panel 9), each shot matching the framing, shot size and camera angle of its panel.
7- Continuity: same woman, same coat, same lantern, same archive architecture and scale throughout.
8- Negative locks: no hand morphing, consistent natural fingers, no face distortion, no duplicated or floating limbs, no text or captions anywhere.
9
10TIMELINE (drive from image 3; all light motivated by the lantern and the glowing book):
110-2s SHOT 1 (panel 1), high oblique aerial looking down, slow drift in.
122-3.5s SHOT 2 (panel 2), centered symmetric wide, slow push-in along the axis.
133.5-5s SHOT 3 (panel 3), extreme worm's-eye craning up toward one glowing book.
145-6.5s SHOT 4 (panel 4), off-center medium close-up, her eyes lifting in quiet awe.
156.5-8s SHOT 5 (panel 5), oblique corner, she reaches upward into deep negative space.
168-9.5s SHOT 6 (panel 6), tight macro, her fingers slide one book free.
179.5-11.5s SHOT 7, first-person POV, her hands hold the book open and warm light erupts toward the lens.
1811.5-13.5s SHOT 8 (panel 8), high angle, loose pages drift through the aisle around her.
1913.5-15s SHOT 9 (panel 9), elevated oblique extreme-wide, she stands tiny in the vast archive now glowing warm.

Seedream 5.0 15秒工作流程的實際成本

整個短片生成成本比一杯咖啡還少。這個管線在一個畫面動起來之前會產生一堆圖片:角色卡、風格卡、建立鏡頭、深度圖,以及九面板故事板,所以是五次圖片生成加上一次影片處理。這就是為什麼每張圖片的價格不是一個四捨五入的誤差,而是決定你是自由迭代還是節省嘗試次數的關鍵。以下所有數字均來自Atlas Cloud 模型定價頁面,截至2026年7月下旬驗證。

工作流程中的資產模型等級定價現在(8折)
角色卡Seedream 5.0 Pro1.5K$0.05$0.04
風格/色卡Seedream 5.0 Pro1.5K$0.05$0.04
建立鏡頭Seedream 5.0 Pro2K$0.09$0.07
深度圖Seedream 5.0 Pro1.5K$0.05$0.04
3x3深度故事板Seedream 5.0 Pro2K$0.09$0.07
圖片小計$0.32$0.25
15秒影片(720p)Seedance 2.0720p$0.2419 / 秒每秒

在720p下,每秒$0.2419,完整的15秒序列大約增加$3.63,因此整個短片(圖片和動態一起)生成成本低於$4。Seedance 2.0在輸入影片時也會降至每秒$0.1486,而其參考到影片模式最多接受九張參考圖片,這正是故事板驅動序列所需要的。

圖片經濟學是讓我多看了一眼的部分。Seedream 5.0 Pro的畫面在1.5K等級起價$0.045,約為GPT Image 2高品質價格的五分之一(GPT Image 2高品質:1024像素畫面$0.21572,1536x1024為$0.16964)。即使在完整的2K等級(GPT Image 2達不到),Seedream 5.0 Pro每張圖片$0.09,仍然不到一半的價格,同時輸出更大的畫面。每部短片生成五張圖片,或一批生成五十張,這個差距會迅速累積。

注意時間點:Seedream 5.0 Pro目前正在進行20%的限時折扣,因此在促銷期間,1.5K畫面實際為$0.036,2K畫面為$0.072。價格和促銷會變動,所以在規劃大量批次前,請查看模型頁面。

將Seedream 5.0工作流程應用於單一短片之外

深度與風格的拆分不僅適用於電影製作。同樣的原則——先鎖定結構,將連續性視為獨立層——正是讓Seedream 5.0 Pro能夠一次在多個面板中保持一致設計系統的原因。這在需要一組屬於同一系列的圖片(而非單一主角畫面)時非常有用。

建築師勒·柯比意從1907年到1965年的生平與作品時間線

一個九面板傳記時間線佈局,每張卡片都有一致的設計系統

多面板知識卡片或輪播與九鏡頭故事板是同樣的問題:在內容逐格變化的同時,保持字體、間距和顏色語言完全相同。Seedream 5.0系列在這裡特別強大,因為它處理結構化佈局和文字渲染的能力很好,這正是它與其他前沿影像模型相抗衡的類別。一次建立網格,保持規範鎖定,每個面板都讀起來像是一個整體。

深度圖、故事板、色卡——所有這些其實只是一種在投入之前先思考鏡頭的方式。工具越來越強大,也越來越便宜。但它們仍然無法決定你真正想講述的故事是什麼。

常見問題

用一句話說明Seedream 5.0 15秒工作流程是什麼?

這是一個六步驟管線,將幾何與風格分離:你先將場景的構圖鎖定為灰階深度圖,將九個攝影機角度建立為單一3x3深度網格,然後將它加上鎖定的角色和色調,輸入影片模型,渲染出15秒的短片。

這個工作流程需要Stable Diffusion或ComfyUI嗎?

不需要。深度圖的概念來自ControlNet工具,但這個工作流程完全在託管API模型上運行。你所有的圖片都在Seedream 5.0 Pro上生成,影片在Seedance 2.0上生成,都透過一個API金鑰,因此無需本機安裝、無需GPU、無需維護節點圖。

Seedream 5.0 15秒工作流程每部短片花費多少?

截至2026年7月下旬,定價約為$4生成。五個圖片資產在Seedream 5.0 Pro上總計約$0.315,而15秒720p片段在Seedance 2.0上以每秒$0.2419增加約$3.63。目前有20%折扣進一步降低圖片成本。請務必在即時模型頁面上確認。

Seedream 5.0 15秒工作流程能保持角色一致嗎?

可以,而且這就是第一步的重點。你建立一個五角度角色卡,刪除重複的臉部,使每個角度只有一個乾淨的參考,然後將該圖像鎖定為故事板和影片提示中的唯一身份來源。深度故事板處理構圖,角色卡處理臉部。

為什麼深度圖在這裡比線稿或3D白色模型更好?

線稿鎖定姿勢但不承載空間,因此畫面保持扁平。3D白色模型增加體積,但會將灰色黏土外觀強加於最終影像。深度圖只編碼距離(白色近、黑色遠),因此提供準確的幾何,且零風格污染,讓你可以隨意為同一構圖重新設定風格。

最新模型

一個 API,暢享全模態 AI。

探索全部模型
Seedream 5.0 15s 工作流程:製作電影級短片,一個API金鑰