I built this over a weekend: a lone keeper walks into an endless archive, opens a book, and light pours out of the pages. Fifteen seconds. Character design, storyboard, motion, all of it generated, all of it from a single API key. No render farm, no three-app pipeline, no plugin zoo.
The trick was not a cleverer prompt. It was splitting the job in two. Most people ask one model to nail color, light, texture, and geometry all at once, then wonder why the result looks flat. This is the Seedream 5.0 15s workflow that fixes that: lock the geometry first with a depth map, paint the style on afterward. Below is the whole pipeline, with every prompt I actually used.
Key takeaways
- The core move in the Seedream 5.0 15s workflow: separate geometry (a depth map) from style (a color card), so you can restyle a shot endlessly without the composition falling apart.
- Nine camera angles get storyboarded up front as one 3x3 depth grid, which keeps a video model on continuity instead of inventing nine unrelated frames.
- Verified late-July 2026 pricing: a Seedream 5.0 Pro image runs from $0.045 at the 1.5K tier, roughly one-fifth of GPT Image 2 at high quality. Images on Seedream 5.0 Pro, video on Seedance 2.0, one key.
Why Depth Maps Come First in the Seedream 5.0 15s Workflow
Geometry is the first thing an image model drops. Ask it to solve color, lighting, material, and 3D space in one pass and space is what collapses, which is why so much AI art reads as a pretty sticker with no room inside it. The fix borrows a well-worn idea from diffusion tooling: hand the model a control image and let it build on top of that fixed structure (ControlNet, ICCV 2023).
There are three control images people reach for, and the gap between them is bigger than it looks. Here is how they compare inside this workflow.
| Control image | What it encodes | Why it helps | Where it fails |
|---|---|---|---|
| Line art | Edges, pose, composition | Locks the layout and the character's pose | No depth information, so the frame still reads flat |
| 3D white model | Volume and silhouette | Real sense of mass | Over-controls: the clay-render look bleeds into the final image and will not scrub out |
| Depth map | Distance only (white near, black far) | Accurate space, zero style to contaminate | Needs a separate style pass, which is the whole point |

A line-art storyboard: edges and poses are locked, but every panel reads flat with no depth
Line art locks the pose, and nothing else. The model cannot tell what is near or far, so the frame stays flat.

The same storyboard as a 3D white-model render, where the gray clay look takes over the image
A 3D white model has volume, but it over-controls. The model treats that gray clay skin as the picture itself, and you end up fighting a plastic render look through every later step.
A depth map encodes exactly one thing: distance from the camera. White is near, black is far, one continuous ramp of gray in between (Depth Anything V2, 2024). That makes it more accurate than line art, because it carries real space, and cleaner than a white model, because it has no material and no style to leak. Drop any look on top and the geometry never gets contaminated. Lock the geometry clean, style it after. That single habit is what makes the rest of the Seedream 5.0 15s workflow hold together.
One API Key Runs the Entire Seedream 5.0 15s Workflow
The reason this stays a weekend project and not a systems-integration headache is that every model in the chain lives behind one endpoint. Images run on Seedream 5.0 Pro, video runs on Seedance 2.0, and both answer to the same API key on Atlas Cloud. No switching platforms between the character card and the final render, no reconciling two billing dashboards, no second SDK.
That matters more than it sounds. A single 15-second short in this pipeline touches at least five image generations and one video pass. Spread those across three vendors and half your time goes to plumbing. Keep them under one key and the friction is close to zero.
Getting the key takes about a minute: open the console, go to API Keys, create one, copy it. Then point any HTTP client at the OpenAI-compatible image endpoint and you are generating.
The Seedream 5.0 15s Workflow, Six Steps From Face to Film
The pipeline is six steps, and only the first two build reusable assets. Everything after that is composition and motion. Feed the character and style forward, keep them locked, and the model stops drifting. Here is the whole thing at a glance before we go step by step.
| Step | What it locks | Model | Input to output |
|---|---|---|---|
| 1 Character card | Identity, wardrobe | Seedream 5.0 Pro | Multi-view reference to a five-angle character sheet |
| 2 Style card | Color grade, look | Seedream 5.0 Pro | Film stills to a palette color card |
| 3 Establishing frame | The first shot | Seedream 5.0 Pro | Character card plus style card to frame one |
| 4 Depth map | Pure geometry | Seedream 5.0 Pro | Establishing frame to a grayscale depth pass |
| 5 3x3 storyboard | Nine camera angles | Seedream 5.0 Pro | Frame plus depth map to a nine-panel depth grid |
| 6 Video | Motion, continuity | Seedance 2.0 | Face plus frame plus storyboard to a 15s sequence |
- Character card: lock the identity. Feed a multi-view reference to Seedream 5.0 Pro and get a sheet that holds the same face across five angles. One tip that saves a lot of pain later: if you already have a clean front headshot, delete the face on the full-body view. One face per angle keeps the video model from drifting into a stranger halfway through a shot.

A five-angle character card of the archivist: full body, back, front headshot, and two three-quarter views
- Style card: lock the grade. Instead of letting the model freestyle the look, pull a reference. Grab a few frames from a film you like, extract the palette into a color card, and feed that back as the style anchor. Look comes from here, composition does not. I handed Seedream 5.0 Pro a few Blade Runner 2049 stills and let it pull the palette, that warm amber rolling into deep ink-teal shadow that Roger Deakins built for the film (American Cinematographer, 2017).

A color card built from Blade Runner 2049 stills, with a six-swatch amber-to-teal palette

The extracted six-color palette strip: dark brown through amber to warm gray
- Establishing frame. Character card plus style card gives you the first frame of the scene. This is where the color read from step 2 shows up in full.

The establishing shot: the archivist small and centered in a symmetric one-point-perspective aisle
Plain1Cinematic establishing shot, composed in the style of Denis Villeneuve / Roger Deakins: rigorous central one-point-perspective symmetry, monumental sense of scale, strong negative space, quiet and reverent. 2 3SCENE: the interior of a colossal, seemingly infinite library-archive at cathedral scale. Towering vertical bookshelves recede symmetrically on both sides down a single central aisle toward a distant glowing warm vanishing point. Volumetric amber god-ray light shafts angle down through fine floating dust and a few drifting loose pages. Three clear depth layers: sharp foreground shelf edges near camera, the hero in the midground aisle, and an infinitely receding golden-hazed background. 4 5HERO, identity and wardrobe locked to image 1. Use image 1 ONLY as the identity and wardrobe reference; do NOT reproduce its reference-sheet layout or its gray studio background. The same exact woman: her face and features unchanged, floor-length deep ink-teal wool coat over a cream turtleneck, low ponytail. She holds a small lit brass oil lantern that casts a warm pool of light around her. She stands small and centered far down the deep aisle, seen from a low angle, dwarfed by the towering shelves. 6 7GRADE and OPTICS, match image 2: warm burnished amber-gold highlights rolling into deep desaturated ink-teal shadows, low contrast with softly lifted clean blacks; tall oval anamorphic bokeh, a faint horizontal lens flare off the brightest shaft, gentle Black Pro-Mist bloom on the highlights, fine 35mm Kodak film grain, 2.00 anamorphic widescreen framing. 8 9Cinematic 35mm film still, photoreal, ultra-detailed, monumental. Not CGI, not a 3D render, not a game-engine look.
- Depth map: the crux of the Seedream 5.0 15s workflow. Convert that frame into a pure grayscale depth pass. White near, black far, style stripped, geometry only. Now the composition is its own layer, and you can restyle it as many times as you want without touching the structure.

The establishing frame converted to a grayscale depth map: white foreground shelves fading to black depth
Plain1Reinterpret this image as a single-channel linear depth pass, a grayscale Z-depth map of the kind a 3D renderer writes from its depth buffer, or a LiDAR range image. Brightness encodes distance from camera only: pure white on the nearest visible surface, pure black at the farthest, one continuous monotonic ramp of mid-grays across everything between. Anchor the scale once to the whole frame so identical distances read as identical gray anywhere in the image; never per-region auto-contrast. Hold geometry exact: rounded volumes get smooth continuous gradients, overlapping objects break at crisp hard-edged occlusion boundaries, thin structures and silhouettes stay legible, connected surfaces keep stable values. Distance is the only variable: flat depth-driven gray with no albedo, no texture, no cast shadows, no directional light, no outlines, no ambient occlusion. Output the clean depth pass and nothing else.
- The 3x3 depth storyboard. Only need one image? Skip this. Making a video? You need a whole storyboard, nine shots built up front that share one depth convention and one continuity, not nine frames that each go their own way. Feed two images: the color establishing frame (what is in the scene) and the depth map (how distance is encoded). The prompt below forces a distinct camera on every panel, which is what stops a video model from defaulting to the same centered aisle nine times.

A 3x3 grid of nine grayscale depth passes, each a different camera angle of the same archive
Plain1You are a depth-map storyboard generator. Output ONE image: a clean 3x3 grid of nine sequential shots from one continuous scene, every panel a grayscale linear depth pass (white = nearest, black = farthest) and nothing else. 2 3INPUT, Image 1 is a centered symmetric establishing depth pass. Use it for TWO things: (a) the grayscale linear-depth CONVENTION and tonal range to match across ALL nine panels; (b) the EXACT composition to reproduce in panel 2. Read the character, wardrobe (floor-length coat, brass lantern) and the archive design language from it too. 4 5CRITICAL, COMPOSITIONAL VARIETY (top priority): every panel MUST use a DISTINCT camera, different shot size, height, azimuth, tilt. ONLY panel 2 may use the centered symmetric one-point-perspective aisle; ALL other panels are FORBIDDEN from using a centered symmetric aisle. Embrace cinematic framing: oblique corners, raking diagonal colonnades, worm's-eye verticals, high top-down angles, strongly off-center asymmetric framing, Dutch tilts, deep negative space. 6 7STORY (nine panels = one continuous event; a lone woman archivist in a long coat with a brass lantern discovers one book and reaches it; read left to right, top row, middle, bottom): 81. ESTABLISHING, high aerial angle near the vaulted ceiling looking obliquely down, the aisle running diagonally, hero tiny on the diagonal. 92. ARRIVAL, the centered symmetric establishing composition from Image 1. 103. DISCOVERY, extreme worm's-eye looking almost straight up a towering shelf toward one small target book high above. 114. REACTION, medium close-up, strongly off-center: hero's face on the right third looking up-left, slight Dutch tilt. 125. PREPARATION, an oblique corner composition, hero small at the turn reaching upward, a vast dark negative-space void filling one half. 136. INSERT DETAIL, extreme macro; the shelf runs as a steep diagonal, her hand and one book spine sharp in the near corner. 147. MAIN ACTION, oblique view down a colonnade of tall repeating vertical fins raking diagonally, the opened book held large in the near foreground. 158. CONSEQUENCE, high angle looking down as loose pages scatter and fall through several depth layers. 169. RESOLUTION, elevated oblique extreme-wide from a high corner, revealing the archive as a vast asymmetric structure. 17 18CONTINUITY: keep identical across panels, character identity, coat, lantern, hair, the shelf and architecture design, scene scale. ONLY camera and pose change. 19 20DEPTH RENDER: every panel a single-channel linear depth pass, grayscale Z-depth, brightness = distance only. Pure white nearest, pure black farthest, one continuous monotonic ramp; anchor scale once and apply identically to all nine. No color, no texture, no directional light, no cast shadows, no glow, no ambient occlusion, no depth-of-field blur, no grain. 21 22GRID: exactly nine panels, three equal rows and three equal columns, identical aspect ratio, thin uniform gutters. No captions, numbers, labels, arrows, borders, or color anywhere.
- Video. Feed the face, the establishing frame, and the 3x3 storyboard to Seedance 2.0, and it renders the sequence shot by shot from your storyboard. The depth grid is doing real work here: because each panel encodes distance, the model can simulate parallax and hold object positions steady from one shot to the next instead of reshuffling the scene every cut.

A frame from the finished film: a top-down shot of the archive as light bursts and pages scatter
Plain1A cinematic 15-second single continuous piece, 2.00 anamorphic widescreen, photoreal 35mm film look with fine Kodak grain, no AI gloss. SCENE: the interior of a colossal, seemingly infinite library-archive of towering bookshelves receding into warm darkness. VISUAL STYLE: warm burnished amber-gold key light rolling into deep desaturated ink-teal shadows, volumetric god-ray shafts through fine floating dust, a Villeneuve / Deakins epic look. DIRECTOR THESIS: a lone keeper walks the infinite archive of every story ever written, finds one glowing book and opens it, and light and worlds pour out of its pages. 2 3LOCKS: 4- Identity: image 1 is the SOLE source of the woman's face and identity, keep her face identical in every shot, never drift. 5- Wardrobe, first frame and grade: image 2 is the FIRST FRAME and the look, match its warm amber-and-ink-teal grade, anamorphic optics and 35mm grain across the whole film. 6- Composition and camera: image 3 is a 3x3 depth-map storyboard of nine shot compositions; drive the sequence shot by shot from image 3 IN ORDER (panel 1 through panel 9), each shot matching the framing, shot size and camera angle of its panel. 7- Continuity: same woman, same coat, same lantern, same archive architecture and scale throughout. 8- Negative locks: no hand morphing, consistent natural fingers, no face distortion, no duplicated or floating limbs, no text or captions anywhere. 9 10TIMELINE (drive from image 3; all light motivated by the lantern and the glowing book): 110-2s SHOT 1 (panel 1), high oblique aerial looking down, slow drift in. 122-3.5s SHOT 2 (panel 2), centered symmetric wide, slow push-in along the axis. 133.5-5s SHOT 3 (panel 3), extreme worm's-eye craning up toward one glowing book. 145-6.5s SHOT 4 (panel 4), off-center medium close-up, her eyes lifting in quiet awe. 156.5-8s SHOT 5 (panel 5), oblique corner, she reaches upward into deep negative space. 168-9.5s SHOT 6 (panel 6), tight macro, her fingers slide one book free. 179.5-11.5s SHOT 7, first-person POV, her hands hold the book open and warm light erupts toward the lens. 1811.5-13.5s SHOT 8 (panel 8), high angle, loose pages drift through the aisle around her. 1913.5-15s SHOT 9 (panel 9), elevated oblique extreme-wide, she stands tiny in the vast archive now glowing warm.
What the Seedream 5.0 15s Workflow Actually Costs
The whole short costs less than a coffee to generate. This pipeline produces a stack of images before a single frame moves: character card, style card, establishing frame, depth map, and the nine-panel storyboard, so five image generations plus one video pass. That is where per-image price stops being a rounding error and starts deciding whether you iterate freely or ration your attempts. All figures below are verified from the Atlas Cloud model pricing pages as of late July 2026.
| Asset in the workflow | Model | Tier | List price | Now (20% off) |
|---|---|---|---|---|
| Character card | Seedream 5.0 Pro | 1.5K | $0.05 | $0.04 |
| Style / color card | Seedream 5.0 Pro | 1.5K | $0.05 | $0.04 |
| Establishing frame | Seedream 5.0 Pro | 2K | $0.09 | $0.07 |
| Depth map | Seedream 5.0 Pro | 1.5K | $0.05 | $0.04 |
| 3x3 depth storyboard | Seedream 5.0 Pro | 2K | $0.09 | $0.07 |
| Image subtotal | $0.32 | $0.25 | ||
| 15s video (720p) | Seedance 2.0 | 720p | $0.2419 / second | per second |
At $0.2419 per second for 720p, a full 15-second sequence adds roughly $3.63, so the entire short lands under $4 to generate, images and motion together. Seedance 2.0 also drops to $0.1486 per second when you feed it a video input, and its reference-to-video mode accepts up to nine reference images, which is exactly what a storyboard-driven sequence wants.
The image economics are the part that made me look twice. A Seedream 5.0 Pro frame starts at $0.045 at the 1.5K tier, about one-fifth of GPT Image 2 at high quality, which runs $0.21572 for a 1024 pixel frame and $0.16964 at 1536 by 1024. Even at the full 2K tier, where GPT Image 2 does not reach, Seedream 5.0 Pro is $0.09 per image, still under half the price while outputting a larger frame. Generate five images per short, or fifty across a batch, and that gap compounds fast.
Note the timing: Seedream 5.0 Pro is running a 20 percent limited-time discount right now, so the 1.5K frame is effectively $0.036 and the 2K frame $0.072 through the promo window. Prices and promos move, so check the models page before you budget a big batch.
Taking the Seedream 5.0 Workflow Past a Single Short
The depth-and-style split is not only for cinematic film. The same discipline, lock structure first and treat continuity as its own layer, is what lets Seedream 5.0 Pro hold a consistent design system across many panels at once. That is useful anywhere you need a set of images that belong together rather than one hero frame.

A nine-panel biographical timeline layout with a consistent design system across every card
A multi-panel knowledge card or a carousel is the same problem as a nine-shot storyboard: keep the type, spacing, and color language identical while the content changes cell to cell. The Seedream 5.0 family is unusually strong here because it handles structured layout and text rendering well, which is the one category where it holds its own against the other frontier image models. Build the grid once, keep the convention locked, and every panel reads as one piece.
Depth maps, storyboards, color cards, all of it is really just a way to think a shot through before you commit. The tools keep getting wilder and cheaper. What they still cannot do is decide what story you actually want to tell.
Frequently Asked Questions
What is the Seedream 5.0 15s workflow in one sentence?
It is a six-step pipeline that separates geometry from style: you lock a scene's composition as a grayscale depth map, storyboard nine camera angles as a single 3x3 depth grid, then feed that plus a locked character and color grade to a video model to render a 15-second short.
Do I need Stable Diffusion or ComfyUI for this workflow?
No. The depth-map idea comes from ControlNet tooling, but this whole workflow runs on hosted API models. You generate every image on Seedream 5.0 Pro and the video on Seedance 2.0, both through one API key, so there is no local install, no GPU, and no node graph to maintain.
How much does the Seedream 5.0 15s workflow cost per short?
Around $4 to generate at list price in late July 2026. The five image assets total about $0.315 on Seedream 5.0 Pro, and a 15-second 720p clip on Seedance 2.0 adds roughly $3.63 at $0.2419 per second. A current 20 percent discount lowers the image cost further. Always confirm on the live models page.
Can the Seedream 5.0 15s workflow keep one character consistent?
Yes, and that is the point of the first step. You build a five-angle character card, delete duplicate faces so each angle keeps one clean reference, then lock that image as the sole identity source in the storyboard and video prompts. The depth storyboard handles composition while the character card handles the face.
Why are depth maps better than line art or a 3D white model here?
Line art locks pose but carries no space, so frames stay flat. A 3D white model adds volume but forces a gray clay look into the final image. A depth map encodes only distance, white near and black far, so it gives accurate geometry with zero style to contaminate, which lets you restyle the same composition as many times as you like.






