Wan 3.0 Multi-Shot Consistency: I Counted Every Cut in 110 Official Clips
You wrote a four-shot script. Shot 1 lands perfectly: straw hat, black beaded necklace, white linen dress, light exactly right. Shot 2 arrives and it is a different woman in a different hat.
That gap is the whole reason people are staring at Wan 3.0 multi-shot consistency, the promise of one prompt, six shots, thirty seconds, same character throughout.
So I stopped reading spec sheets. I downloaded all 110 demo files Alibaba published with its own Wan 3.0 creator handbook, ran ffmpeg over every one, parsed all 64 paired prompts, and then did the thing nobody in the search results has done: I counted the actual cuts in the delivered videos and compared them against the shot count each prompt asked for.
Three findings contradict what you will read everywhere else.

Cinematic colour-grading suite where a wall of monitors shows the same straw-hat character held across six different shots, the reference point for Wan 3.0 multi-shot consistency
A colourist's reference wall is the honest mental model for multi-shot consistency: one identity, six framings, checked side by side. Generated with openai/gpt-image-2.
Key takeaways
- 6 shots is real but not guaranteed: 3 of Alibaba's own multi-shot prompts hit their declared count exactly, 2 under-delivered.
- The worst case asked for 4 shots and rendered 2.
- No 4K. Zero of the 110 official files exceed 1080p, and only 4 are exactly 1920x1080.
- Props and camera drift before faces do. Watch the hat, not the eyes.
- No public Wan 3.0 API outside Alibaba's own platform yet, so the tutorial below rebuilds it on models you can call today.
Here is the best-case result, from Alibaba's own handbook. Six shots declared in the prompt, six shots delivered, same two characters, same wardrobe, same rain:

Six-shot anime sequence on a rainy train platform, the same girl with a transparent umbrella and boy with a khaki backpack held consistently across every cut
Alibaba's official Wan 3.0 beta demo (handbook clip v013, 1280x720, 20.05s). The prompt numbered shots 1 to 6; ffmpeg finds exactly 5 hard cuts, so all 6 shots landed. Not generated by Atlas Cloud.
What Wan 3.0 Multi-Shot Consistency Actually Is
Short version: it is three separate promises wearing one name, and Alibaba is unusually specific about which three.
In its launch write-up, Alibaba splits consistency into characters ("facial features, hairstyle and hair color, body shape, clothing, and accessories"), props ("multi-angle appearance, hardware structure, logos, and material details"), and space ("character blocking and camera perspective") (Alibaba Cloud, 2026). That three-layer split is the grading rubric for everything below. Same face, wrong necklace, is a fail.
The model generates up to 30 seconds in one job, at 480P, 720P or 1080P (Alibaba Cloud, 2026).
Now the part the search results get wrong.
| Common claim in search results | What the first-party material actually shows |
|---|---|
| "Native 4K generation" | The official post lists 480P / 720P / 1080P only. Across 110 official demo files, nothing exceeds 1080p, and just 4 files are exactly 1920x1080. The usual "1080p" landscape output is 1920x1072. |
| "Up to 6 independent shots per generation" | No shot-count spec appears in Alibaba's own launch post. Empirically it holds up: of 8 explicitly multi-shot prompts in the handbook, the highest any declares is 6, and two of those delivered 6. |
| "Up to 12 reference images" | Hands-on testing reports up to 10 images plus 5 video clips plus 5 audio clips (Curious Refuge, 2026). Alibaba's post gives no number at all. |
| "Open weights, Apache-2.0" | Earlier Wan releases were open-weight; Wan 3.0 ships as a platform service and no weights have appeared (Curious Refuge, 2026). |
| "The API is out" | It runs on Alibaba Cloud Model Studio. Third-party inference platforms do not have it yet. |
And this is the table that took the longest to build. Every number is mine, from ffmpeg scene detection on the delivered file versus the shot count written in the paired prompt:
| Official demo | Prompt declares | Hard cuts found | Shots delivered | Verdict |
|---|---|---|---|---|
| v013 rainy platform, anime | 6 shots | 5 | 6 | Exact |
| v055 golden retriever diary | 6 shots, 30.0s | 5 | 6 | Exact |
| v021 ramen shop and kitten | 4 shots | 3 | 4 | Exact |
| v008 forest lake, 3D | 4 shots | 2 | 3 | One short |
| v085 straw-hat woman | 4 shots | 1 | 2 | Half |
| v041 hallway to magic alley | 3 segments, "no hard cut" | 0 | 1 continuous take | Exactly as asked |
| v045, v084, v102 | 3 / 5 / multi-subject | Too many to resolve | Not countable | Fast cutting and ink-wash strobing defeat detection |
Read the middle rows again. Three prompts hit their number dead on. Two did not. The shot count you type is a request, not a contract.
Why Wan 3.0 Multi-Shot Consistency Breaks: 4 Failure Modes
Verdict first: it rarely breaks on the face. It breaks on the objects and the camera, and it breaks hardest when you change several things at once.
Here is the clip that taught me the most, because it is the worst performer in Alibaba's own showcase.
GtZy2TUK4ak
Alibaba's official Wan 3.0 beta demo (handbook clip v085, 1240x742, 30.08s). The prompt scripts four timecoded shots. ffmpeg finds one real cut, at 9.3s. Shots 3 and 4 collapse into a single held take that fades to black. Not generated by Atlas Cloud.
Failure mode 1: props drift before faces do. In that clip, look at the opening take alone, before any cut happens. At 0.6s the boater hat has a thin black band and the pendant is a faceted navy stone. At 9.2s the band is a wide black ribbon and the pendant is a smooth glossy black drop. Her face is stable the whole time. The accessories are not.

Two frames from the same continuous nine-second take, side by side, showing the hat band widening and the necklace pendant changing shape and colour
Both frames come from the same unbroken take of the official v085 demo, 8.6 seconds apart, no cut between them. Hat band and pendant both change. This is the props layer failing while the character layer holds.
That is why the acceptance test that actually works is a prop checklist, not a vibe check. For this character it is three items: hat shape and band, beaded necklace and pendant, dress neckline.
Failure mode 2: multi-subject scenes hold individually and smear on contact. Each character stays recognisable alone. Put them in frame together, interacting, and the identities bleed. Independent testing found the same thing: "Character consistency started to break down, and the compositing occasionally looked more like a bad green screen effect than a naturally generated environment," with lip sync called out as the single biggest weakness (Curious Refuge, 2026).
Failure mode 3: you changed four variables in one line. New location plus new camera angle plus new lighting plus new wardrobe state is the reliable way to lose an identity. The handbook's own prompts fight this by restating the entire costume in every segment. Across the 64 paired prompts, 9 explicitly say "stays unchanged", 5 pin the camera position, and 4 demand the character stay identical.
Failure mode 4: 30 seconds in one job is not 30 seconds of one shot. These are different products. A 30s multi-shot job cuts between framings. If what you want is one continuous unbroken take, chaining continuations degrades: identity error compounds at every handoff, so the clean single take is short by nature.
The handbook has exactly one prompt that gets seamlessness right, and it is worth copying the sentence structure:

Three prompt segments rendered as one unbroken continuous take, from an apartment hallway through a door into a lamplit gothic alley
Alibaba's official Wan 3.0 beta demo (handbook clip v041, 720x1280, 15.04s). Three prompt segments, and ffmpeg finds zero hard cuts, which is exactly what the prompt demanded. Not generated by Atlas Cloud.
Its handoff line, translated, is the most reusable thing in the entire handbook: continue seamlessly from the last frame of the previous segment; the character, the wardrobe, the location and the camera position stay completely identical; no hard cut and no abrupt scene transition. Only one prompt in 64 uses it. Steal it.
The Multi-Shot Consistency Workflow You Can Run Today
Question is: where do you actually run this?
Right now Wan 3.0 lives on Alibaba's own Model Studio. No third-party inference platform hosts it. On Atlas Cloud it appears only as a coming-soon entry with zero live endpoints, which is the honest state of play and worth saying plainly rather than pretending otherwise.
So you have two options: wait, or stop asking one prompt to do six things.
The second option is better anyway, and my cut-counting is the argument for it. A single multi-shot job gave you a shot count that missed by half in Alibaba's own showcase. Splitting the job puts that control back in your hands:
- Lock the look once, as a still image.
- Clone that identity into a keyframe per shot, changing exactly one variable each time.
- Animate each shot separately with the reference locked.
- Stitch locally.
Slower to set up, dramatically more predictable. And every model in the chain is callable today.
| Step | Model | Role in the chain | Listed price |
|---|---|---|---|
| 1 | bytedance/seedream-v5.0-pro/text-to-image | Lock the character look | $0.045 / image |
| 2 | google/nano-banana-2/edit | Clone identity into each shot's keyframe | $0.08 / image |
| 3 | alibaba/wan-2.7/reference-to-video | Animate each shot with the reference locked | $0.10 / second |
| Alt | alibaba/wan-2.7/text-to-video | One prompt, many shots (least controllable) | $0.10 / second |
| Fix | alibaba/wan-2.7/video-edit | Repair one wrong prop without regenerating | $0.10 / second |
Worth knowing for the "best model for multi-shot consistency" question: reference-to-video is a whole category now. bytedance/seedance-2.0-fast/reference-to-video runs $0.072/s (20% off $0.09, as of August 2026), kwaivgi/kling-video-o3-pro/reference-to-video $0.095/s (15% off $0.112), minimax/h3/reference-to-video $0.10/s, and google/veo3.1/reference-to-video $0.20/s. All prices verified against the live catalogue on 21 August 2026.
One warning before the steps: listed prices are a floor, not a quote. Resolution and duration move the real number, so read the Run button before you commit.
How to Build Wan 3.0-Style Multi-Shot Consistency, Step by Step
We are rebuilding the exact beat sheet that Alibaba's own model fumbled: the straw-hat woman, four shots, garden to car. Same script, per-shot control, and this time you get four shots because you rendered four.
Let's get started.
Step 1: lock the look. One image becomes the identity anchor for everything downstream. Generate it on Seedream v5.0 Pro, 16:9, highest resolution the page offers. Name the three checkpoint props explicitly, because those are what drift.
plaintext135mm film still, vintage sunlit rose garden. An elegant young woman wearing a 2natural straw boater hat with a black band, a black beaded necklace with a single 3teardrop pendant, and a white sleeveless linen dress. She stands before a wall of 4blooming pink and cream roses, one hand lightly touching the brim of her hat, 5gentle and thoughtful gaze just off camera. Soft warm natural light, dappled sun 6flare, shallow depth of field, fine film grain, medium close-up.

Seedream v5.0 Pro playground on Atlas Cloud with the straw-hat prompt entered and the finished character keyframe rendered in the output panel
Step 1 completed: the identity anchor for the whole four-shot sequence.
Step 2: clone the identity into shots 2 to 4. One keyframe per shot, one run each, using Nano Banana 2 Edit with the Step 1 output as the reference. The discipline that matters: restate the full wardrobe every single time, then change exactly one thing.
Shot 2, the emotional turn:
plaintext1Keep the exact same woman from the reference image: identical face, identical 2natural straw boater hat with black band, identical black beaded necklace with a 3teardrop pendant, identical white sleeveless linen dress. New shot: cinematic 4close-up, she slowly lifts the straw hat off her head and lowers her chin, a faint 5wistful expression. Same rose garden background. Strictly preserve the same 6lighting, skin tone, film grain and colour grade as the reference.
Shot 3, the cut to the car:
plaintext1Keep the exact same woman from the reference image: identical face, identical black 2beaded necklace with a teardrop pendant, identical white sleeveless linen dress, 3the straw boater hat now resting on her lap. New shot: back seat of a moving car, 4side profile, her hand propped against the door, gazing out a rain-beaded window, 5blurred green foliage rushing past. Overcast light with a warm interior fill, 6shallow depth of field, elegant melancholy. Strictly preserve the same face, the 7same colour grade and the same 35mm film grain as the reference.
Shot 4, the last look:
plaintext1Keep the exact same woman from the reference image: identical face, identical black 2beaded necklace with a teardrop pendant. New shot: extreme facial close-up beside 3the car window, eyes slowly closing, a calm and released expression. Raindrops 4slide down the glass in front of her, focus shifting onto the droplets, background 5falling into bokeh. Wong Kar-wai style cinematic mood, fine film grain. Strictly 6preserve the same face, the same lighting and the same colour grade as the 7reference.

Nano Banana 2 Edit playground on Atlas Cloud with the Step 1 keyframe loaded as reference and the shot 2 keyframe rendered, identity and wardrobe held
Step 2 completed: shot 2's keyframe, same woman, hat now in her hands.
Step 3: animate shot 1 with the reference locked. Now each shot becomes video independently on Wan 2.7 reference-to-video. Settings: 720P, 5 seconds, and drop the Step 1 keyframe into the Images slot. Check the Run button quote before you fire it.
plaintext1Shot 1 of 4. Reference image 1 defines the character: keep the same face, the same 2natural straw boater hat with black band, the same black beaded necklace with a 3teardrop pendant and the same white sleeveless linen dress identical for the entire 4clip. 35mm film look, vintage sunlit rose garden. The woman lightly touches the 5brim of her hat as a soft breeze lifts strands of her hair; the camera pushes in 6very slowly. Warm natural light, dappled sun flare, shallow depth of field, fine 7film grain. The camera position stays stable. No hard cut and no scene change.

Wan 2.7 reference-to-video playground on Atlas Cloud with the keyframe in the Images slot and the finished five-second clip playing in the output panel
Step 3 completed: the reference-locked clip rendered in the output panel.
And here is what came out:

Five-second reference-locked clip of the straw-hat woman in the rose garden, hat band and necklace holding steady throughout
Our own run on Atlas Cloud, not an Alibaba demo. Shown as a silent GIF ; the delivered file carries an audio track.
Step 4: chain shots 2 to 4, then stitch. Repeat Step 3 for each remaining keyframe, and paste the handoff sentence from the handbook at the top of every follow-on prompt:
plaintext1Continue seamlessly from the last frame of the previous segment. The character, 2the wardrobe, the location and the camera position stay completely identical. 3No hard cut and no abrupt scene transition.
Then concatenate locally with ffmpeg and run the three-point acceptance check: hat shape and band, beaded necklace and pendant, and whether the framing progression between shots actually makes sense. If exactly one prop is wrong in one shot, do not regenerate the shot. Send it through wan-2.7/video-edit and change only that, borrowing the handbook's phrasing: keep the actions, the wardrobe and the rest of the frame unchanged.
Variations: Multi-Subject Scenes and Style Locks
Three patterns from the handbook worth stealing.
Character cards for multi-subject shots. When two or more people share the frame, the official prompts assign lettered character cards and re-declare each one's full costume in every segment. Verbose, ugly, and it works better than pronouns.
Global style declaration for stylised work. Ink-wash, cel animation and other heavy styles need the style pinned once for the whole piece, then a timecoded beat sheet underneath it.

Ink-wash martial arts sequence with a bamboo-grove swordsman, style locked across a timecoded five-beat fight
Alibaba's official Wan 3.0 beta demo (handbook clip v084, 20.05s, cropped). Five timecoded beats under one global ink-wash style declaration. The strobing brushwork is why automated cut detection cannot count this one. Not generated by Atlas Cloud.
Fix, do not regenerate. A video-edit pass on one wrong prop is cheaper than a fresh render and cannot break the shots that were already right.
What a Four-Shot Sequence Costs
Concrete numbers for the sequence built above, at listed rates:
- 1 identity anchor on Seedream v5.0 Pro: $0.045
- 3 shot keyframes on Nano Banana 2 Edit: 3 x $0.08 = $0.24
- 4 clips at 5s, 720P, on Wan 2.7 reference-to-video: 20s x $0.10 = $2.00
- Total: about $2.29 for a 20-second, four-shot, identity-locked sequence
Stretch it to a six-shot 30-second piece and the video line becomes $3.00, landing near $3.44 all in.
Two honest caveats. First, listed prices are a floor: resolution tiers and duration change the real charge, and the playground's own Run quote is the only number that binds, which is why the screenshots above matter more than this list. Second, on Wan 3.0 itself, Alibaba prices Model Studio video generation per second by resolution tier, but I could not confirm the exact Wan 3.0 per-second rates from a first-party page, so I am not quoting figures I could not verify.
Worth noting what you are buying with the split-job approach: not a lower price, but a shot count that matches your script. On the evidence of Alibaba's own showcase, that is not something a single 30-second job can promise you yet, and it is the practical answer to Wan 3.0 multi-shot consistency until the API opens up.
Frequently Asked Questions
How many shots can Wan 3.0 keep consistent in one generation?
Six, on current evidence, with a caveat. Alibaba's launch post publishes no shot-count spec. Across the 8 explicitly multi-shot prompts in its official handbook, the highest any declares is 6, and two of those delivered exactly 6 cuts-verified shots. Treat 6 as an observed ceiling, not a guarantee.
Does Wan 3.0 really generate native 4K?
No. The official post lists 480P, 720P and 1080P only. Across all 110 official demo files nothing exceeds 1080p, only 4 files are exactly 1920x1080, and the common landscape output is 1920x1072. Frame rates split 24fps for 63 files and 30fps for 46.
Is the Wan 3.0 API available, and where can I run it?
It runs on Alibaba Cloud Model Studio. No third-party inference platform hosts it yet; Atlas Cloud lists it as a coming-soon entry with no live endpoints. If you need multi-shot output this week, the Wan 2.7 reference-to-video chain above is the working substitute.
Why do my characters still change between shots even with a reference image?
Usually because one prompt changed the location, the camera angle and the lighting simultaneously. Change one variable per shot, and restate the entire wardrobe in every prompt. Also check the right things: in Alibaba's own demo the face held steady while the hat band and the necklace pendant both changed inside a single unbroken take.
Are the Wan 3.0 weights open source?
No weights have been released. Earlier Wan versions were downloadable and runnable locally, whereas Wan 3.0 ships as a hosted platform service. Anyone citing an Apache-2.0 licence for Wan 3.0 is repeating something with no first-party source behind it.
What is the fastest way to get multi-shot consistency without Wan 3.0 access?
Four moves: lock one identity image, clone it into one keyframe per shot changing a single variable each time, animate each shot with reference-to-video, then stitch locally and check your props. Roughly $2.29 and about half an hour for a four-shot sequence.
Methodology: all 110 demo files and 64 paired prompts come from Alibaba's official Wan 3.0 creator handbook, archived locally on 11 August 2026. Metadata was read file by file with ffmpeg; shot counts come from ffmpeg scene-change detection at a 0.2 threshold, cross-checked against extracted frames; prompt structure was parsed programmatically. Atlas Cloud prices were verified against the live model catalogue on 21 August 2026.
Sources: Alibaba Cloud (August 2026); Curious Refuge (2026).






