Seedance 2.5 Now Live — First on Atlas Cloud

How to Write Wan 3.0 Prompts: What Three Official Cases Teach That No Parameter List Can

How to write Wan 3.0 prompts, reverse-engineered from official 30-second cases: the three-layer structure, performance writing, sound syntax, and motion control.

The Wan line is already the most-forked open video family on the internet. The Wan 2.2 and Wan 2.1 repositories sit at roughly 17,000 stars each (GitHub, August 2026), and the Wan-AI page moves hundreds of thousands of downloads a month across its checkpoints (Hugging Face, August 2026). So when Wan 3.0 entered early access with native 30-second generation, up to 20 reference inputs, and audio generated in the same pass as the picture, the first question was not whether people would use it. It was whether anyone would know how to write for it.

Because a 30-second single generation is not a bigger version of a 5-second clip. It is a different writing problem. A 5-second prompt describes a picture that moves. A 30-second prompt directs time: who changes, when, on which cut, with what sound under it.

The Wan 3.0 creator handbook circulating with early access ships nineteen official prompt-and-result cases. I studied all nineteen, then picked the three that teach the most transferable technique and pulled each one apart: what the prompt actually controls, why it is built the way it is, and how to reuse its skeleton for your own shots. Every video below is the real, unedited output for the case being discussed. The original case prompts are written in Chinese; the skeletons here are English restatements of their structure, and the Wan line accepts prompts in both languages.

Key takeaways

  • Long Wan 3.0 prompts share one architecture: reference bindings first, global rules second, a timestamped shot list third. Short test prompts can drop layers, production prompts should not.
  • Write performance as a chain of visible actions, never as an emotion label. "Anxious" is a hope. "Chewing slows, brows draw in, gaze drops to her hands" is an instruction.
  • Sound is part of the prompt, not a post step. Dialogue sits in curly braces, effects and music get their own sentences, and silence needs to be requested explicitly.
  • The official cases repeat their most important constraint and ban the failure they expect. One handbook prompt forbids "smooth glide" in three different phrasings, because the model defaults to smooth.
  • Motion can be quantified: frames per burst, percent of frame width per dash, number of repetitions. The dragonfly case below is the proof.

One generation, 30 seconds, ten numbered storyboard beats and a full score arc written into a single prompt: a fishing boat, a storm, and a sea monster that surfaces on schedule. Official Wan 3.0 handbook case, shown here unedited.

What Makes Writing a Wan 3.0 Prompt Different

Three capabilities change the job, and each one adds a writing skill that shorter models never asked for.

Native 30 seconds means structure. When a model renders 5 seconds, a prompt can be one dense sentence. At 30 seconds, an unstructured prompt drifts: characters swap clothes, the camera forgets its rule, the ending has nothing to do with the beginning. Every long case in the handbook fights drift the same way, with explicit stages and end states. That structure is the core of this guide.

Joint audio means sound direction. Wan 3.0 generates the soundtrack with the picture. Dialogue, effects, ambience and music are all promptable, which also means an unwritten soundtrack is an improvised one. The official cases treat audio with the same discipline as light: stated, scoped, and sometimes deliberately silenced.

Up to 20 references means binding syntax. Faces, outfits, locations, voices and even storyboard images can be pinned to uploaded materials. The handbook cases bind each reference to a named subject once, then reuse the name in every later shot. If you have tested the same discipline on the current generation, say on Wan 2.7 reference-to-video, the 3.0 version will feel familiar rather than new.

None of this requires writing a novel every time. The handbook's shortest good case is a single sentence about a UFO stealing a cow in a hand-painted style. The skill is knowing which layers your shot needs, which is exactly what the three cases below are for.

The Wan 3.0 Prompt Stack: Three Layers Every Long Case Shares

Strip the nineteen official cases to their bones and the long ones are all the same three-layer stack.

Layer 1: reference bindings. One line per uploaded material, stating what it is and what may be taken from it. Define the woman in image 1 as Subject 1. Take only the outfit from image 2. Voice for Subject 2 follows audio 1. Bind once, then refer by name forever after. Vague plurals like "the images define the characters" are exactly what this layer exists to prevent.

Layer 2: global rules. Everything that must stay true for the whole clip, written before any action: visual style and its reference points, a named color system, light behavior, camera grammar, performance discipline, audio discipline, and a list of banned failures. This layer is where the handbook cases are most aggressive. One martial-arts case caps slow motion at two uses of under a second each, bans on-screen subtitles, bans background music entirely, and forbids the camera from sitting still for more than 1.5 seconds.

Layer 3: the timeline. Timestamped or numbered beats, each carrying one main event and a visible end state. Not "they fight for a while" but "0 to 1.5 seconds: he stands at the reed line, backlit, wind in his coat, and says his one line." End states are what keep second fifteen consistent with second five.

A useful way to hold it in your head: Layer 1 is casting, Layer 2 is the rulebook, Layer 3 is the shot list. The three cases that follow each stress one layer hardest.

Case 1: How to Write a Wan 3.0 Prompt for Performance, One Face for 30 Seconds

The least flashy case in the handbook is the one most people should study first: a single fixed close-up of a young woman in a blue graduation gown, held for a full 30 seconds while she talks, worries, and finally looks down at her hands.

Thirty seconds, one shot, no cut. The gown tells you the scene, the brow tells you the stakes, and the gaze drop lands exactly where the prompt put it. Official handbook case, unedited output.

Here is what the official prompt does, move by move, in my words rather than its own:

  1. Camera before subject. It opens by locking the frame: frontal close-up, head and upper body, and one deliberate imperfection, a second person on the left edge who is fully defocused. Backgrounds you name and then blur are backgrounds the model will not sharpen mid-shot.
  2. Wardrobe as exposition. One blue graduation gown replaces a paragraph of scene-setting. The model infers the ceremony, the crowd, the day. Pick the one garment or prop that implies the world, and skip describing the world.
  3. Emotion as a chain of visible acts. The prompt never asks for "anxiety." It writes the mouth moving as she speaks, the brow slightly knit, the serious set of the face. Each item is something a camera can verify. Emotion labels are hopes. Muscle movements are instructions.
  4. The micro-move that sells realism. Two details do more than everything else combined: the camera holds a faint handheld breathing sway, and her gaze travels from ahead to down at her hands, on a stated trajectory. Give any long take one slow, deliberate change like that and the shot reads as directed rather than generated.

Reusable skeleton, in English, structure identical to the case:

Plain
1Frontal [shot size] on [subject], [framing detail]. A [defocused element] sits at the
2[edge] of frame. [Subject] wears [one garment that implies the whole scene].
3[Subject] is [action], mouth moving, [brow / eyes / jaw detail]. Light falls from
4[direction], keeping [skin / fabric texture note]. The camera is locked but carries a
5faint handheld breathing sway. Across the shot, [one slow visible change: a gaze
6trajectory, a posture shift, an object lowered].

Fill that template with a subject of your own before touching the 30-second epics. If the face holds for thirty seconds, your structure is sound.

Case 2: How to Structure a 30-Second Wan 3.0 Prompt Like a Director

The handbook's absurdist short is the best structural teaching object in the whole document: a cowgirl rides a real horse into a lonely Route 66 gas station, picks up the pump, and fills her horse with fresh grass while the attendant's worldview quietly collapses.

Thirty seconds of deadpan setup and one perfect sight gag. The comedy is written into the structure: fixed observational framing, then a sudden extreme close-up exactly when the absurd thing happens. Official handbook case, unedited output.

Its prompt is organized as five named modules, and the module list itself is the lesson:

  1. A logline. Two sentences of what happens and what the joke is. Not visuals, story. Writing this first forces you to know what the 30 seconds are for.
  2. A named color system. Not "colorful" but a two-color law: teal sky and orange earth, then every major object assigned to one side of it, the faded red pump, the orange scarf, the terracotta mountains. Name the system, then distribute it. Consistency in AI video is usually a palette problem before it is a character problem.
  3. A cast list with narrative jobs. Each character gets two or three visible traits and, crucially, a function: the attendant is explicitly designated as the reaction-shot carrier of the film. Assigning a job tells the model where the camera's attention belongs in every beat that character appears.
  4. A named camera philosophy. The case invents a rule and defines it in one line: calm observation plus absurd close-up, meaning fixed mid-shots and slow pans as the default, broken by a sudden extreme close-up only when the absurd event fires. Naming your camera rule and then referring to it beats re-describing camera moves in every beat. It is also the entire comedic mechanism: the flat gaze makes the grass-pump close-up land.
  5. A four-act timeline with end states. Setup, escalation, punchline, aftermath, each with a time range, one main event, and a frozen final image, down to the toothpick finally dropping from the attendant's mouth in the last beat.

The transferable model is bigger than comedy. Logline, palette law, cast with jobs, one named camera rule, four acts with end states: that is a production brief, and Wan 3.0 is the first generation of this family with enough runway, thirty seconds of it, for a brief to matter. The sea-monster showcase at the top of this article is the same skeleton at ten beats with a scored music arc, written the same way: each beat one event, each beat a visible end state.

Case 3: Writing Motion Physics Into a Wan 3.0 Prompt

The hardest thing to get from any video model is motion that is not smooth. Models average toward gentle, continuous movement, which is why every AI fairy floats like a slow balloon. The handbook's fairy does not float. She moves like a dragonfly: dead hover, instant dash, dead stop, new direction.

Tiny fairy with pink hair and wings hovers over carpet

Hover, dash, hard stop. The blur in the middle of each burst is one or two frames wide, exactly as the prompt budgeted. Official handbook case, shown as a silent GIF; the delivered clip carries its own ambience.

This prompt wins with four techniques you can lift directly:

  1. Priority declaration up top. The first line announces the single highest-priority requirement, the dragonfly movement law, before any scenery. Attention is a budget. Spend the opening tokens on the rule that decides success, not on the carpet color.
  2. A real-world physics anchor. Instead of adjectives like "fast and agile," the prompt names an animal whose motion signature the model has seen thousands of times: point hover, burst launch, full stop, sharp retarget. One familiar anchor outperforms ten abstract adverbs.
  3. Numbers where adjectives would blur. The dash must cross 60 to 90 percent of the frame width within 3 to 5 frames. The stop must hold for 2 to 4 frames. At least six clearly separated bursts must occur in the first eight seconds. Residual motion blur is budgeted at 1 to 2 frames. You are not writing poetry at this point, you are writing a spec, and the model honors it. Look at the strip below: the mid-dash frame is pure streak, the frames on either side are pin sharp.
  4. Banning the default. The prompt explicitly forbids the failure it predicts, in multiple phrasings: no slow fairy floating, no uniform-speed cruising, no continuous smooth gliding, and then a corrective sentence that says what the motion is not, "not a continuous flight path, but short, sharp, almost frame-skipping displacements." When you know what the model will do on autopilot, write the autopilot out of the shot.

Pink-haired toy fairy flying among teddy bears and colorful blocks

Four consecutive stills from the dragonfly case: hover beside the teddy bear, mid-dash motion blur, arrival above the blocks, dead-stop hover

Four frames from the output above. Sharp, streaked, sharp, sharp: the blur budget from the prompt, visible frame by frame.

The same four moves generalize to any motion problem: whip pans that must not become cuts, impacts that need weight, mechanical movement that must not turn organic. Anchor, quantify, prioritize, and ban the default.

Dialogue, Sound, and Reference Syntax in Wan 3.0 Prompts

The three cases above are silent on the mechanics that make Wan 3.0 prompts talk, so here is the syntax layer, compiled from the rest of the handbook's cases.

Dialogue goes in curly braces. Spoken lines are wrapped as {Which muscle group today} rather than left in open prose, which keeps the model from narrating your dialogue as a caption. Lines land with lip sync, and a dialogue scene is written as alternating action and braced lines, exactly like a screenplay.

Voices can be cast. The reference syntax extends to audio: define the woman in image 1 as Subject 1, then state that Subject 1's voice follows audio 1. Face from one file, voice from another, both locked for the duration.

Audio needs a discipline statement. The strongest cases declare their soundscape the way they declare their palette. An animation case states its audio types up front: ambience and music only, no speech. The martial-arts case goes further, banning background music outright, keeping only breathing, cloth friction, strikes and wind, and it gets exactly that. Music, when wanted, is written as an arc across the timeline, from low dread drones through string escalation to a brass impact and a long doom note, mapped to the beats it must hit.

Silence is a request too. Unprompted, the model fills the track. If your shot needs only room tone and one sound event, say so. Sound direction in Wan 3.0 is not garnish. It is the second half of the medium.

Where to Practice the Wan 3.0 Prompt Formula Today

Wan 3.0 is still on the way, and users cannot use it right now. The structure above still compounds now, because nothing in it is version-locked: bindings, global rules, timelines, braced dialogue and banned defaults all run on the current Wan generation.

The whole current family is live on Atlas Cloud, Wan 2.7 text-to-video, image-to-video, reference-to-video and video editing, on one key with per-run pricing shown before you press run. A practical drill: take the Case 1 skeleton, fill it with your own subject, and run it as a short take on Wan 2.7 text-to-video. If the performance chain holds there, the same file is your Wan 3.0 draft on day one, with the duration dialed up instead of rewritten. Atlas Cloud has a habit of carrying Day-0 access when this family updates, so the practice environment and the destination are likely to be the same page.

Frequently Asked Questions

How long should a Wan 3.0 prompt be?

As long as the shot needs and no longer. The official handbook ranges from one sentence, a hand-painted UFO gag, to prompts over a thousand words for a ten-beat monster film. The deciding variable is duration and cast size: a 5-second single-subject test needs one dense sentence, a 30-second multi-character story needs all three layers, bindings, global rules and a timeline.

Does Wan 3.0 generate sound from the prompt?

Yes, audio is generated in the same pass as the picture. Dialogue is written in curly braces and lands with lip sync, sound effects and ambience are written as prose, and music can be directed as an arc across the timeline. The discipline cuts both ways: if you want near-silence, or ambience with no score, you have to say so, or the model will fill the track on its own.

What is the curly-brace syntax in Wan 3.0 prompts?

Braces mark words that are spoken aloud by a character, as in {See you at the finish line}. Keeping dialogue inside braces separates it from scene description, so the model performs the line instead of illustrating it. For multilingual work, state the language before the line and the delivery style with it.

How do I keep one character consistent across a 30-second Wan 3.0 video?

Bind the character to a reference once, with scope: define the person in image 1 as Subject 1, take only face, hair and outfit. Then repeat the subject's name in every timeline beat they appear in, and restate the two or three locked traits at any high-risk moment such as a costume-adjacent action or a lighting change. The handbook's fight-scene case even restates trait locks per module, which is the belt-and-suspenders version.

Can I practice Wan 3.0 prompt writing before I have access?

Yes, and you should. The three-layer structure, braced dialogue, reference bindings and banned-default technique all run on the current family today. Wan 2.7 on Atlas Cloud covers text-to-video, image-to-video, reference-to-video and video editing on one key, so a template proven there transfers to 3.0 as a duration change rather than a rewrite.

Conclusion

Nineteen official cases say the same thing three different ways: a Wan 3.0 prompt is a production document, not a caption. Cast your references, legislate your style, schedule your beats, and treat sound as half the picture. Start with the graduate close-up skeleton, graduate to the five-module brief, and save the frame-budgeted motion spec for the shot that needs it. The models will keep changing. The discipline of writing time instead of describing pictures is the part you get to keep.

Latest Models

One API for All Media AI.

Explore all models