
Two editing monitors side by side playing the same desert gas station shot, one with an unbroken 30-second timeline below it and one with a timeline split into two segments
At second 20, the fuel nozzle sprays out fresh green grass. Still beaded with water. The horse squints with pleasure. The attendant's toothpick drops out of his mouth.
That joke only works inside one continuous 30-second shot. Cut four 8-second clips together and the horse changes coat, the sunglasses change shape, the sky shifts a stop warmer, and the punchline dies somewhere in the seams.
So when Wan 3.0 and Seedance 2.5 both landed in the same window claiming native 30 seconds, one pass, no stitching, it stopped being a benchmark story. It became the first time an AI video model could hold a setup, a turn, and a payoff in a single take.
I pulled the specs apart, ran a spec check against Alibaba's own published demo files, and found the thing nobody writing about this has printed: both models say 30 seconds, but the resolution underneath those 30 seconds is not the same product at all.
Key Takeaways
- Both models really do generate 30 seconds in a single pass. That part is not marketing. Wan 3.0 covers 2 to 30 seconds with a smart-duration option, Seedance 2.5 covers 4 to 30 seconds.
- Only Wan 3.0 renders native 1080p at 30 seconds. Seedance 2.5 generates natively at 480p and 720p only. Its 1080p, 1440p and 4K options are post-generation upscales of a 720p source, and its own API documentation says the 4K tier is "not native 4K generation".
- Seedance 2.5 wins on reference volume: up to 30 images plus 10 videos plus 10 audio files, 50 in total. Wan 3.0's creator handbook caps references at 20.
- Wan 3.0 has one input nobody else has:
.doc,.xls,.ppt,.pdf,.txtfiles and live web pages, fed in directly as creative references. - Only one of the two is actually runnable outside Alibaba today. Wan 3.0 sits in public beta on Alibaba Cloud Model Studio and Qwen Cloud with third-party API access still rolling out. Seedance 2.5 is live on Atlas Cloud with 30 in the duration control right now.
- Want to pressure-test 30-second story structure before you pay for a single second of it? Atlas Cloud's free AI ad skit generator gives you one free run: a 15-second finished spot with spoken dialogue and sound effects, watermark-free download.
Alibaba's official Wan 3.0 demo, from the Tongyi Wan 3.0 creator handbook. One take, no cuts, 30.07 seconds. Measured with ffprobe: 1920x1072, 24fps, AAC stereo audio baked in. Sound on for the tail flick and the toothpick hitting the ground at second 29. Used here as reference material for comparison and commentary only.
Wan 3.0 vs Seedance 2.5 Native 30s: What Actually Changed This Month
Fifteen seconds was the wall for a long time. Wan 2.7 tops out there. Seedance 2.0 topped out there. Every "one-minute AI film" you saw on your feed for the last year was four to eight clips glued together in an NLE, with a colourist pass to hide the joins.
Two things broke that wall in the same few weeks. Alibaba opened Wan 3.0 public beta on 6 August 2026 with native 30-second generation and full multimodal reference input (AlphaSignal, August 2026). ByteDance doubled Seedance from 15 to 30 seconds in 2.5, positioning it explicitly as a way to cut clip stitching and the visual drift that comes with it.
Native matters here in a specific technical sense. It means one forward pass over one latent sequence, not last-frame extension where the model takes the final frame of clip A and starts clip B from it. Extension is why your character's jacket collar quietly changes shape between shots. A native pass has the whole 30 seconds in the same context, so the horse stays the same horse.
The gas station demo is the cleanest possible test of that. The structure is four beats: 0 to 8 seconds of nothing happening, 8 to 16 of something odd happening, 16 to 24 for the punchline, 24 to 30 for the walk-off. The joke lands at second 20. Which means it only lands if the horse, the sunglasses, the teal-and-orange grade and the deadpan locked-off camera all survive the first 19 seconds untouched. Stitching cannot do that. Not because the editor is bad, but because clip B never saw clip A.
Wan 3.0 vs Seedance 2.5 Native 30s: Where the Picture Actually Breaks
Thirty seconds is twice fifteen. The drift risk is not twice as large, it is worse than that, and it moves to a different failure mode.
The stitched-workflow failures were between clips. Native 30s failures are inside the clip. In a hands-on Seedance 2.5 short-film production, the reviewer found decoherence and morphing showing up as "visual spikiness or instability in character models", concentrated in fast action sequences, and specifically noted that these artefacts cannot be cleaned up by upscaling (MindStudio, August 2026). That is the part that matters for anyone planning to render at 720p and upscale to 4K afterwards: the upscaler sharpens the morph, it does not remove it.
The same production logged a 3.2-to-1 shoot ratio. Roughly 450 seconds of raw generation to cut a 2 minute 22 second final. Plan long-take work like a shoot with coverage, not like a render queue where every job ships.
Two more findings from that production are worth stealing outright. Visual identity across the whole film rested on three reference images, two character portraits and one location plate, even though the system accepts up to 50. And in voice casting, writing "American" fixed an unintended British accent while writing "American English" did not, because the word "English" nudged delivery back toward British.
The prompt structure that survives 30 seconds is the one Alibaba uses in its own handbook: explicit timestamped beats. Not a paragraph of vibes. Write 0-8s, 8-16s, 16-24s, 24-30s, give each block its own camera behaviour and its own event, and end with a single audio line covering the whole clip. You will see that exact skeleton in Step 4 below, because I kept Alibaba's structure and only translated it.
Wan 3.0 vs Seedance 2.5 Native 30s: Specs, Price, and What You Can Run Today
Here is the honest state of play, and the part where the two models stop being comparable products.
Read the resolution row twice, because it is the whole article. Atlas Cloud's own Seedance 2.5 Text-to-Video API documentation states that 720p-esr enhances a 480p source, that 1080p-esr, 1440p-esr and 4k-esr all enhance a 720p source, and that the 4K option delivers a 2160-pixel short edge "not native 4K generation". So a 30-second Seedance clip labelled 4K is 720p worth of real detail, resampled. Wan 3.0's price sheet, by contrast, has a straight 1080p tier at $0.20 per second, and Alibaba's own published 30-second demo files measure 1920x1072.
The upside of that constraint: the entire test runs in one browser tab, on three models, with no local setup.
| Role in this test | Model | Listed price |
|---|---|---|
| Opening frame | GPT Image 2 Text-to-Image | from $0.009 / image |
| The native 30s run | Seedance 2.5 Text-to-Video | from $0.134 / SEC |
Listed prices verified on the Atlas Cloud model pages, August 2026. Read them as floors, not quotes. Every one of these models bills by what you actually ask for, and the Run button prices your exact settings before you click it. In this test the button quoted $0.1745 for one GPT Image 2 frame at quality high and 2048x1152, and $1.514799 for a five-second Seedance 2.5 job at 720p and 16:9, which works out to roughly $0.30 per second rather than the listed $0.134. The listed figure is the 480p rate. Check the button, not the catalogue. Seedance 2.5 also ships a reference-to-video endpoint at the same $0.134 per second if you want to push the 50-reference ceiling.
How to Reproduce This Native 30s Test on Seedance 2.5
Five steps, three models, all in the browser. Copy the prompts verbatim.
Step 1: Lock the Timestamped 30s Skeleton
Do not run anything yet. Read the structure first, because this is the part that decides whether your 30 seconds holds.
The Wan 3.0 gas station demo at the top of this article is built as four labelled blocks with a fixed camera doctrine stated once at the head: calm objective observation, locked-off medium wides, sudden extreme close-ups only on the absurd beats. Every event is pinned to a time range. Audio is one line at the end covering all 30 seconds.
Here is the skeleton, empty. Fill it and you have a prompt that a native 30s model can actually track:
Plain1A 30-second [genre] set in [place]. [Visual style, grade, saturation]. Camera language: [doctrine], punctuated by [exception]. 2 30-8s: [setup, wide, establish the normal world, introduce the anomaly on the horizon] 4 58-16s: [the anomaly acts, still treated as normal, one POV or reaction insert] 6 716-24s: [the payoff, extreme close-up on the thing itself, hard cut to the reaction face] 8 924-30s: [the walk-off, a small physical button that closes the joke] 10 11Audio: [ambience, three or four specific diegetic sounds, one final small sound]. No music, no dialogue, no on-screen text.
The measurable spec of Alibaba's reference file, for anyone who wants to check my numbers: 30.07 seconds, 1920x1072, 24fps, AAC stereo. That is the bar the Seedance run is being measured against.
Step 2: Generate the Opening Frame
You need a fixed visual anchor so the 15-second and 30-second runs start from the same world. Open GPT Image 2 Text-to-Image.
Settings: quality high, aspect ratio 16:9, 1 image.
Plain1A wide, dead-still establishing shot of a lonely Route 66 style gas station in a bright desert. Retro yellow-orange-and-blue building, paint faintly sun-bleached; a faded vintage-red fuel pump out front; a small convenience store beside it with a washed-out cola vending machine by the door. Cracked orange dry earth in the foreground, warm terracotta mountains far behind, cloudless clean cyan-blue sky. A gas station attendant in greasy blue coveralls and oversized black sunglasses slouches half-asleep in a chair by the door, toothpick in his mouth. Strict bright teal-and-orange color grading, high saturation, high brightness, cinematic photoreal, locked-off camera, 16:9.

GPT Image 2 playground on Atlas Cloud with quality set to high and 16:9, the generated Route 66 gas station establishing shot in the output panel
GPT Image 2 on Atlas Cloud: quality high, 16:9, one image. The Run button quoted $0.1745 for this frame, which is what a 2048x1152 high-quality render actually costs. The $0.009 on the model page is the floor.

Route 66 gas station establishing frame generated with GPT Image 2, teal sky and orange cracked earth, attendant dozing by the door
The opening frame. This is the input for Step 3 and the visual target for Step 4. Generated with openai/gpt-image-2.
Step 3: Hit the 15-Second Wall
Settings: first frame = your Step 2 output, duration dragged to its maximum of 15, resolution 1080P.
Plain1Locked-off wide shot holds. A cowgirl in a wide-brim hat rides a real horse in from the horizon at a walk. The dozing attendant is woken by hoofbeats, pushes his sunglasses up, sits upright, chewing his toothpick, unimpressed. She dismounts cleanly, leads the horse straight to the vintage fuel pump. Bright teal-and-orange grade, high saturation, photoreal, calm observational camera, no cuts.
The dropdown stops at 15. That is the whole point of this step. Your punchline is scheduled for second 20, and this pipeline cannot reach second 20 in one piece. To get there you would generate a second clip from the last frame, and the moment you do that, the horse is a new horse.

Step 4: Run the Full Native 30s Pass
Open Seedance 2.5 Text-to-Video.
Settings: duration = 30, resolution = 720p, ratio = 16:9, generate_audio = on, output_format = mp4, watermark off.
Pick 720p deliberately, not 1080p-esr. 720p is the model's real ceiling, so this is a like-for-like pixel comparison instead of a comparison against an upscaler.
The prompt below is Alibaba's own gas station demo prompt, translated faithfully from the Wan 3.0 creator handbook and restructured into nothing. Same beats, same timings, same camera doctrine, same audio list.
Plain1A 30-second absurdist deadpan comedy short set at a sun-bleached Route 66 style gas station in a bright desert. Photoreal, strict bright teal-and-orange grading, high saturation and high brightness, so the absurd event happens in a completely normal, almost pretty world. Camera language: calm, objective observation, mostly locked-off medium wides and slow lateral drifts, no emotional steering, punctuated by sudden extreme close-ups on the absurd beats. 2 30-8s: Wide, locked-off. Cyan sky, drifting dust. The retro yellow-orange-and-blue station sits alone by the road; an attendant in greasy blue coveralls and huge black sunglasses dozes in a chair, toothpick in mouth. Only wind. On the horizon a cowgirl on a real horse appears at a walk. Cut to medium: hoofbeats wake him, he pushes his sunglasses, sits up, reads the pair as "another one asking for directions." 4 58-16s: She says nothing. She dismounts cleanly, leads the horse straight to the old fuel pump. The horse casually puts its head next to the pump like it has done this before. Brief slightly fish-eyed POV from behind his sunglasses, the woman and horse look even stranger. Close: his brow furrows, he takes the sunglasses off, about to shout. Slow-motion close-up: as the glasses come off, she aims the fuel nozzle at the horse's mouth and squeezes the trigger. 6 716-24s: Extreme close-up on the nozzle. Out comes not gasoline but jets of fresh bright-green grass, still beaded with water. Hard cut to an extreme close-up of the attendant's face, frozen: eyes wide as saucers, mouth slightly open, toothpick forgotten mid-chew. Medium: the horse eats happily, squints with pleasure, then gives one satisfied powerful flick of its tail, and the dust it kicks up lands squarely on the still-stunned attendant's face. 8 924-30s: "Tank full" of grass. She hangs the nozzle back on the pump, digs coins from her pocket, walks to the store door and drops them clinking into a tin can on the counter. She glances back at the petrified attendant; under the hat brim the corner of her mouth lifts a fraction. She remounts and rides off in a trail of dust. The camera pushes slowly in on him: same frozen posture, face covered in dust, sunglasses still pinched in his hand, and finally the toothpick drops from his mouth with a small clack. 10 11Audio: dry desert wind throughout, hoofbeats, the mechanical clunk of the pump handle, wet grass spurting, the tail flick, coins in a tin can, and one small clack as the toothpick hits the ground. No music, no dialogue, no on-screen text.
One warning that will save you a wasted bill, and it is the same trap as Step 3: drag the duration slider to 30, do not type 30 into the number box. The box will happily show 30 while the slider stays on its five-second default, and the Run button will quote you for five seconds. Check the quote before you click. At 720p and 16:9 a five-second job quotes $1.514799, so a real 30-second pass quotes roughly six times that.
There is no finished-run screenshot for this step. Three automated capture attempts all submitted the slider default instead of 30 seconds, and rather than show you a five-second clip labelled as a 30-second one, the capture is left out. The prompt and settings above are exactly what to paste and set.
Step 5: Read the Drift, Honestly
Here is the Seedance 2.5 side of that exact prompt, generated at native 30 seconds, 720p, 16:9, audio on.
Seedance 2.5, the same Wan 3.0 gas-station prompt copied verbatim: native 30s in one generation, 1280x720, 24fps, audio baked in. Same four beats, cowgirl and horse arriving, the fuel-nozzle gag, the attendant's cognitive-dissonance close-up. Put it next to the Wan 3.0 clip at the top for a like-for-like read.
Put the two clips side by side and check five specific things. Do not check "which looks better", that is a taste argument and it will waste your afternoon.
- Coat and costume identity. Is the horse the same colour at second 29 as at second 6? Are the sunglasses the same shape after they come off and go back into his hand?
- Grade stability. Sample the sky at second 2 and second 28. Teal-and-orange grading drifts warm over long generations more often than people expect.
- The physical beat at second 20. Grass out of a fuel nozzle is a physics request. Does it arc and fall with weight, or does it fade in as a texture?
- Camera discipline. The prompt says locked-off. Count how many times the camera invents a push-in it was not asked for. This is the most common long-generation failure: the model runs out of scheduled events and fills time with movement.
- Fast-motion morphing. The tail flick and the dust kick are the fastest motion in the clip. Per the MindStudio production notes, this is exactly where decoherence concentrates, and exactly what an upscale will not fix.
Score those five, not the vibe. That is a comparison you can defend to a client.
Three More Wan 3.0 vs Seedance 2.5 Native 30s Matched Prompts
One demo is an anecdote. Here are three more official Wan 3.0 30-second files from the same handbook, each stressing a different part of the 30-second problem. For the sea monster and the dark fantasy monologue, the Seedance 2.5 side was generated on the same prompt at native 30 seconds and is embedded right after each, so you can watch both rather than take my word for it.
| Prompt | What it stress-tests | Wan 3.0 official file, measured | Seedance 2.5 side, same prompt |
|---|---|---|---|
| Sea monster disaster film | 10 scripted shots plus an orchestral score inside 30s | 1920x1072, 24fps, 30.07s, AAC | generated at 30s, embedded below; one IP name and the boat brand text were removed so Seedance's content filter would pass it, the storyboard is otherwise identical |
| Dark fantasy English monologue | native English voice and lip sync across a slow continuous push | 1280x720, 24fps, 30.05s, AAC | generated at 30s from the prompt verbatim, embedded below |
| 2D sports anime, two-line prompt | can the model fill 30s with almost no instruction | 1280x720, 24fps, 30.05s, AAC | not generated here; shown as a Wan-only frame grid to read structure across time |
| Whip-pan comedy short (excluded) | 8 whip pans and 8 emotional beats without a cut | 1280x720, 24fps, 30.04s, AAC | not run: dialogue density is Mandarin-specific, an English rewrite would not be a fair like-for-like |
Official Wan 3.0 demo: ten scripted shots and a full orchestral build inside one 30-second pass, at 1920x1072. This is the hardest argument for native 1080p, because the water volume and the creature's wet skin detail are exactly what a 720p upscale invents rather than resolves. Alibaba Tongyi creator handbook, used for comparison and commentary.
Seedance 2.5, the same ten-shot disaster prompt at native 30s, sound on. Storm-tossed fishing boat, the terrified deckhand, the swell heaving up, then the deep-sea leviathan breaking the surface in the lightning. One IP reference and the on-hull brand text were stripped so Seedance's content filter would pass it; the storyboard is otherwise identical, which is what makes the water-volume and wet-skin detail a fair like-for-like against the Wan 3.0 clip above.
Official Wan 3.0 demo, sound on. Native English villain dialogue, "Ripe enough for the void", delivered in a single slow upward tilt with no cut. This is the clip English-speaking teams should test against, because voice, lip sync and a 30-second continuous camera move all have to hold at once.
Seedance 2.5, the same dark-fantasy prompt copied verbatim, native 30s, sound on. Same violet-eyed villain, the star-rune markings, the shadow-nebula forming in his open palm, and the two English lines. Watch whether the lip sync and the single continuous push hold across the full 30 seconds the way they do on the Wan 3.0 side.
If you run the matched Seedance 2.5 side of this one, keep the two English lines verbatim and remember the accent quirk from the production notes: write "American", not "American English". The word "English" pulls the delivery back toward a British read.
The third one is the quiet surprise. The 2D sports anime prompt in Alibaba's handbook is two lines long. No timestamps, no shot list, just a boxer hitting a heavy bag, tired, in pain. Thirty seconds from that is a genuine test of whether the model can invent its own structure. Here it is sampled across its own runtime:

Four frames from the official Wan 3.0 2D sports anime demo sampled at roughly 1, 10, 20 and 29 seconds, showing the boxer's design and the gym background across the full 30 seconds
Four frames from the official Wan 3.0 2D sports anime demo, sampled at roughly 1, 10, 20 and 29 seconds. Same character design, same tank top, same heavy bag, same row of lockers, across a wide-to-close-up change the prompt never asked for. A still grid rather than a clip, because the comparison here is across time inside one file, not motion.
Practical read: with a two-line prompt the model invented its own coverage, moving from a locked wide to a sweat-and-exhaustion close-up, and held the character design while doing it. That is the good news. The trade is that you gave up control of when anything happens. If you need 30 seconds to land a specific beat at a specific second, write the timestamps.
Test Native 30s Storytelling Free Before You Pay for Either
A 30-second 720p run on Seedance 2.5 quotes around $9 once you price the real resolution rather than the catalogue floor. A 30-second 1080p run on Wan 3.0's price sheet is $6.00. Neither is expensive by production standards, but at a 3.2-to-1 shoot ratio you are not buying one clip, you are buying three or four.
Before you spend anything, it is worth answering a cheaper question: does my script actually hold up as one continuous shot? Most 30-second failures are not model failures. They are structure failures, three beats crammed where there was room for two, or a payoff written at second 26 with nothing carrying seconds 8 through 20.
Atlas Cloud's free AI ad skit generator answers that for nothing. You give it one product description and up to four product photos, each under 8MB, pick one of six styles (funny meme, plot twist, sitcom, heartwarming, luxury, or hard sell), and it writes a two-character script first, with a three-second hook, a conflict and a twist, then builds every shot around that script. Characters speak their lines out loud and sound effects are rendered into the video.
The honest limits: it is 15 seconds, not 30, you need to be logged in, and you get one free generation before you top up credits. But 15 seconds with a hook, a conflict and a twist tells you whether your three beats breathe. If the structure works there, take the same beats, stretch them onto the 0-8s / 8-16s / 16-24s / 24-30s skeleton from Step 1, and go spend the nine dollars on a real native 30-second pass.
What a Native 30s Clip Actually Costs on Wan 3.0 vs Seedance 2.5
| Setup | Rate | Duration | One clip |
|---|---|---|---|
| Wan 3.0, 1080p native (Alibaba Cloud only) | $0.20 / sec | 30s | $6.00 |
| Wan 3.0, 720p native (Alibaba Cloud only) | $0.10 / sec | 30s | $3.00 |
| Wan 3.0, 480p native (Alibaba Cloud only) | $0.05 / sec | 30s | $1.50 |
| Seedance 2.5, 720p native, on Atlas Cloud | $0.30 / sec measured at 720p, listed from $0.134 | 30s | about $9.09 |
| GPT Image 2 opening frame, quality high, 2048x1152 | listed from $0.009 / image | 1 image | $0.1745 quoted |
The comparison that actually decides it: Wan 3.0 at 720p is cheaper per second than Seedance 2.5 at 720p, and at 1080p it is the only one of the two selling real pixels. Seedance 2.5's answer is 50 reference slots, live availability, and a working API you can ship against this afternoon.
Sourcing and Licensing Note for These Native 30s Clips
Every Wan 3.0 clip and every Wan 3.0 prompt in this article comes from Alibaba's publicly distributed Tongyi Wan 3.0 creator handbook, reproduced here for comparison and commentary, credited, unmodified and not de-watermarked. Commercial use of Seedance 2.5 output falls under ByteDance and Volcengine terms; API billing and data handling on the Atlas Cloud side fall under Atlas Cloud's own terms. No real person's likeness is reproduced anywhere in this test. If you are shipping client work, read the model terms for the specific endpoint you call, not a blog summary of them.
Frequently Asked Questions
Is Wan 3.0's native 30s actually one pass, or stitched clips?
One pass. Wan 3.0 generates 2 to 30 seconds in a single generation with a smart-duration option, so it can also decide the clip does not need the full 30. That is different from last-frame extension, where a second clip is seeded from the final frame of the first and identity drifts across the seam.
Which one is really 1080p at 30 seconds, Wan 3.0 or Seedance 2.5?
Wan 3.0. It has a native 1080p tier on its price sheet, and Alibaba's own 30-second demo files measure 1920x1072. Seedance 2.5 generates natively at 480p and 720p only; its 1080p, 1440p and 4K options are enhancements of a 720p source, and its API docs describe the 4K tier as "not native 4K generation".
Does the picture still fall apart at second 25?
It can, but the failure changed shape. Instead of a visible seam between clips, you get morphing and decoherence inside the shot, concentrated in fast-motion passages, and upscaling does not remove it. A documented Seedance 2.5 short-film production ran a 3.2-to-1 shoot ratio, so plan long takes with coverage.
Which model takes more reference material?
Seedance 2.5, by a wide margin: 30 images plus 10 videos plus 10 audio files, 50 in total, with reference video and audio each capped at 30 seconds combined per request. Wan 3.0's handbook caps references at 20, but it is the only one that accepts .pdf, .ppt, .xls, .doc, .txt files and live web pages as creative input.
Will Wan 3.0 have open weights?
There is no official commitment. As of the August 2026 public beta, no Wan 3.0 weights, Hugging Face checkpoint or ComfyUI node have been published, and official open weights for the flagship video line stop at Wan 2.2 (Kingy AI, August 2026). Treat anything else you read about a 3.0 weight drop as speculation until Alibaba says otherwise.






