TYLKO DWA TYGODNIE | 20% ZNIŻKI na Seedream 5.0 Pro!

Seedance 2.5 vs MiniMax H3: Everyone Tested the 30-Second Flex. Nobody Tested Shot 4.

Seedance 2.5 vs MiniMax H3, same character sheet, same prompt. Real specs, a copy-paste 5-step short-drama workflow, and which one holds a face across four shots.

MiniMax shipped H3 on Friday. Inside twenty-four hours Seedance 2.5 was in people's hands, and the timeline collapsed into one conversation: thirty seconds, native 4K, fifty reference inputs. H3's launch post sat underneath, quietly, with almost nobody running a real test on it.

I get why. Thirty seconds is a great number. It fits in a headline.

But if you actually make vertical short drama, you know the thing that ruins your night is not second thirty. It is shot 4. The male lead's jawline shifts half a centimetre, his cravat loses a fold, and the viewer swipes away before the hook lands. Nobody churns because a clip was fifteen seconds instead of thirty. They churn because the guy in the close-up is not the guy from the wide.

So I skipped the spec-sheet duel. I built one character turnaround sheet, one six-panel location board, wrote one 15-second gothic short-drama prompt, and sent the identical package to both sides. Then I stared at the face.

Key takeaways

  • Both are real, but not equally reachable. H3's three endpoints (text, image and reference to video) are live and callable right now. Seedance 2.5's API is out at ByteDance, but on the multi-model platform where I ran everything it is still a Day-0 page, so the Seedance side of my head-to-head is Seedance 2.0 Reference-to-Video. I flag that everywhere it matters.
  • The specs split cleanly. Seedance 2.5 = up to 30 seconds in one generation, native 4K, up to 50 multimodal references, region-level editing. H3 = 5 to 15 seconds, native 2K at 24fps, 9 images + 3 videos + 3 audio clips (12 files max), stereo audio on every output, whole-clip instruction editing.
  • Character stability is a packaging problem, not a model-generation problem. With one composite turnaround sheet plus one lighting board, both models held the same face and the same wardrobe. Neither generation number saved or ruined the shot; the pack did.
  • The model nobody tested is already rated third. On the Artificial Analysis image-to-video arena, H3 sits at 1,185 Elo behind Seedance 2.0 at 1,196. Seedance 2.5 is not on the board yet at all.
  • Both generate their own audio. H3 also accepts up to three audio references to pin a voice to a character, which deletes a whole TTS-and-sync stage from a short-drama pipeline.
  • Ask about pixels per frame, not seconds. Same reference pack, same runtime: H3's reference-to-video returns 1440p, Seedance 2.0's reference-to-video returns 720p. That reframes the whole cost question.

Seedance 2.5's own paired numbers are not in yet. I will add them the day the endpoint opens where I test.

The Seedance 2.5 vs MiniMax H3 Result, Before Any Theory

Before any theory, here is what the workflow produces. These are four shots pulled out of a finished 15-second vertical clip and tiled. It is MiniMax's own H3 reference run built from exactly the character sheet and location board you will see in Step 1 and Step 2, at native 1440x2560. My own runs of the same pack come later, in the tutorial, with the playground state visible.

 

 

She reaches the carved door, which is panel 4 of the location board, then the corridor stand-off, then a tight close-up on him, then the two-shot. His hairline parts the same way in the profile and in the close-up. Her half-up braided crown survives a full-face close-up, which is exactly where most models quietly rebuild a face. The cream lace bodice keeps the same neckline and sleeve cuff from the first frame to the last.

That is the whole test. Not thirty seconds. Four shots and a jawline.

Why Seedance 2.5 vs MiniMax H3 Broke the Same Week, and Why Most Character Tests Are Useless

MiniMax released H3 on 30 July, a general-purpose multimodal video model that reads text, images, video and audio as one context and returns 5 to 15 second clips at up to 2K with native stereo sound. It can also edit existing footage and transfer motion from one clip to another, and MiniMax said weights would follow within days (Reuters, July 2026).

Seedance 2.5 landed in the same window with a louder number: 30 seconds in a single generation, native 4K, and up to 50 multimodal reference inputs in one job (The Next Web, July 2026). Its API is already in developers' hands, with localized editing that redraws part of a frame while leaving performance, lighting and camera behaviour intact, reference-to-video that accepts green-screen plates or 3D white-model blockouts, eleven languages, and an experimental long-video mode reaching 180 seconds (CineD, July 2026).

 

 

The attention gap is the interesting part. Almost every post I could find was a 2.5 clip. Meanwhile the public arena numbers say something else: on the Artificial Analysis image-to-video leaderboard, Seedance 2.0 at 720p leads with 1,196 Elo, Gemini Omni Flash follows at 1,195, and MiniMax H3 is already third at 1,185, four days out of the gate. Seedance 2.5 has no rating there yet (Artificial Analysis, July 2026). The model with the least chatter is the one with the highest ratio of score to attention.

Now the part that actually decides your output, because it has almost nothing to do with which logo you pick.

Most character-consistency tests fail before the model is even involved. Four failure modes, and all four are yours to fix:

Failure modeWhat you see in the outputThe fix
Multi-view references sent as separate filesEvery shot's face is "nearly right" and no two agreeComposite the views into one image, then assign roles in the prompt ("the man on the LEFT", "the woman on the RIGHT")
Rewriting the character description per shotWardrobe and accessories drift shot to shotWrite the identity block once and reuse it verbatim; only camera and action change per shot
Letting lighting move with the sceneThe face gets structurally rebuilt in the new lightBuild a separate location board that locks light direction, fog and floor reflection, and feed it as reference
Treating max duration as a quality metricOne drift inside 30 seconds kills all 30Cut to the granularity you can re-roll, then talk about length

Row one is the one people get wrong most. Send six single-view images and the model treats them as six candidate people and averages them. Send one clean composite and it treats them as one person seen from three sides. Same pixels, completely different result.

 

 

Seedance 2.5 vs MiniMax H3 Spec Sheet, and What You Can Actually Call Today

Split capability from availability and the tension gets clear. On raw capability, Seedance 2.5 leads on duration, resolution ceiling, reference count and editing granularity. On availability, H3 is the one with three live endpoints on a general-purpose model platform this week. If you are doing AI video model comparison for 2026 planning, that second column matters as much as the first.

 MiniMax H3Seedance 2.5Seedance 2.0 Ref-to-Video (what I ran)
Max length, one generation5 to 15 sup to 30 s, no stitching (beta mode to 180 s, experimental)short-clip menu, see model page
Resolutionnative 2K, 1440p short side at 16:9 to 9:16; a 768p tier is listed as comingnative 4K, 10-bit colourup to 720p on this endpoint
Frame rate24 fpsnot separately publishednot separately published
Reference inputs9 images + 3 videos + 3 audio, 12 files totalup to 50 multimodal references9 images + 3 videos + 3 audio
Native audioyes, every output, stereoyes, plus 11 subtitle and voice languagesyes
Editing granularitywhole-clip instruction edits (swap character, background, lighting, dialogue)region-level edits that redraw part of a framewhole-clip instruction edits
Prompt cap7,000 charactersnot publishednot published
Open weightsannounced for days after launchnot announcednot announced
Artificial Analysis I2V Elo, 31 Jul 20261,185 (3rd)not rated yet1,196 (1st, Dreamina 720p entry)
Callable on the platform I testedyes, 3 endpointsnot yet, Day-0 pageyes

Full disclosure on the test setup. Seedance 2.5 was not open on my platform when I ran this, so the Seedance side is Seedance 2.0 Reference-to-Video, the currently shipping member of the same family with the same quad-modal reference system. Where 2.5 would change the answer, I say so. The Seedance 2.5 page is a Day-0 notice for now.

Everything below runs in one browser tab, on one key, three models total:

StepModelWhat it producesKey settingsWhat drives the cost
1OpenAI GPT Image 2 Text-to-Imagethe character turnaround sheetquality high, 16:9per image, pay as you go, current rate on the model page
2same model, second runthe six-panel location and lighting boardquality high, 16:9per image, pay as you go
3MiniMax H3 Reference-to-Videothe 9:16 clip with native audio9:16, 2K, duration 5 for a test and 15 for the real cutseconds x resolution, billed per second
4Seedance 2.0 Reference-to-Videothe control run, identical inputs9:16, top available resolution, same durationseconds x resolution, billed per second
5H3 Reference-to-Video, clip as inputone fixed shot without re-rolling everythingsame resolution, short durationseconds x resolution, billed per second

H3 also has text-to-video and image-to-video entries if you want a pure text baseline or first-and-last-frame control instead.

One more thing worth knowing before you plan a batch: on-screen text. This 15-second H3 title sequence holds English credits through eight hard cuts without a single misspelling, which is one of the hardest things to keep stable in generated video.

 

 

The Seedance 2.5 vs MiniMax H3 Short-Drama Workflow, Step by Step

Five steps. Copy the prompts exactly as written, including the ugly parts. I did not clean them up for the article and you should not clean them up for the model.

Step 1: Build the character sheet with GPT Image 2

Make the reference pack before you make any video. One image, two characters, three views each, pure white seamless background, even studio light. That removes every ambiguity except the one you care about: what does this person look like.

plaintext
1A full-body character turnaround reference sheet on a pure white seamless background, two characters side by side, six figures total, even neutral studio lighting, no shadows on the floor, no text, no labels, no watermark.
2
3LEFT HALF, male lead, three views in a row: front, right profile, back. Late 20s, pale complexion, dark wavy medium-length hair, sharp jawline. Wearing a floor-length black velvet frock coat with subtle tone-on-tone embroidery on the lapels and cuffs, a black brocade waistcoat, a black silk cravat, straight black trousers, polished black leather derby shoes. Arms relaxed at his sides, neutral expression.
4
5RIGHT HALF, female lead, three views in a row: front, right profile, back. Early 20s, fair complexion, long wavy chestnut-brown hair with a half-up braided crown. Wearing a cream Victorian cotton-lace dress, long sleeves with lace cuffs, fitted bodice with small buttons, full ankle-length skirt, cream ankle-strap low-heel shoes. Arms relaxed at her sides, neutral expression.
6
7Photographic realism, consistent scale between all six figures, feet aligned on the same baseline, sharp focus edge to edge.

Settings: model OpenAI GPT Image 2 Text-to-Image, Quality high, Aspect ratio 16:9, everything else default. Total: 11 words. Perfect. AI image generator interface showing prompt input and generated character GPT Image 2 Text-to-Image on Atlas Cloud, run completed: the turnaround prompt on the left, six figures on one white sheet in the OUTPUT panel.

This is the target shape. Six figures, one baseline, nothing in the background to argue about: Let's The reference pack that drove the demo clip: male lead in black velvet, female lead in cream Victorian lace, three views each, one file. Front, side, and back views of Victorian black and cream outfits My own GPT Image 2 run of the Step 1 prompt, for comparison with the reference above.

Step 2: Lock the location with a six-panel scene board

You held the face. Light will still betray you. The second reference image nails down six camera positions, the direction of the moonlight, how much fog sits in the air and how wet the floor reads. You are telling the model this film has exactly one lighting setup.

plaintext
1A six-panel cinematic location board, 2 rows by 3 columns, thin black gutters between panels, each panel captioned in small elegant white all-caps letterspaced type at the bottom of that panel. Subject: the interior of an abandoned gothic castle at night. Cold blue moonlight, drifting low fog, wet reflective stone floor, tall stained-glass windows, iron candle sconces with small warm flames, deep desaturated blacks, heavy velvet curtains. No people in any panel.
2
3Panel 1, caption "WIDE ESTABLISHING VIEW - CORRIDOR": a long pillared corridor receding to a lit stained-glass window.
4Panel 2, caption "LOW ANGLE - MOONLIGHT ACROSS FLOOR": low camera, a shaft of moonlight raking across wet flagstones.
5Panel 3, caption "DOORWAY & STAIRCASE COMPOSITION": an arched doorway framing a curved stone staircase.
6Panel 4, caption "CLOSE DETAIL - CARVED DOOR": tight shot of a heavy carved wooden door with an iron handle.
7Panel 5, caption "WINDOW-SIDE VIEW - FOG & CURTAINS": tall window, fog rolling in, curtain edge in foreground.
8Panel 6, caption "FINAL POSTER FRAME - EMPTY HALL": symmetrical empty hall, patterned runner rug leading to the window.
9
10Photographic, cinematic, film grain, consistent colour grade across all six panels.

Settings: same model, Quality high, Aspect ratio 16:9. Six storyboard panels of dark gothic castle interiors with moonlight The location board that drove the demo clip: six gothic castle interiors, one colour grade, captions readable. Six cinematic camera angles of a dark gothic castle interior My own GPT Image 2 run of the Step 2 board prompt.

Step 3: Generate the shot with MiniMax H3 reference-to-video

Two reference images plus one prompt carrying camera positions and audio direction. Sound is produced in the same pass, so there is no separate voice or score step. Keep the duration short while you are tuning, then push it to 15 for the real cut.

plaintext
1Reference image 1 is the character sheet: the man on the LEFT is the male lead, the woman on the RIGHT is the female lead. Keep both identities exactly as shown: his jawline, dark wavy hair, black velvet frock coat, black brocade waistcoat and black silk cravat; her chestnut half-up braided hair and cream Victorian lace dress. Reference image 2 is the location and lighting board: match its cold blue moonlight, drifting fog, wet reflective stone floor and desaturated black tones.
2
39:16 vertical, premium live-action short-drama look, shallow depth of field, cinematic contrast, no subtitles, no on-screen text, no watermark.
4
5Shot 1: Medium close-up, camera slowly pushing in. The female lead stands in the moonlit corridor, breathing shallow, eyes fixed on something off-frame right. Fog moves past her at knee height.
6Shot 2: Hard cut. Over-the-shoulder from behind her, the male lead steps out of the dark archway into the moonlight, coat catching the light, expression cold and curious.
7Shot 3: Tight close-up on his face, half-lit by the window, then a matching close-up on hers as she refuses to look away.
8
9Audio: low sustained bass drone, distant stone echo, her unsteady breath, one heavy footfall as he steps into the light, a single soft string swell on the last cut.

Settings: model MiniMax H3 Reference-to-Video, Resolution 2K, Duration 8 (the slider default; push to 15 for the real cut), Aspect Ratio adaptive, and both files in Reference Materials. The panel counts them as 2 of 9, so you have plenty of room to add wardrobe or prop plates later.

Worth noting what happened with the ratio. I left Aspect Ratio on adaptive and only wrote "9:16 vertical" inside the prompt. H3 came back vertical anyway, so it read the framing out of the text rather than needing the form field. AI video generator interface showing text prompt and completed video preview MiniMax H3 Reference-to-Video on Atlas Cloud, run completed: both reference files loaded as 2 of 9, resolution 2K, and the generated vertical clip playing in the OUTPUT panel.

The first frame is the female lead in the cream lace bodice, standing in the corridor from board panel 1, with fog at knee height and the moonlight hitting the wet floor exactly where the board put it. Both reference images landed.

Step 4: Run the Seedance 2.5 vs MiniMax H3 control with the identical pack

Do not change a word. Same two reference images, same prompt, different model. I left each page's own duration default alone rather than forcing a match, and I say what each one returned below. Re-wording a prompt "for" the other model is exactly how comparisons start lying.

plaintext
1Same prompt as Step 3, verbatim. Do not re-word it for the other model.

Settings: model Seedance 2.0 Reference-to-Video, Resolution 720p, same two reference images in Reference Images (again 2 of 9, with separate slots for up to 3 reference videos and 3 reference audio files). Duration left at the page default, which came back as a 10-second clip. AI video generation interface showing a text prompt and completed video Seedance 2.0 Reference-to-Video on Atlas Cloud, run completed with the identical reference pack and prompt from Step 3, and a 10-second result in the OUTPUT panel.

Here is what the two OUTPUT panels actually showed, and I am reporting it flat rather than picking a winner.

Both runs held the character. Same cream lace dress with the same neckline and sleeve cuffs, same chestnut hair, same fog-and-moonlight corridor lifted off the board. Neither one invented a different woman. The reference pack did that work on both sides, which is the finding I care most about.

The differences were about the form, not the face. This Seedance endpoint delivered 720p; H3 delivered 2K off identical inputs. And the framing split: H3 read "9:16 vertical" straight out of the prompt text and returned vertical, while the Seedance run came back 16:9 because that page carries its own aspect-ratio control and the prompt text alone did not override it. If you are cutting for TikTok, that difference will cost you a run the first time you forget the form field. Seedance 2.5 would likely change the resolution half of this and add region-level re-rolls, and I am not going to pretend I measured that.

Step 5: Fix one shot without re-rolling the whole clip

The real cost of a short drama is not generation one. It is revision seven. H3 accepts an existing clip as input and edits against instructions, holding what you did not mention. Seedance 2.5's pitch here is finer: redraw a region of the frame and leave performance, lighting and camera alone.

plaintext
1Using the supplied clip: keep the two characters, their wardrobe, the camera moves and the cutting rhythm exactly as they are. Change only the environment, replace the moonlit stone corridor with the same castle's ballroom, tall arched windows, a lit chandelier, and warm candlelight, and re-light both characters so their key light comes from the chandelier at frame left instead of the window. Rewrite her only spoken line to "You were never asleep, were you." Keep her performance timing on the same beats.

Settings: MiniMax H3 Reference-to-Video with the Step 3 clip as the reference input, Duration 5, Resolution 2K.

 

 

Beyond the Seedance 2.5 vs MiniMax H3 Test: Four Ways to Push a Reference Pack

Same character, next episode. Save the Step 1 sheet as an asset and change only the shot list and story block in the prompt. Identity comes from a file, not from your memory of how you phrased it last Tuesday.

 

 

Pin the voice too. H3 takes up to three audio references, 2 to 15 seconds each, and they must accompany an image or video input. Give it a voice sample and the character speaks in that timbre, so you are not re-casting a voice actor between episodes.

 

 

Block the camera before you pay for the render. Seedance 2.5's reference-to-video accepts 3D white-model blockouts and green-screen plates, so you can lock framing on grey geometry and only then commit to a finished look. This example ran on Seedance 2.1, an earlier member of the family, but the shape of the workflow is the same.

 

 

Localise one master cut instead of reshooting it. Region-level replacement swaps a face, a product or a spoken language while camera, timing and lip duration stay put. Here one master beauty cut is re-cast and re-voiced for Spanish, Korean and Hindi, with the room, the framing and the lip duration held in place.

 

 

And because someone will ask what motion transfer looks like when you stop being sensible: three men in suits, replaced by three photoreal capybaras, following the original clip's motion path beat for beat until they stack into a pyramid.

 

 

There is also the plain version of all this, the one nobody posts: a static poster becomes a moving poster, and the layout survives. Cartoon boy in pink dinosaur hoodie reaching forward with giant claw The static source poster. Feed it to reference-to-video with "keep the frame, the palette and the layout, animate the type" and you get a moving version of the same design.

What a Seedance 2.5 vs MiniMax H3 Short Film Actually Costs You

No numbers here, because the numbers move and the structure does not. Both families bill per second, pay as you go, no seat and no subscription. That leaves three variables: seconds, resolution tier, and re-rolls.

The third one is the whole bill. A 15-second clip that lands on the first try is one charge. The same clip on the seventh try is seven. So the money-saving move is not picking the cheaper per-second model, it is getting the reference pack right before you generate anything. Those two image calls in Step 1 and Step 2 are the cheapest line in the entire pipeline and the one that decides how many video calls you make. Highest return on spend in the whole chain, by a distance.

Then correct the per-second instinct. Same reference pack, same runtime: H3's reference-to-video hands back 1440p, this Seedance reference-to-video endpoint hands back 720p. Per second is the wrong unit. Pixels per frame, and whether you then have to pay for a separate upscale pass, is the right one.

Audio belongs in the arithmetic too. Both sides produce native synced sound in the same pass. If your current pipeline is video model plus TTS plus an editor aligning tracks, what you save is two API calls and a person's afternoon, and that saving appears on no price list anywhere.

Which one for which job:

JobPickWhyCaveat
Vertical short drama, 9:16, TikTok or ReelShortMiniMax H31440p vertical with native stereo, callable now, 15 s is a full hookyou cut at 15 s, so plan beats accordingly
One continuous 30-second brand filmSeedance 2.5single generation, no stitching, native 4Kcheck availability on your platform first
Fixing one shot inside an approved cutSeedance 2.5region-level redraw leaves the rest aloneH3 whole-clip editing gets you most of the way today
Product and e-commerce cutdownseitherboth hold on-screen text and brand marks wellverify text at the resolution you ship
Game or interface demosMiniMax H3UI and typography stability at 2K is strong15 s ceiling for a single pass
Multi-language versions of one masterSeedance 2.5region replacement plus multi-language voicelocalisation quality still needs a native check
Shipping tonightMiniMax H3three endpoints live, plus all of Seedance 2.0Seedance 2.5 access varies by platform

Before you ship. MiniMax said H3's weights would follow within days of launch, and open weights are not the same thing as a commercial licence, so read the terms before you self-host. Watermark and commercial-use policy also differ between consumer apps and API access on both sides, and I could not confirm either one from a primary terms page while writing this, so check the provider's terms rather than taking my word for it. And nothing in any model licence covers likeness or trademark. If a real person, a real brand or someone else's IP is in your reference pack, that is a separate legal question and the model's terms will not answer it.

Every model in the tables above runs on the same key, and the model pages carry the current rates plus copy-paste code, including whatever promotions happen to be running on the Seedance line this week.

Frequently Asked Questions

Is Seedance 2.5 better than MiniMax H3 for character consistency?

On paper it should be, with up to 50 multimodal references against H3's 12 files, plus region-level re-rolls. But as of 31 July 2026 there is no paired public test I can point to, and Seedance 2.5 has no arena rating yet. In the runs I could actually make, the variable that decided identity stability was reference pack structure, not model generation: with one composite turnaround sheet plus one lighting board, both sides held the character. Get the pack right before you spend time shopping models.

Can I use Seedance 2.5 right now, or do I have to wait?

Its API is live at ByteDance. Availability on third-party model platforms is uneven, and on the one I tested it is still a Day-0 notice. If you need output tonight, the callable options are MiniMax H3's three endpoints and the whole shipping Seedance 2.0 line.

Does MiniMax H3 really do 30-second video?

No. The specification is 5 to 15 seconds per generation at 24fps. Anything longer is extension or stitching, which is a different thing with different drift risk. The 30-second single generation belongs to Seedance 2.5. Do not let the two claims blur.

Do both models generate audio, and can I keep one voice across episodes?

Both produce native synced audio in the same pass, and H3 outputs stereo on every result. For a persistent voice, H3 accepts up to three audio references of 2 to 15 seconds each, which must be submitted alongside an image or video input rather than on their own.

How many reference images should I actually send?

Fewer than you think, structured better than you think. One well-composed composite generally beats six loose crops, because separate files invite the model to average them into a person who does not exist. Give it two clear jobs, identity and lighting, and do not include style samples that fight each other.

Which model should I use for TikTok or ReelShort-style vertical short drama?

Shipping this week, in 9:16, at high resolution with sound already mixed: MiniMax H3. Planning a 30-second continuous scene where you expect to re-roll single regions rather than whole clips: build for Seedance 2.5 and check when your platform opens it.

Najnowsze modele

Jedno API do całej multimedialnej AI.

Przeglądaj wszystkie modele