Seedance 2.5 現已上線 — 首發於 Atlas Cloud

創作者報價875美元 vs 3.16美元:五個MiniMax H3 UGC廣告,按秒計價

五個視覺上不同的 MiniMax H3 UGC 廣告,分別基於五個基本框架:實際每秒計費、15槽並發上限,以及品牌文字斷點。

Twenty clips on the meeting room screen. The client watched for eleven seconds, stopped on the one where the guy pushes the bottle toward the camera, and asked what the label said.

Nobody could tell. The four letters had turned into a smear somewhere between his hand and the lens. That clip cost fifty cents. It was worth nothing.

This is the full ledger for that fifty cents. Five base frames, five completely different shots, one wave of 20 jobs fired at a documented 15-slot ceiling, and native-resolution crops showing exactly where brand text holds and where it goes soft. Every number below came off a real bill this week, not off a pricing page.

The brand is invented on purpose. More on that at the end.

Key takeaways

  1. Five finished MiniMax H3 UGC ads, from five different base frames, cost $3.16 all-in. Five human UGC creators at the market median would be $875 before usage rights (Influee, 2026).
  2. 768P bills at $0.10/s and 2K at $0.14/s on the same endpoint. Both spelled the brand correctly on a static end card. 768P broke when the product moved toward the lens.
  3. H3 is metered by CONN, max concurrent tasks: 2 free, 15 paid (MiniMax Docs, 2026). My burst of 20 all completed, in 4 minutes 13 seconds end to end.
  4. resolution and duration are required fields, not optional. Leave them out and you are buying 2K by accident.
  5. Both tiers ship AAC stereo at 32 kHz. Sound does not improve with resolution, so paying for 2K buys pixels and nothing else.

Five MiniMax H3 UGC Ads From Five Completely Different Shots

Here is the payoff first. One invented skincare brand, ORVA, five archetypes that look nothing alike: a bathroom mirror selfie with a face in it, a faceless macro drop test, a golden-hour street testimonial with a different person entirely, a top-down desk unboxing with only hands, and a typographic end card.

The five finished spots, hard-cut into one reel. Sound on: the dialogue, the room tone and the paper rustle were all generated inside the same pass as the picture, not layered on afterwards. Generated with minimax/h3/image-to-video.

Thumbnails matter more than the reel here, because the whole point is that these five do not look like the same shoot with five different scripts:

Five different ad formats for ORVA skincare

Five MiniMax H3 UGC ad archetypes side by side with aspect ratio and per-clip price

One frame lifted from each finished spot, labelled with its aspect ratio and what that clip actually cost.

Five aspect ratios, and none of them cost extra. H3's image-to-video endpoint only accepts adaptive for ratio, which sounds like a limitation until you realise it means the first frame decides the shape. You buy the aspect ratio once, at image prices, in step 1.


Why Most MiniMax H3 UGC Ads Never Reach the Ad Account

The gap between "I made a video" and "I put this in a live ad account" is where almost every AI UGC experiment dies. Three specific blockers, in the order they usually bite.

BlockerWhat it actually costs you
Free consumer tier: watermark plus a resolution capThe file is unusable as a deliverable. You cannot hand a client a corner logo, and stripping it violates the terms you agreed to.
Credit-metered toolsYou can report a monthly subscription figure to your boss. You cannot report cost per usable clip, which is the only number a media buyer can act on.
One job at a timeTwenty hooks becomes an overnight job. Testing stops being a Tuesday afternoon and starts being a sprint ticket.

The first one is a product boundary, not a bug. The free consumer app and the API are different products with different terms. If the deliverable leaves your building, you are on the paid path, full stop.

The second one is quieter and worse. Most ai ugc video generator products sell a monthly subscription and meter the work in credits, and a credit is not a unit anyone can defend in a budget review. Ask a performance marketer what a single 5-second hook test costs and you will get a shrug or a division problem. The whole reason this article prices things per second is that per second is the only unit that survives contact with a spreadsheet.

The third one is the one people discover last. Batch is the entire value proposition of AI creative. You are not making one perfect ad, you are making twenty mediocre hooks so the algorithm can find the two that work. That means concurrency is a first-class spec, and almost nobody publishes theirs.

Then there is the failure mode nobody prices at all:

Hand holding a phone with a sticky note reading Rejected Wrong Name

A phone showing an AI-generated UGC ad frame where the product wordmark is misspelled, with a rejection note stuck beside it

The rejection nobody budgets for. The lighting is fine, the performance is fine, the four letters on the bottle are wrong, and the spot is worth zero.

This is exactly the ground MiniMax picked a fight on. Their launch post claims H3 excels at "instruction following, accurate text and brand rendering" and says the model is "built for advertising, branding, e-commerce" (MiniMax, July 2026). That is a testable claim, and step 4 tests it.

It matters because readability is the variable that moves the needle. In a study of Taboola platform data covering more than 500 million impressions and 3 million clicks, researchers from Columbia, Harvard, TUM and Carnegie Mellon found AI-generated ads averaged a 0.76% CTR against 0.65% for human-made ads, and that once strict statistical controls were applied the two were essentially level (Taboola, January 2026). The strongest single predictor was not "made by AI" or "made by a human", it was whether the creative carried a large, clear human face and read as authentic. Nobody trusts a bottle whose label is spelled wrong.


The MiniMax H3 UGC Ads Workflow, and What Each Second Costs

The chain is three moving parts. An image model nails the shot down and buys the aspect ratio. A video model buys seconds on top of that frame. A rival model at a near-identical listed price runs the control, so the comparison is not just me quoting one vendor.

All three live in the same catalogue on Atlas Cloud, which is the only practical reason I ran it this way: one key, one invoice, and the per-second rate is the same number in the bill as on the model card.

Chain stepModelPriceDiscounted?
Five base framesgoogle/nano-banana-2/text-to-image1k $0.08 / 2k $0.12 / 4k $0.16 per imageNo
Drafts and finalsminimax/h3/image-to-video768P $0.10/s, 2K $0.14/sNo
No base frame at allminimax/h3/text-to-videoSame two tiersNo
Same face across a seriesminimax/h3/reference-to-videoSame two tiers; refers takes any mix of images, video and audioNo
Control modelbytedance/seedance-2.5/image-to-videoListed from $0.134/s, caps at 720p, 4 to 30s. Read the variations section before you budget with that numberNo
Cheap hook-test tierbytedance/seedance-2.0-mini/image-to-video$0.039/s, down from $0.056Yes, 30% off as of August 2026
Keep the take, add pixelsatlascloud/video-upscaler$0.018/s to 1080p, $0.024/s to 2K, 5s minimum chargeNo

Webpage interface displaying three MiniMax H3 video generation model cards

The three MiniMax H3 endpoints on the Atlas Cloud model catalogue, each showing the From $0.10 per second price card

The three H3 endpoints in the catalogue. The card price is the floor, which is the 768P tier at $0.10 a second; 2K is $0.14 a second on the same endpoint.

One honest footnote before anyone fact-checks me: MiniMax's own pay-as-you-go page lists 2K at $0.13/s and 768P at $0.08/s, so buying direct is a cent or two cheaper per second. What the extra cent buys here is that the image model, both video models and the upscaler sit behind one key, which is the difference between a workflow and four workflows.


How to Make MiniMax H3 UGC Ads: The Four-Step Run

Four steps. Every prompt below is the exact string I pasted, and every setting is the exact setting I picked.

Step 1: Build Five Different Base Frames, Not Five Hooks Off One

This is where the article's whole argument lives. The common failure is one product photo and six different scripts, which produces five clips that are visually identical and one bored reader. Instead, spend $0.60 and buy five genuinely different shots: with a face and without, indoors and outdoors, eye-level and top-down, two different people.

Model: google/nano-banana-2/text-to-image, resolution: 2k, $0.12 each. I picked this over a general photo model for one reason: the only acceptance criterion at this stage is that the four letters on the bottle are spelled correctly. Put the best text-rendering model first and you fix spelling at $0.12 instead of at $0.70.

Check every frame for three things before you move on. Wordmark spelled right. No hand covering the wordmark. Enough headroom and footroom that captions and platform UI will not sit on the product.

A. Mirror morning routine (aspect_ratio: 9:16)

text
1Vertical phone photo shot in a small bathroom mirror: a 29-year-old woman with damp hair and
2no makeup holds a 30 ml amber glass dropper bottle up beside her cheek, looking at her own
3reflection. The bottle's front label reads "ORVA" in bold cream sans-serif, with a smaller
4line "10% NIACINAMIDE · 30 ML" beneath it. Warm tungsten vanity light from above, slightly
5steamy mirror edge, white tile, a toothbrush cup and a folded grey towel in frame.
6Photoreal, handheld, imperfect exposure, no on-screen text overlay, no watermark,
7no logo other than the bottle label.
8

B. Macro drop test (aspect_ratio: 1:1)

text
1Extreme macro photograph, no face in frame: a hand squeezes a glass dropper and one clear
2viscous drop is about to fall onto the back of the other hand. Cool north-window daylight,
3shallow depth of field, visible skin texture and fine hairs, the amber "ORVA" bottle
4standing softly out of focus behind. Clean pale-grey background. Photoreal product macro,
5no text overlay, no watermark.
6

C. Golden-hour street testimonial (aspect_ratio: 9:16)

text
1Vertical handheld selfie video still, front camera: a 34-year-old man with short beard and a
2navy crewneck walks along a city sidewalk holding the small amber "ORVA" dropper bottle up
3toward the lens, talking to camera. Strong golden-hour backlight with lens flare, blurred
4shopfronts and parked bicycles behind him, slight motion blur at the frame edges.
5Photoreal, natural skin, no colour grading, no text overlay, no watermark.
6

D. Top-down desk unboxing (aspect_ratio: 16:9)

text
1Straight top-down flat-lay photo, hands only, no face: two hands lift a small amber dropper
2bottle out of an open kraft-paper box lined with cream tissue paper, on a pale oak desk.
3The box's inner lid is printed with "ORVA" in cream sans-serif and a smaller line
4"NIACINAMIDE SERUM · MADE IN SPAIN". A folded card, scissors and a linen cloth sit beside it.
5Soft diffused daylight from the left, real shadows. Photoreal e-commerce unboxing,
6no text overlay, no watermark.
7

E. End-card still (aspect_ratio: 9:16)

text
1Vertical product still: the 30 ml amber "ORVA" dropper bottle standing centred on a seamless
2warm-sand backdrop, single soft key light from camera right, long clean shadow.
3Typographic layout in the lower third: "ORVA" in bold cream sans-serif, under it
4"20% OFF YOUR FIRST BOTTLE" in a lighter weight, and at the very bottom "orvaskin.com".
5Crisp letter edges, generous spacing, no clutter, no extra logos, no watermark.
6

One disclosure on this step. I tried three times to capture the Nano Banana 2 playground mid-run for this article and all three attempts died on an upstream connection error from the model provider on the test environment, so there is no browser screenshot here. The five frames below came from the API instead, at the same settings and the same $0.12 per image the playground quotes.

Five different ad creative options for ORVA skincare

The five ORVA base frames side by side, each labelled with its aspect ratio and $0.12 cost

All five base frames. Five aspect ratios, $0.60 total, and the video step inherits every one of those shapes for free. Generated with google/nano-banana-2/text-to-image.

Step 2: Buy Seconds and Turn Each Frame Into a 768P Draft

Now each frame becomes motion. Model: minimax/h3/image-to-video. Settings that matter:

  • image: the matching base frame from step 1, as a URL or a base64 data URL.
  • resolution: "768P" for drafts. This is a required field. The default is 2K, and the two tiers differ by 40% per second.
  • duration: 5. Also required. Leaving it out does not give you a free short clip.
  • ratio: ignore it. The enum on this endpoint is adaptive only, and the first frame decides the shape.

That pair of required fields is the single most expensive thing to get wrong in this whole workflow. A wave of 20 jobs submitted on defaults is 20 clips at 2K and 8 seconds, which is $22.40 instead of $10.00, for drafts you were going to throw away.

text
1A: Handheld vertical phone video in the mirror. She tilts the bottle so the label faces the
2mirror, then looks back at her own eyes and says, in a dry, slightly amused tone:
3"Two weeks in. My skin is boring now. That's the whole review." Bathroom room tone, a faint
4extractor fan hum, no music. One continuous take, no cuts. Keep the bottle label sharp and
5unchanged.
6
7B: Locked-off macro shot, no camera move. The drop detaches from the dropper, falls, and
8spreads slowly on the back of the hand; the fingers tilt to catch the light. No speech.
9Only sound design: a soft liquid touch, a faint room tone. One continuous take.
10
11C: Handheld walking selfie video, natural bounce. He keeps eye contact with the lens, glances
12down at the bottle once, and says in a casual American accent: "I have never had a skincare
13routine. I have one step now. This is the step." Street ambience, distant traffic, wind on
14the mic, no music. One continuous take, no cuts.
15
16D: Locked-off top-down shot. The hands finish lifting the bottle clear of the box, set it
17upright on the oak desk, then fold the tissue paper back down. No speech. Paper rustle,
18a soft glass-on-wood tap, quiet room tone. Keep the printed text on the box lid sharp,
19still and unchanged.
20

Screenshot of an AI video generator interface showing input and output

MiniMax H3 image-to-video playground on Atlas Cloud with the base frame loaded and the finished clip playing in the output panel

The MiniMax H3 image-to-video playground, prompt pasted and base frame loaded, the finished clip in OUTPUT. This capture is deliberately left on the untouched defaults, and the Run button prices them for you: 2K at 8 seconds is $1.12. The drafts in this tutorial are 768P at 5 seconds, which is $0.50. That gap is the whole reason both fields are marked required.

Delivered geometry from this run, read straight off the files. 768P puts 768 pixels on the short side: 768x1376 vertical, 1376x768 landscape, 768x768 square. 2K puts 1440 there: 1440x2560 vertical, 1440x1440 square. Both at 24 fps, both carrying AAC stereo at 32 kHz. The sound spec does not change with the tier, so the extra four cents a second buys pixels and nothing else.

Step 3: Fire the Whole Wave Into the 15-Slot Ceiling

Do not iterate one clip at a time. That habit is a holdover from tools that could only do one thing at a time, and it is the reason people think AI creative is slow.

H3 is not metered in requests per minute like the older Hailuo models. It is metered in CONN, maximum concurrent tasks: 2 on the free tier, 15 on the paid tier (MiniMax Docs, 2026). On paper that is a very specific consequence: 15 is your batch width, and job 16 waits its turn instead of erroring.

So I submitted 20 jobs in one burst and timestamped every one of them. Four archetype spots, the end card at both resolutions, and 14 variants that change the line, the gesture or the camera move. Eighteen of the twenty were 768P at 5 seconds; the two end-card jobs were 4 seconds, one at 2K and one at 768P.

text
1Variant recipe: keep the base frame and the shot identical, change exactly one thing.
2- Swap the spoken line only (same framing, same gesture): 4 variants off A and C.
3- Swap the gesture only (same line, same framing): 3 variants off A and D.
4- Swap the camera move only (locked-off to slow push, or handheld to steadier):
5  3 variants off B and D.
6- Swap the closing beat only (bottle lowered, bottle set down, cut on the smile):
7  4 variants off A, C and D.
8Submit all of them in one loop. Do not await one before submitting the next.
9

Horizontal bar chart of completion times for twenty MiniMax H3 jobs

Timeline of 20 MiniMax H3 jobs submitted in one burst, each bar measured from submit to completed

Every bar is one job, measured from my submit call to the completed status. All 20 went in inside 8.1 seconds and all 20 came back.

What came back was not what the ceiling made me expect.

All 20 submissions returned HTTP 200 inside 8.1 seconds. All 20 jobs completed. The fastest landed at 114.3 seconds, the median at 175.6 seconds, and the slowest at 253.0 seconds. Eighteen of the twenty finished inside a tight 114 to 196 second band. The 2K end card took 222.7 seconds, which matches everything else I have measured about 2K on this model: it is slower, but not dramatically. Exactly one 768P job, a mirror-scene line variant, came in at 253 seconds, and it is the only bar in the chart that looks like it queued.

End to end, twenty 5-second UGC clips in 4 minutes 13 seconds, for $9.96. That is the number to hold onto: a hook-testing wave that used to be an overnight job is now shorter than a coffee break.

Two things this does and does not prove. It does not prove the 15-slot ceiling is fictional, because a routed provider can pool capacity in ways a single direct account cannot, and one wave on one afternoon is one data point. It does prove that the practical batch width is high enough that you should stop writing sequential loops. Submit the wave, then go read something.

One caution, because I have measured this endpoint on several different days: concurrency behaviour is not a constant. On an earlier run six overlapping jobs all completed cleanly; on another, an overlapping job came back with 429 rate limit exceeded (task concurrency). Treat the numbers above as what my wave did on the day, and measure your own before you build a scheduler on top of it.

Step 4: The End Card, Where MiniMax H3 UGC Ads Live or Die on Text

The end card is the exam. A wordmark, an offer line and a URL, all of which have to be readable on a phone at arm's length, and one of which is a domain where a single wrong character sends the click nowhere.

Run it twice off the identical base frame with the identical prompt: once at resolution: "2K", duration: 4, once at resolution: "768P", duration: 4.

text
1The bottle rotates about 12 degrees to camera right while every letter stays perfectly still:
2letter shapes locked, no warping, no flicker, no re-lettering. One soft specular highlight
3travels across the glass. Silence for three seconds, then one short clean UI chime.
4

Locking the first frame is not optional here, and this is the least intuitive thing in the whole workflow. H3's 2K tier is in-context regeneration, not upscaling. Give the two tiers the same text prompt with no base frame and you get back two different films, with different set dressing and different label positions, which tells you nothing. Pin both runs to the same first frame and the only thing left that can differ is detail. Then the comparison means something.

Side by side comparison of Orva ads at different video resolutions

Side by side end card, 768P on the left and 2K on the right, same first frame and same prompt

Same base frame, same prompt, four seconds each. Left 768P, right 2K. Shown here as a silent GIF; both delivered files carry AAC stereo at 32 kHz.

And here is the part that actually settles it:

Let's try: "Side-by-side comparison of 768P and 2K video resolution

Pixel-level crop comparison of the ORVA wordmark, the ingredient line and the URL at 768P versus 2K

Top row is the bottle label, bottom row is the end-card type block. Native crops were 292x220 and 691x372 pixels at 768P against 547x410 and 1296x691 at 2K; both are shown at one display width, which is how a phone renders them.

Here is the honest result, and it is not the one I expected to write.

768P did not misspell anything. ORVA, 20% OFF YOUR FIRST BOTTLE and orvaskin.com all came out correct, stable and legible at both tiers. Nothing flickered, nothing re-lettered, nothing warped as the bottle rotated. On a near-static end card, MiniMax's "accurate text and brand rendering" claim held at the cheap tier.

What 768P loses is edge quality. Look at the counters inside the O and the R, and at 10% NIACINAMIDE / 30 ML on the label. At 2K the letterforms have clean edges; at 768P they are soft, and the small line on the bottle is on the edge of guessing. On a phone at arm's length you would probably pass both. On a tablet, in a client review, or the moment anyone pauses and zooms, the 768P version reads as a cheap render.

Where 768P actually broke was motion, not resolution. In the wave, one street variant ends with the bottle pushed toward the lens. As it fills more of the frame and tilts, the four letters dissolve into a smear that reads like a different alphabet. Same tier, same model, same brand, unusable clip. That is the real rule: text survives when it is small, flat and still, and it breaks when it moves fast through the frame. Buying 2K does not fix a shot that asks the model to re-letter a rotating label mid-move. Designing the shot does.


MiniMax H3 UGC Ads vs the Best AI Video Generator for Ads at the Same Listed Price

Anyone shopping for the best ai video generator for ads ends up comparing these two, because the catalogue puts them within half a cent of each other per second. So I ran base frame B, the macro drop test with the hardest small print in the set, through both with an identical prompt, and then read both bills.

Two comparison images of a dropper dripping serum onto a hand

Same base frame animated by MiniMax H3 at 2K on the left and Seedance 2.5 at 720p on the right

Identical first frame, identical prompt, five seconds each. Left MiniMax H3 at 2K, right Seedance 2.5 at 720p. Both clips came from the API: the Seedance playground run was still showing "In progress" after three capture attempts, so there is no browser screenshot for this step either.

Generated with minimax/h3/image-to-video and bytedance/seedance-2.5/image-to-video.

MiniMax H3 (image-to-video)Seedance 2.5 (image-to-video)
Listed price$0.14/s at 2K, $0.10/s at 768PFrom $0.134/s
What this 5s clip actually billed$0.70 (2K, 1440x1440)$1.51 (720p, 960x960)
Resolution ceiling2K, 1440 on the short side720p, which was 960x960 from this square source
Clip length4 to 15s4 to 30s
Reference materialsAny mix of images, video and audioUp to 30 images, images only
Native audioYes, 32 kHz stereo, generated with the pictureYes
Small print on the label, this runWordmark and ingredient line both readable at 2KWordmark readable, ingredient line soft
Latency, this clip190.5s224.7s

Read the first two rows together, because that is the whole surprise. The listed prices are almost identical. The bills are not.

H3 billed exactly its listed rate, both times: $0.70 for five seconds at 2K, $0.50 at 768P, no rounding, no surprise. Seedance 2.5 billed $1.5148 for five seconds at 720p, which is about $0.30 a second, not $0.134. So I checked why, and ran the same shot again at 480p: $0.5397 for four seconds, or $0.1349 a second, matching the listed number almost to the cent.

That explains it. The catalogue figure is the 480p rate, and 720p on a square source is 960x960 against 640x640, which is 2.25 times the pixels and, in my bill, 2.25 times the price. Nothing dishonest about a "From" price, but "From" is doing real work in that sentence, and it is the kind of thing you find in the invoice rather than the comparison table.

Which flips the decision cleanly. On this run H3 delivered 1440 pixels on the short side for $0.70 while Seedance 2.5 delivered 960 for $1.51. If your ad is a person talking and nothing on screen has to be read, Seedance 2.5 is still worth having: it goes to 30 seconds against H3's 15, which matters for longer storytelling formats, and it accepts far more reference images. The moment a wordmark, an offer line or a URL has to survive a phone screen, H3 is both the sharper and the cheaper option, and it is not close.

The general lesson outlives both models: run one clip at your real settings and read the actual bill before you put a per-second figure in a spreadsheet.

Three cheaper variations worth knowing, all verified on the catalogue this month:

Split the buy. Two 5-second clips beat one 10-second clip for hook testing, because you get two openings instead of one. Same money, twice the top-of-funnel.

Draft on the cheap tier. For pure hook tests where nothing has to be read, bytedance/seedance-2.0-mini/image-to-video is $0.039/s, currently 30% off its $0.056 list price as of August 2026. That makes a 5-second hook about 20 cents, which is cheap enough to be careless with.

Lock the face with reference-to-video. Feeding the same reference images into minimax/h3/reference-to-video keeps one presenter consistent across scenes, which is how you get a month of content out of one character instead of a new stranger every week. It bills at the same two tiers. I did not run it for this article, so I am not going to show you a result I do not have.

Keep the performance, add pixels. When a 768P take is perfect and only the resolution is wrong, atlascloud/video-upscaler preserves that exact take at $0.024/s to 2K with a 5-second minimum charge. It cannot invent text the source never resolved, so it rescues a good performance, not a broken label.


What MiniMax H3 UGC Ads Cost Per Usable Clip

Sticker price first, then the number that actually matters.

Line itemThis runHuman UGC creator equivalent
Five base frames, 2k5 x $0.12 = $0.60Included in the shoot
Four spots, 768P, 5s each4 x $0.50 = $2.004 x $175 = $700
End card, 2K, 4s$0.56$175
Five deliverable spots$3.16$875 before usage rights
Paid-ad usage rightsIncluded+30% to +50%, so $1,138 to $1,313
The 768P end-card control that proved the point$0.40Not a thing you can buy
The whole 20-job hook-testing wave$9.96Not a thing you can buy

The creator column uses the market median of about $175 per video, against a typical range of $150 to $212, with a further 30% to 50% on top for paid distribution rights (Influee, 2026).

Now the honest part, because sticker price assumes every clip is usable and no clip ever is.

I watched all 20 clips from the wave and rejected three. One had the wordmark smear as the bottle came toward the lens. One had the presenter's fingers wrapped across the label for the whole take. One drifted so far that the printed box lid, which was the entire point of that shot, left the frame.

That is 17 usable out of 20, an 85% hit rate, on $9.96 of spend. Cost per usable clip: $0.59, against an average sticker price of $0.50.

Two things worth saying about that number. First, 85% is high, and I would not promise it holds on a harder brief; a wave with faster motion or denser packaging text would fail more. Second, every one of the three rejects was a composition problem I could have prevented in the prompt, not a model defect. Rejects are a design tax, not a rendering tax.

That is the number to put in the deck: cost per usable clip, not cost per clip. And you cannot know it until the whole wave has finished and you have watched all of it. Anyone quoting you a per-clip figure without a usability rate is quoting you the cheaper half of the equation.

Worth saying plainly: this does not make human creators obsolete. It makes the test phase cheap. Find the two hooks that work for a few dollars, then spend the $875 on the version that goes behind real budget, with a real person whose face people will follow.


Disclosure and Brand Rules for MiniMax H3 UGC Ads

Three things, none of them optional, all of them short.

Declare it. Both Meta and TikTok require advertisers to flag creative that is AI-generated or substantially AI-modified, and both run automated detection, including C2PA Content Credentials embedded in files by many generation tools. Being AI-made is not a rejection reason. Failing to declare it is. Policy wording changes often enough that the only sane instruction is: read the current text in your own Ads Manager before you ship, not a blog post's summary of it.

Do not put a real trademark in a prompt. This is why ORVA is invented. MiniMax's API terms include a defence-and-indemnity clause that covers patent and copyright claims arising from model output, and it does not extend to trademark or likeness. So the one category of risk you would most want covered when generating brand-heavy advertising is precisely the category that is not. Use your own marks, which you have the right to use, or invented ones.

Open weights are not a jurisdiction workaround. H3-Base is genuinely downloadable, and it is a real 33B release, not a stripped demo. But the community licence currently excludes the EU, the UK, South Korea and the United States, with a separate formal licensing channel for those territories. And self-hosting runs a 768-pixel short side, with the Context-IR and Regenerate-2K stages staying API-side. So the local path cannot deliver the one thing this article proved you need for readable brand text.


MiniMax H3 UGC Ads: Frequently Asked Questions

How much does one MiniMax H3 UGC ad cost?

A 5-second vertical spot costs $0.50 at 768P or $0.70 at 2K on minimax/h3/image-to-video, plus $0.12 for the base frame if you generate one. A 4-second end card at 2K is $0.56. Buying direct from MiniMax is slightly cheaper per second, at $0.13/s for 2K and $0.08/s for 768P on their pay-as-you-go page, against $0.14 and $0.10 here.

How many MiniMax H3 UGC ads can I generate at once?

H3 is metered by concurrent tasks, not requests per minute: 2 simultaneous tasks on the free tier, 15 on the paid tier. Submitting more does not throw an error, it queues. In my run, 20 jobs submitted in one burst all completed, the median in 175.6 seconds and the whole wave in 4 minutes 13 seconds. At that rate a working day is several hundred 5-second clips, far more creative than any team can actually review. Measure your own wave before you build a scheduler on it, because concurrency behaviour shifts.

Do MiniMax H3 UGC ads come out watermarked?

Not from the API. I stepped through the delivered files frame by frame and found no corner mark on any of them. The free consumer app is a separate product with a watermark, a resolution cap and terms that forbid removing the mark. If the file is a client deliverable, use the API path.

Will MiniMax H3 render my brand name and CTA text correctly?

Mostly yes, with one clear condition. On a near-static end card, both 768P and 2K rendered ORVA, a discount line and orvaskin.com correctly and held them steady through the shot; 2K only bought sharper letter edges and readable small print on the bottle. Where it failed in my run was motion: when the product was pushed toward the lens and tilted, the wordmark smeared into nonsense at 768P. Three rules that made the difference in this run: put critical text on a near-static end card rather than in a moving shot, never let the camera travel across the letters, and check on the base frame that no hand or finger crosses the wordmark before you spend a cent on video.

Is MiniMax H3 the best AI video generator for ads, or is something cheaper enough?

For pure talking-head hook tests where nothing has to be read, the cheap tier is genuinely enough: bytedance/seedance-2.0-mini/image-to-video is $0.039/s, currently 30% off its $0.056 list price as of August 2026, which makes a 5-second hook about 20 cents. The moment a wordmark, an offer line or a URL has to be legible, the 2K tier's four extra cents per second pays for itself the first time a spot does not get sent back.

Will Meta or TikTok reject an AI-generated UGC ad?

Not for being AI-generated. Both platforms ask advertisers to declare AI-generated or substantially AI-modified creative, and both run automated detection including C2PA Content Credentials. The rejection risk sits in not declaring, not in the technique. Check the exact current wording in your own Ads Manager, since both platforms revise this language frequently.

最新模型

一個 API,暢享全模態 AI。

探索全部模型