TWO WEEKS ONLY | 20% OFF Seedream 5.0 Pro!

MiniMax H3 API Pricing in 2026: $0.13/Second, a Third Tier Nobody Lists, and the 6 Limits That Decide Your Bill

MiniMax H3 API pricing, decoded: $0.13/s 2K, $0.08/s 768P, a $0.05/s tier nobody lists, concurrency 2 vs 15, i2v/r2v exclusivity, 64MB cap, and a real invoice.

You already decided to wire up the API. You want two numbers: cost per second, and how many jobs you can run at once.

The first eight results will hand you the first number. Most of them get the 768P tier wrong. Not one of them gives you the second number, because concurrency limits are bad for conversion.

There is also a clock on this. OpenAI is removing the Videos API and every Sora 2 alias on 24 September 2026 (OpenAI Deprecations, March 2026). If you are migrating off it, you have about seven weeks to pick a replacement and ship.

So here is the full rate card, the limits that will actually bite you, and a line-by-line invoice from a real 10-second spot. We spent $9.23 on 3 August 2026 to write this, including two requests we deliberately broke.

Key takeaways

  • 2K output: $0.13 per second. Official pay-as-you-go rate. A 10-second clip is $1.30.
  • 768P output: $0.08 per second, not $0.09, and there is no beta label on the official page.
  • There is a third tier: 768P to 2K regeneration at $0.05 per second. Almost nobody lists it.
  • $0.08 + $0.05 = $0.13. Drafting cheap then upgrading costs exactly the same as shooting 2K directly.
  • Concurrency is 2 on free, 15 on paid. That is concurrent tasks, not requests per minute.
  • Two different defaults for duration. The UI quoted 8 seconds at $1.12; the API, with duration omitted, billed 5 seconds at $0.70.
  • Hard caps: 64MB total request body, 9 reference images, 3-second callback challenge window.

An invoice on a wooden desk with a pen and lemon water

A printed itemised invoice on a dark walnut desk next to a glass of citrus soda with condensation, mechanical pencil resting across the rows of dollar amounts

Before the tables, the thing the tables are about. One 10-second beverage spot, 2560x1440, native stereo audio, generated image-to-video from a single locked frame.

Tall glass filled with ice cubes and a lime slice

The delivered spot. Billed $1.40 on Atlas Cloud, $1.30 at MiniMax list price, 10.13 seconds long, 24fps, stereo AAC. Every call that led here is itemised at the bottom of this article.

MiniMax H3 API Pricing: Every Rate in One Table

MiniMax bills H3 per second of output, at three different rates depending on what you ask for. Two are widely quoted. The third is the interesting one.

T1. The full rate card (MiniMax pay-as-you-go pricing, August 2026):

What you are billed forRate
Standard generation, 2K output$0.13 / second
Standard generation, 768P output$0.08 / second
Regeneration, 768P to 2K upgrade$0.05 / second
Reference audioFree
Reference imagesFirst 5 free, then $0.04 each ($0.025 on regeneration)
Reference videoInput duration billed at the output resolution rate
H3-Context-IR tokens$0.90 / M input, $3.60 / M output

Two corrections worth having, because they change your budget.

The 768P tier is $0.08, not $0.09. The official page lists 768P | Billed per second | $0.08 / second with no beta marker on it. The $0.09 and $0.10 figures circulating in roundups and reseller pages are second-hand numbers copied from each other. If you budgeted at $0.09, you have 12.5% of headroom you did not know about.

There is a third tier nobody lists. MiniMax-H3-Regeneration, which upgrades a finished 768P render to 2K, is billed at $0.05 per second. It exists only as a separate endpoint on MiniMax direct, which is why every spec-sheet copy job misses it.

T2. Price per clip (the table you are going to screenshot):

Duration2K, official768P, official768P then upgrade2K on Atlas Cloud
4s$0.52$0.32$0.52not exposed
5s$0.65$0.40$0.65$0.70
10s$1.30$0.80$1.30$1.40
15s$1.95$1.20$1.95$2.10

Look at column four. $0.08 per second for the draft plus $0.05 per second for the upgrade equals $0.13 per second, which is the direct 2K rate to the cent. The cheap-draft path does not save money on clips you keep. It saves money only on clips you throw away, at exactly $0.05 per discarded second. Bin one 5-second take and you saved $0.25. Bin none and you paid the same and waited twice.

Where This Test Ran, and What Each Endpoint Costs

Three H3 endpoints and one image model did all the work, and they sit in one browser tab so there is no key juggling between the still frame and the video steps.

Job in this spotModelRate
Lock the composition (one still frame)GPT Image 2$0.11861 measured at 1280x720, $0.1745 quoted at 2048x1152, quality high
Test take and delivery takeMiniMax H3 Image-to-Video$0.14 / second
Second visual language, no base frameMiniMax H3 Text-to-Video$0.14 / second
Style transfer, and the exclusivity testMiniMax H3 Reference-to-Video$0.14 / second

Be honest about the trade. Atlas Cloud lists H3 at $0.14 per second against MiniMax's $0.13, so a 10-second clip costs $0.10 more. In exchange you skip the free-tier concurrency ladder, run the image model and all three video endpoints off one key and one balance, and submit-time rejections cost nothing. None of the three H3 endpoints is discounted as of August 2026. Atlas also exposes 2K only, so the draft-then-upgrade path is not available there. Run the numbers for your own volume.

T4. Per-second rates for comparison, all verified on the Atlas model list in August 2026:

ModelPer secondNative audioNotes
MiniMax H3, 2K$0.14Yes, stereoUp to 15s
HappyHorse-1.1$0.14Noi2v, t2v, r2v
Wan 2.7$0.10Noi2v, r2v
Youchuan V8.2$0.086Noi2v
Seedance 2.0 Mini~$0.056NoCheapest per second on the list

Per-second price alone is a bad comparison. H3 at $0.14 ships a native stereo track baked into the same call. If a $0.056 clip then needs a separate sound pass, the gap closes fast.

Why Your MiniMax H3 API Pricing Estimate Comes In Low

Five things inflate the bill past seconds x rate. All five are measured, on 3 August 2026.

Trap 1: duration has two different defaults. Open the playground and leave the duration slider alone: it sits at 8 and the Run button quotes $1.12. Call the API and omit duration entirely: our task came back billed at $0.70, which is 5 seconds. Same model, same day, a 60% gap in the per-clip price depending on which door you walked through. Pass duration explicitly on every single call or your forecast will never match your statement.

Trap 2: a broken reference image does not fail, it bills. We pointed image_url at a URL that returns 404 and submitted. The task did not error. It reported completed, rendered a video anyway, and charged the full $0.70. Nothing in the response flags that your reference never loaded. Validate that your URLs resolve before you spend.

Trap 3: reference video is billed twice. Input video is billed by input duration at the output resolution rate. Feed a 5-second reference clip to produce a 5-second 2K output and you pay $0.65 for the input plus $0.65 for the output. $1.30 for a five-second deliverable.

Trap 4: image six onward costs money. First five reference images are free, then $0.04 each. Run the full nine and you add $0.16 per call. On the regeneration tier extra images drop to $0.025.

Trap 5: draft-then-upgrade has a break-even point. Covered above, and worth repeating because the arithmetic is counterintuitive: savings are $0.05 per second of footage you discard, and zero on footage you keep.

The one piece of good news. A request rejected at submit time genuinely costs nothing. We sent an invalid ratio value and got HTTP 400 back with outputs: null and no price field on the record at all. Front-load your validation and the failures are free. It is the requests that succeed on paper that cost you.

The MiniMax H3 API Pricing Limits Nobody Publishes

This is the part reseller pages skip, because none of it helps them close. All of it decides whether your architecture works.

T3. Every limit that matters, from the MiniMax API reference and rate limits page (August 2026):

LimitWhat the docs sayWhat it costs you
ConcurrencyVideo Generation V2 | MiniMax-H3CONN (max concurrent tasks)21530 clips means 15 sequential waves on free, 2 on paid. Concurrent tasks, not RPM.
i2v vs r2vMutually exclusive: if any reference_image / reference_video / reference_audio appears, first_frame / last_frame must not, and vice versaThe API does not enforce it. It bills you and drops one of your inputs. See step 5.
Audio alone"audio alone is not allowed, at least one reference video or image is required"An audio-only reference request is rejected. Pair it or lose it.
Request bodyTotal under 64MB; single image under 30MB, video under 50MB, audio under 15MBBase64 inflates payloads by about a third, so 64MB of body is roughly 48MB of real files. Use public URLs.
Reference countsMax 9 images, 3 videos, 3 audioImages 6 through 9 are billable.
Image constraintsJPG, JPEG, PNG, WEBP, HEIC, HEIF; side length 256 to 5760 px; aspect ratio 0.4 to 2.5A 9:21 crop sits at 0.43 and barely clears it. Validate before upload.
CallbackReturn the challenge value unchanged within 3 secondsA cold-starting serverless function misses that window. Answer the challenge first, do the work after.
duration enumOfficial reference accepts 4 through 15Hosted schemas commonly start at 5. Do not assume 4 works everywhere.

One more thing that only shows up when you build billing on top of this: the price field fills in late. Polling until a task reports completed frequently returns no price at all. Query the task ID again a few seconds later to get the real charge. That second lookup is the only way to build an invoice that matches your statement, and it is how the table at the bottom of this article was assembled.

A caution on concurrency: day-to-day behaviour is noisy. We have seen six overlapping H3 tasks finish clean and we have seen a single task get a 429. Design against the published 2 and 15, then measure your own account with two overlapping requests before you ship.

Tutorial: Build the Spot and Watch MiniMax H3 API Pricing Add Up

Every step is reproducible. Settings stay constant: resolution: 2K, ratio: 16:9, and duration always passed explicitly. On text-to-video, ratio is mandatory and adaptive is rejected outright with invalid params, ratio is required for t2va (text-only) and cannot be 'adaptive'.

Step 1. Lock the composition with one still frame. A $0.12 image is roughly ten times cheaper than re-shooting a $1.40 clip, so buying certainty here is a pricing decision, not an aesthetic one. Use GPT Image 2 at quality high, 16:9.

text
1A tall glass of ice-cold citrus soda on a wet black stone bar top, condensation
2beading and sliding down the glass, a single lime wheel on the rim, backlit by a
3warm amber pendant lamp with deep shadows behind, shallow depth of field,
4commercial beverage photography, crisp macro detail on the water droplets,
516:9, photorealistic
6

AI image generator interface with text prompt and generated drink image

GPT Image 2 playground on Atlas Cloud with the beverage prompt entered, quality set to high, 16:9 at 2048x1152, and the generated base frame in the output panel

GPT Image 2, quality high, 16:9. The Run button quotes $0.1745 at 2048x1152. The same prompt at 1280x720 billed us $0.11861.

Step 1 output at 1280x720. This single frame is the input for steps 2, 3 and 5.

Step 2. Buy a cheap test take first. Five seconds, not ten. You are checking whether the physics of the condensation and the ice hold up before committing to full length. Endpoint: H3 image-to-video, image set to the Step 1 frame, duration: 5, resolution: 2K, ratio: 16:9.

text
1The condensation droplets slide down the glass, ice cubes shift and settle with a
2soft clink, the lime wheel trembles slightly, warm amber light flickers across the
3wet stone surface, slow subtle push-in, ambient bar room tone with a faint fizz of
4carbonation
5

A sweating glass of iced water with a lime slice

The $0.70 test take. Half the price of the delivery clip, and it answers the only question that matters at this stage.

Step 3. Shoot the delivery take. Same endpoint, same base frame, duration: 10. Billed $1.40. This is the clip at the top of the article.

text
1Slow cinematic push-in on the citrus soda glass, condensation runs down in three
2distinct rivulets, ice cubes rotate and settle, carbonation bubbles rise in
3continuous streams, the lime wheel catches a specular highlight as the camera
4drifts, warm amber key light with cool rim light from the left, ambient bar tone
5with a crisp carbonation fizz building through the shot
6

AI video generator interface showing numbered steps for input and output

MiniMax H3 image-to-video playground with the base frame uploaded, resolution 2K, the duration slider untouched at 8, a $1.12 quote on the Run button, and the finished clip in the output panel

This is trap 1 in one frame. Nobody touched the Duration slider, so it sits on 8 and the quote reads $1.12 for an 0:08 result. Our API call for the delivered spot passed duration: 10 and billed $1.40; the same call with duration omitted billed $0.70.

Step 4. Prove the rate buys more than one look. Same $0.14 per second, completely different visual language, no base frame at all. Endpoint: H3 text-to-video, duration: 5, resolution: 2K, ratio: 16:9 set explicitly.

text
1Neon-lit late-night ramen counter in heavy rain, steam curling off a fresh bowl,
2a chef's hands dusting scallions in slow motion, reflections of pink and cyan
3signage rippling in a puddle on the counter, handheld documentary feel, shallow
4focus, rain hiss and sizzling broth ambience
5

AI video generator interface showing prompt settings and completed video output

MiniMax H3 text-to-video playground with the ramen prompt, aspect ratio set explicitly to 16:9, resolution 2K, and the generated clip playing in the output panel

H3 text-to-video. Ratio has to be set explicitly here. Leave it on adaptive and the request is rejected before it costs anything.

Steam rising from a bowl of soup on a wooden counter

Text-to-video, 5 seconds, $0.70. Same rate as step 2, nothing in common visually.

Step 5. Trigger the mutual exclusion on purpose. Two runs, because one alone proves nothing. First a control: reference-to-video with refers only, no first frame, 5 seconds, $0.70.

Steaming bowl of ramen on a wet counter under a neon sign

A neon-lit rain-streaked ramen counter with pink and cyan signage reflections, the still used as the style reference

The style reference passed in refers.

text
1A tall glass of citrus soda on a wet bar top, rendered in the neon-and-rain colour
2language of the reference: pink and cyan rim light, rain-streaked reflections on the
3stone, steam drifting behind the glass, slow push-in, ambient rain and carbonation fizz
4

Iced drink with lemon slice on a neon-lit bar counter

Control run. refers works: pink and cyan rim light, wet neon bar top, the citrus soda from the prompt. $0.70.

Now the forbidden request. Same endpoint, but send the Step 1 first-frame image and the same refers style reference together, which the docs say cannot coexist. duration: 10, ratio: 16:9.

text
1The soda glass on the wet bar top, rendered in the neon-and-rain colour language
2of the reference: pink and cyan rim light, rain-streaked reflections on the stone,
3steam drifting behind the glass, slow push-in, ambient rain and carbonation fizz
4

Glass of iced cola on a wet bar with neon reflections

HTTP 200, completed, billed the full $1.40. The neon reference came through, but the locked first frame did not: the pale citrus soda in a tall highball became a dark drink in a short glass. One of the two inputs was silently discarded and the response said nothing about it.

That is the whole point of the experiment. The documentation states the rule; the API does not enforce it. Put the check in your own client, because nothing upstream will stop you paying for an instruction that gets thrown away. No playground screenshot for this step: the reference-to-video page defaults its aspect ratio to adaptive and would not accept an override in our runs, so it went through the API directly.

Variations That Change Your MiniMax H3 API Pricing

One 15-second clip or three 5-second clips. Identical cost at $1.95 official, $2.10 on Atlas. Very different retry exposure. A rejected 15-second render burns $2.10. A rejected 5-second render burns $0.70. If your prompt is unproven, shoot short and cut.

Text-to-video or image-to-video. Text-to-video skips the base-frame cost. Image-to-video adds about $0.12 and buys a locked composition, which is the single biggest lever on retry rate. One avoided 10-second retry pays for eleven base frames.

Reference-to-video for style or character consistency. Cost is output seconds plus $0.04 for every reference image past the fifth. Cheap, as long as you remember it is mutually exclusive with first-frame and last-frame inputs, and that the API will not remind you.

Self-hosting. H3 weights are public, but local inference produces 768p and needs a separate H3-Regenerate-2K pass for 2K. Weigh the GPU and storage bill against $0.13 per second before you commit an engineer to it.

The Real MiniMax H3 API Pricing Invoice for This Spot

Every call, including the two we broke on purpose and the ones behind the screenshots. Charges are the values the API reported after the price field settled.

T5. Line-by-line invoice, 3 August 2026, Atlas Cloud rates:

#EndpointSettingsStatusCharged
1GPT Image 2, text-to-image1280x720, quality high, x3 framescompleted$0.35583
2GPT Image 2, text-to-image1024x1792 vertical variant, quality highcompleted$0.15689
3GPT Image 2, text-to-image2048x1152, quality high (playground run)completed$0.17450
4H3 image-to-videoduration: 5, 2K (test take)completed$0.70
5H3 image-to-videoduration: 10, 2K (delivered spot)completed$1.40
6H3 text-to-videoduration: 5, 2K, 16:9completed$0.70
7H3 reference-to-videoduration: 5, 2K, refers onlycompleted$0.70
8H3 reference-to-videoduration: 10, 2K, image + refers, first frame droppedcompleted$1.40
9H3 image-to-videoduration omitted, 2K (came back as 5s)completed$0.70
10H3 image-to-videoimage_url returning 404, duration: 5, 2Kcompleted$0.70
11H3 image-to-videoinvalid ratio valuerejected at submit$0.00
12H3 image-to-videodefault 8s, 2K (playground run)completed$1.12
13H3 text-to-videodefault 8s, 2K (playground run)completed$1.12
Total$9.23

One delivered 10-second 2K spot with native audio, a test take, a second creative direction, a working style transfer, and four experiments, for $9.23. Strip the experiments out and the marginal cost of one more delivered 10-second spot is $1.40 plus one base frame, about $1.52.

At 30 spots a month that is roughly $46 in render cost. Concurrency is what decides whether those 30 spots take two waves or fifteen.

Licensing, if you self-host. The MiniMax Community License took effect on 2 August 2026 and it is stricter than most people expect. Excluded Territories cover the EU, the UK, South Korea and the United States. Use of the Works and their Outputs outside the permitted regions is prohibited, organisations above $20M in annual revenue need written authorisation, the UI must display "MiniMax H3", outputs cannot be used to train other models, and Hong Kong law governs. The repo's own license Q&A frames the territory list as not yet open rather than permanently closed. Critically, all of this binds the downloadable weights only. Calling a hosted API is outside its scope.

MiniMax H3 API Pricing: Frequently Asked Questions

How much does 10 seconds of MiniMax H3 video cost?

$1.30 at MiniMax's official 2K rate of $0.13 per second, or $0.80 at the 768P rate of $0.08 per second. On Atlas Cloud, which exposes 2K only at $0.14 per second, the same clip billed us exactly $1.40. Reference images past the fifth add $0.04 each.

Is MiniMax H3 API pricing actually cheaper on the 768P tier?

At $0.08 per second, yes, if you ship 768P. But upgrading a 768P render to 2K costs $0.05 per second, and $0.08 plus $0.05 is exactly the $0.13 direct 2K rate. Drafting cheap only saves money on takes you discard, at $0.05 per discarded second.

How many MiniMax H3 requests can I run in parallel?

Two concurrent tasks on the free tier, fifteen on paid. That is CONN, maximum concurrent tasks, not requests per minute. Thirty clips means fifteen sequential waves on free versus two on paid. Real behaviour varies day to day, so measure your own account with two overlapping requests.

Do failed requests count toward MiniMax H3 API pricing?

Requests rejected at submit time are free: an invalid parameter returns HTTP 400 with no price on the record. But a request that is merely wrong still bills in full. Ours pointed at a 404 image URL, reported completed, rendered anyway, and charged $0.70.

Can I send a first-frame image and reference materials together?

The docs say the two are mutually exclusive. In practice the API accepted both, returned 200, billed the full $1.40, and silently ignored the first frame while honouring the reference. Enforce the rule in your own client, because the API will not enforce it for you.

Is self-hosting cheaper than MiniMax H3 API pricing?

Only at serious volume, and with caveats. Local inference produces 768p and needs a separate regeneration pass for 2K, and the Community License Excluded Territories include the US, EU, UK and South Korea. The hosted API is not subject to that clause.

Latest Models

One API for All Media AI.

Explore all models