

The MiniMax H3 API opens MiniMax's general purpose multimodal video model, which reads text, images, video and audio as one context instead of one task at a time. Clips run 5 to 15 seconds at 24 FPS across aspect ratios from 21:9 to 9:16, and one prompt can swap characters, replace backgrounds, rewrite dialogue or clone a voice from a reference clip. Atlas Cloud serves it all through one OpenAI-compatible endpoint. Start building today.
MiniMax H3 is developed by MiniMax. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.
Atlas Cloud provides you with the latest industry-leading creative models.
Every MiniMax H3 API endpoint reads text, images, video, and audio differently, so scan the rows below and match a modality to what you are building.
| Modality | Description |
|---|---|
| MiniMax H3 T2V API (Text to Video) | Write a prompt of up to 7,000 characters and the model returns a 5 to 15 second clip at 24 FPS in 1440p, with native stereo sound generated in the same pass. Aspect ratio is set per request across 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, so one call covers cinematic trailers and vertical social cuts alike. |
| MiniMax H3 I2V API (Image to Video) | One image sets the opening frame and two lock both the first and the last, with H3 filling in the motion between them at the aspect ratio of the input image. Sources from 256 to 5760 pixels per side are accepted, which makes key art, product stills, and storyboard frames straightforward to put into motion. |
| MiniMax H3 Omni Reference API (Reference to Video) | Need several sources working together? Up to nine images, three video clips, and three audio tracks fit in a single request under a 12 file cap, all read as one multimodal context. Carry a character, a camera move, or a voice timbre across shots, or edit an existing clip by swapping subjects, backgrounds, and spoken lines. |
Every clip the MiniMax H3 API returns runs up to fifteen seconds at 1440p with native stereo sound, built from any mix of text, image, video and voice references, and the same call can also edit characters, scenes and dialogue inside footage you already have.
One MiniMax H3 API request accepts up to nine reference images, three video clips and three audio tracks, capped at twelve files. Rather than reading them as separate slots, the model treats text, picture and sound as one context, pulling a character from a photo, a camera move from a clip and a mood from a track into the same scene. Teams with existing material get a direct route to a finished shot.
Every result ships with sound. Dialogue, ambience and music are produced in native stereo during the same pass that renders the picture at 24 frames per second, so nothing has to be scored or dubbed afterwards. Because timing is decided while the shot is generated, footsteps, speech and cuts stay locked to the action. Short ads and social spots come out ready to publish.
Need a cat swapped for a dog, a green screen replaced, or one line of dialogue rewritten? The MiniMax H3 API applies edits like these to video you supply, covering characters, objects, backgrounds, lighting and effects, while parts you did not mention stay close to the original. Prompts run to 7000 characters, so a dozen changes can be stacked into one pass. That makes iterating on an approved cut practical.
Send a voice sample alongside your images or footage and the generated character speaks in that timbre. Up to three audio references are allowed per request, each between two and fifteen seconds, and audio must always accompany a visual input rather than arrive on its own. Existing dialogue can also be replaced and the performance adjusted to match. Series work keeps one recognizable voice across every episode.
Clips run from five to fifteen seconds, and the 1440p mode puts 1440 pixels on the short side between 16:9 and 9:16, or roughly 3.7 megapixels at wider ratios such as 2976 by 1248 for 21:9. Six ratios are selectable, from cinematic 21:9 to vertical 9:16, and the MiniMax H3 API can also choose one for you. That range covers a trailer, a product loop and a vertical drama without switching models.
If your source files come straight off a camera or an editing timeline, they go in as they are: H.264 and H.265 video, JPG, PNG, WEBP, HEIC and HEIF stills, plus WAV and MP3 audio. Per-file limits sit at 50MB for video, 30MB for images and 15MB for audio, while passing assets by URL keeps requests within the 64MB body limit. On Atlas Cloud the whole set runs through one OpenAI-compatible key with pay-as-you-go billing.
Every clip in this set comes from one identical prompt sent to the MiniMax H3 API and two other video models hosted on Atlas Cloud, so motion, sound, and instruction fidelity can be compared without changing a single word.
15 seconds, 16:9 landscape short video. Live-action footage of a late-night self-service laundromat, blended with hand-drawn glowing animation into a mixed-media image. A small self-service laundromat, its fluorescent lights faintly flickering; inside are running washing machines, plastic laundry baskets, and an old bench, with a single sock lying on the floor. The whole space is quiet, carrying a faint, nostalgic mood. It has the texture of one-handed handheld phone footage, with noticeable camera shake; the white fluorescent light causes the exposure to fluctuate between bright and dim; glass surfaces carry ambient reflections; there's a focus lag when the lens moves close to objects. The image should not be as polished and orderly as a commercial ad — the overall feel should be like a genuine documentary snapshot, as if you stumbled in by chance late at night and grabbed the shot while chasing some strange, dreamlike vision.
Generated with MiniMax H3 on Atlas Cloud
Generated with Seedance 2.0 on Atlas Cloud
Generated with Wan-2.7 on Atlas Cloud
First-person perspective · eye-level height · handheld gaming camera Scene: The shot simulates a player operating a modern-warfare FPS game, both hands holding an assault rifle while slowly advancing along the outer perimeter of a military base. The player moves forward along a road beside cover, the crosshair sweeping across the passage ahead; after a brief pause, they fire a few rounds toward a distant objective, then continue pushing forward — like the live gameplay footage of an ordinary player. Lighting: The cool-toned natural light of a modern military base interweaves with smoke and muzzle fire. The image is realistic and crisp, with the metallic weapon and the battlefield dust and haze carrying a AAA-game quality. Camera work: The camera has a slight handheld sway as the player moves — first advancing slowly, then making small left-and-right sweeps to observe, with a subtle recoil shake when firing, before finally continuing to push steadily forward.
Generated with MiniMax H3 on Atlas Cloud
Generated with Seedance 2.0 on Atlas Cloud
Generated with Wan-2.7 on Atlas Cloud
From brand films and vertical drama to product cuts, game visuals and precise edits of existing footage, the MiniMax H3 API covers each scenario through one multimodal request that returns video with native stereo audio.
Feed storyboard frames and a shot list into one call for 1440p trailers, TVC spots and fashion campaigns at 24 FPS. Native stereo audio ships with every result, so brand teams screen finished cuts.
Vertical 9:16 output covers scripted drama scenes, from costume mystery to family confrontation, with dialogue voiced in the same pass. Studios building short drama libraries get 15 second hooks without booking actors or stages.
If a shot needs a specific face and motion, up to nine images, three videos and three audio clips can guide one call. Character identity, camera work and vocal timbre hold steady across episodes.
Need a change after the fact? Existing footage can be edited by prompt: swap a subject, replace a background, adjust lighting or rewrite a spoken line without reshooting a frame.
Product photos become motion: one reference image turns into a 360 degree showcase, a feature explainer or a paid social cut. Aspect ratios from 21:9 to 9:16 let one asset set feed every placement.
Stylized output holds up for game CG, character PV, anime openings and interface demos where menus, HUD elements and text overlays must stay readable. Art teams use it for concept validation before production.
Line the MiniMax H3 API up against the other video models hosted on Atlas Cloud and see how input modalities, reference limits, length, resolution, and audio output actually differ before you commit to an endpoint.
| Model | Input Modalities | Max Reference Files | Output Duration | Max Resolution | Native Audio |
|---|---|---|---|---|---|
| MiniMax H3 | Text, image, video, audio | 9 images, 3 videos, 3 audio clips, 12 files total | 5s to 15s | 1440p at 24 FPS | √ Native stereo audio on every output |
| Seedance 2.0 Reference-to-Video | Text, image, video, audio | 9 images, 3 videos, 3 audio clips | Up to 15s | 720p | √ Stereo dialogue, effects, and music generated in one pass |
| Veo3.1 Reference-to-video | Text and image | 3 reference images | 4s, 6s, or 8s | 4K at 8s length only | √ Dialogue, ambience, and sound effects aligned to the timeline |
| Wan-2.7 Reference-to-video | Text, image, video, audio | 5 images or video clips, plus one voice clip | 2s to 15s | 1080p | √ Music and effects generated, or drive lip sync with your own audio |
| Kling v3.0 Pro Image-to-Video | Text and image | - | Up to 15s | 1080p | √ Multilingual dialogue with lip sync in five languages |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced MiniMax H3 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run MiniMax H3, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
The MiniMax H3 API gives developers programmatic access to MiniMax H3, an open general purpose multimodal video model that treats text, images, video, and audio as one shared context. Rather than splitting generation, editing, and reference into separate task models, H3 reads the full input set and returns a finished clip with sound. On Atlas Cloud it runs behind a single API key with pay-as-you-go pricing.
Brand films, trailers, vertical short drama, product and ecommerce spots, game and UI motion demos, and stylized animation all sit inside its range. Because the model handles on screen text, subtitles, and brand assets, teams also use it for concept validation, storyboard previews, and visual pitches before committing production budget.
Create an Atlas Cloud account, generate an API key, then send a prompt plus any reference files to the video generation endpoint and poll for the finished result. Passing media as hosted URLs is recommended over inline uploads, since the request body is capped at 64MB. Start building today.
Billing is pay-as-you-go, so you pay per call instead of buying a subscription or a seat license. Cost tracks what you actually render, which means resolution and clip length drive the total for a batch. Check the model page for the current per generation rate before planning large volume runs.
Yes. Every H3 result is delivered with sound in native stereo, so dialogue, effects, and ambience arrive in the same pass as the picture. Supply a reference audio clip and the model can carry that timbre onto a character, which removes a separate voice synthesis step from the pipeline.
Clips run from 5 to 15 seconds at 24 FPS. In 1440p mode the short side renders at 1440 pixels for ratios between 16:9 and 9:16, and outside that band the frame holds roughly 3.7M pixels in total, for example 2976 by 1248 at 21:9. Text to video and omni reference requests accept 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, while first and last frame requests inherit the aspect ratio of the input image.
Up to nine images, three video segments, and three audio clips can travel with one prompt, capped at twelve files in total. Video and audio each stay within 15 seconds combined, and audio must accompany an image or a video rather than arrive on its own. Prompts reach 7000 characters, which leaves room for shot by shot direction covering camera, performance, and sound.
Accepted inputs include H.264 and H.265 video, JPG, JPEG, PNG, WEBP, HEIC, and HEIF images, and WAV or MP3 audio, with AAC or MP3 for the audio track inside a video file. Per file ceilings are 50MB for video, 30MB for images, and 15MB for audio. Because the request body is limited to 64MB, hosted URLs remain the safer route for heavy assets.
Editing is one of its core modes. Send a source clip with instructions and the MiniMax H3 API can swap characters or objects, replace backgrounds and lighting, layer in visual effects, and rewrite dialogue while keeping untouched regions stable. Compound instructions are handled in one request, so several changes land together instead of across repeated round trips.
Most video models take one prompt plus one image and return a silent clip. H3 instead consumes a mixed set of references, reads character, motion, camera, and sound intent across all of them, then returns a clip with native audio. If your pipeline currently stitches a video model, a voice model, and an editor together, H3 collapses those stages into a single call.
Guides, tutorials, and product updates to help you get the most out of Atlas Cloud.