Seedance 2.5 Now Live — First on Atlas Cloud
Hero background 1Hero background 2

MiniMax H3 API for Video Generation and Editing

The MiniMax H3 API opens MiniMax's general purpose multimodal video model, which reads text, images, video and audio as one context instead of one task at a time. Clips run 5 to 15 seconds at 24 FPS across aspect ratios from 21:9 to 9:16, and one prompt can swap characters, replace backgrounds, rewrite dialogue or clone a voice from a reference clip. Atlas Cloud serves it all through one OpenAI-compatible endpoint. Start building today.

MiniMax H3 is developed by MiniMax. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.

Explore the Leading MiniMax H3

Atlas Cloud provides you with the latest industry-leading creative models.

From Text to Nine References: MiniMax H3 API Modalities

Every MiniMax H3 API endpoint reads text, images, video, and audio differently, so scan the rows below and match a modality to what you are building.

ModalityDescription
MiniMax H3 T2V API (Text to Video)Write a prompt of up to 7,000 characters and the model returns a 5 to 15 second clip at 24 FPS in 1440p, with native stereo sound generated in the same pass. Aspect ratio is set per request across 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, so one call covers cinematic trailers and vertical social cuts alike.
MiniMax H3 I2V API (Image to Video)One image sets the opening frame and two lock both the first and the last, with H3 filling in the motion between them at the aspect ratio of the input image. Sources from 256 to 5760 pixels per side are accepted, which makes key art, product stills, and storyboard frames straightforward to put into motion.
MiniMax H3 Omni Reference API (Reference to Video)Need several sources working together? Up to nine images, three video clips, and three audio tracks fit in a single request under a 12 file cap, all read as one multimodal context. Carry a character, a camera move, or a voice timbre across shots, or edit an existing clip by swapping subjects, backgrounds, and spoken lines.

Sight, Sound and Editing in One MiniMax H3 API Call

Every clip the MiniMax H3 API returns runs up to fifteen seconds at 1440p with native stereo sound, built from any mix of text, image, video and voice references, and the same call can also edit characters, scenes and dialogue inside footage you already have.

Twelve Reference Files, One MiniMax H3 API Call

One MiniMax H3 API request accepts up to nine reference images, three video clips and three audio tracks, capped at twelve files. Rather than reading them as separate slots, the model treats text, picture and sound as one context, pulling a character from a photo, a camera move from a clip and a mood from a track into the same scene. Teams with existing material get a direct route to a finished shot.

Native Stereo Sound in Every Clip

Every result ships with sound. Dialogue, ambience and music are produced in native stereo during the same pass that renders the picture at 24 frames per second, so nothing has to be scored or dubbed afterwards. Because timing is decided while the shot is generated, footsteps, speech and cuts stay locked to the action. Short ads and social spots come out ready to publish.

Prompt-Level Edits Through the MiniMax H3 API

Need a cat swapped for a dog, a green screen replaced, or one line of dialogue rewritten? The MiniMax H3 API applies edits like these to video you supply, covering characters, objects, backgrounds, lighting and effects, while parts you did not mention stay close to the original. Prompts run to 7000 characters, so a dozen changes can be stacked into one pass. That makes iterating on an approved cut practical.

Voice Timbre Transfer and Dialogue Swaps

Send a voice sample alongside your images or footage and the generated character speaks in that timbre. Up to three audio references are allowed per request, each between two and fifteen seconds, and audio must always accompany a visual input rather than arrive on its own. Existing dialogue can also be replaced and the performance adjusted to match. Series work keeps one recognizable voice across every episode.

1440p and Six Aspect Ratios in the MiniMax H3 API

Clips run from five to fifteen seconds, and the 1440p mode puts 1440 pixels on the short side between 16:9 and 9:16, or roughly 3.7 megapixels at wider ratios such as 2976 by 1248 for 21:9. Six ratios are selectable, from cinematic 21:9 to vertical 9:16, and the MiniMax H3 API can also choose one for you. That range covers a trailer, a product loop and a vertical drama without switching models.

Production Formats In, No Conversion Step

If your source files come straight off a camera or an editing timeline, they go in as they are: H.264 and H.265 video, JPG, PNG, WEBP, HEIC and HEIF stills, plus WAV and MP3 audio. Per-file limits sit at 50MB for video, 30MB for images and 15MB for audio, while passing assets by URL keeps requests within the 64MB body limit. On Atlas Cloud the whole set runs through one OpenAI-compatible key with pay-as-you-go billing.

Same Prompt, Three Engines: MiniMax H3 API Head to Head

Every clip in this set comes from one identical prompt sent to the MiniMax H3 API and two other video models hosted on Atlas Cloud, so motion, sound, and instruction fidelity can be compared without changing a single word.

Prompt

15 seconds, 16:9 landscape short video. Live-action footage of a late-night self-service laundromat, blended with hand-drawn glowing animation into a mixed-media image. A small self-service laundromat, its fluorescent lights faintly flickering; inside are running washing machines, plastic laundry baskets, and an old bench, with a single sock lying on the floor. The whole space is quiet, carrying a faint, nostalgic mood. It has the texture of one-handed handheld phone footage, with noticeable camera shake; the white fluorescent light causes the exposure to fluctuate between bright and dim; glass surfaces carry ambient reflections; there's a focus lag when the lens moves close to objects. The image should not be as polished and orderly as a commercial ad — the overall feel should be like a genuine documentary snapshot, as if you stumbled in by chance late at night and grabbed the shot while chasing some strange, dreamlike vision.

Generated with MiniMax H3 on Atlas Cloud

Generated with Seedance 2.0 on Atlas Cloud

Generated with Wan-2.7 on Atlas Cloud

Prompt

First-person perspective · eye-level height · handheld gaming camera Scene: The shot simulates a player operating a modern-warfare FPS game, both hands holding an assault rifle while slowly advancing along the outer perimeter of a military base. The player moves forward along a road beside cover, the crosshair sweeping across the passage ahead; after a brief pause, they fire a few rounds toward a distant objective, then continue pushing forward — like the live gameplay footage of an ordinary player. Lighting: The cool-toned natural light of a modern military base interweaves with smoke and muzzle fire. The image is realistic and crisp, with the metallic weapon and the battlefield dust and haze carrying a AAA-game quality. Camera work: The camera has a slight handheld sway as the player moves — first advancing slowly, then making small left-and-right sweeps to observe, with a subtle recoil shake when firing, before finally continuing to push steadily forward.

Generated with MiniMax H3 on Atlas Cloud

Generated with Seedance 2.0 on Atlas Cloud

Generated with Wan-2.7 on Atlas Cloud

Where the MiniMax H3 API Fits in Production

From brand films and vertical drama to product cuts, game visuals and precise edits of existing footage, the MiniMax H3 API covers each scenario through one multimodal request that returns video with native stereo audio.

Cinematic Brand Films on the MiniMax H3 API

Feed storyboard frames and a shot list into one call for 1440p trailers, TVC spots and fashion campaigns at 24 FPS. Native stereo audio ships with every result, so brand teams screen finished cuts.

Vertical Short Drama Production

Vertical 9:16 output covers scripted drama scenes, from costume mystery to family confrontation, with dialogue voiced in the same pass. Studios building short drama libraries get 15 second hooks without booking actors or stages.

Multi-Reference Shot Assembly

If a shot needs a specific face and motion, up to nine images, three videos and three audio clips can guide one call. Character identity, camera work and vocal timbre hold steady across episodes.

Editing Existing Footage with the MiniMax H3 API

Need a change after the fact? Existing footage can be edited by prompt: swap a subject, replace a background, adjust lighting or rewrite a spoken line without reshooting a frame.

Product and E-commerce Marketing

Product photos become motion: one reference image turns into a 360 degree showcase, a feature explainer or a paid social cut. Aspect ratios from 21:9 to 9:16 let one asset set feed every placement.

MiniMax H3 API for Game and Anime Visuals

Stylized output holds up for game CG, character PV, anime openings and interface demos where menus, HUD elements and text overlays must stay readable. Art teams use it for concept validation before production.

MiniMax H3 API Compared With Other Multimodal Video Models

Line the MiniMax H3 API up against the other video models hosted on Atlas Cloud and see how input modalities, reference limits, length, resolution, and audio output actually differ before you commit to an endpoint.

ModelInput ModalitiesMax Reference FilesOutput DurationMax ResolutionNative Audio
MiniMax H3Text, image, video, audio9 images, 3 videos, 3 audio clips, 12 files total5s to 15s1440p at 24 FPS√ Native stereo audio on every output
Seedance 2.0 Reference-to-VideoText, image, video, audio9 images, 3 videos, 3 audio clipsUp to 15s720p√ Stereo dialogue, effects, and music generated in one pass
Veo3.1 Reference-to-videoText and image3 reference images4s, 6s, or 8s4K at 8s length only√ Dialogue, ambience, and sound effects aligned to the timeline
Wan-2.7 Reference-to-videoText, image, video, audio5 images or video clips, plus one voice clip2s to 15s1080p√ Music and effects generated, or drive lip sync with your own audio
Kling v3.0 Pro Image-to-VideoText and image-Up to 15s1080p√ Multilingual dialogue with lip sync in five languages

How to Use MiniMax H3 on Atlas Cloud

Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.

Create an Atlas Cloud Account

Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.

Why Use MiniMax H3 on Atlas Cloud

Combining the advanced MiniMax H3 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.

Performance & flexibility

Low Latency:
GPU-optimized inference for real-time reasoning.

Unified API:
Run MiniMax H3, GPT, Gemini, and DeepSeek with one integration.

Transparent Pricing:
Predictable per-token billing with serverless options.

Enterprise & Scale

Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.

Reliability:
99.99% uptime, RBAC, and compliance-ready logging.

Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.

MiniMax H3 API: Answers for Developers

The MiniMax H3 API gives developers programmatic access to MiniMax H3, an open general purpose multimodal video model that treats text, images, video, and audio as one shared context. Rather than splitting generation, editing, and reference into separate task models, H3 reads the full input set and returns a finished clip with sound. On Atlas Cloud it runs behind a single API key with pay-as-you-go pricing.

Brand films, trailers, vertical short drama, product and ecommerce spots, game and UI motion demos, and stylized animation all sit inside its range. Because the model handles on screen text, subtitles, and brand assets, teams also use it for concept validation, storyboard previews, and visual pitches before committing production budget.

Create an Atlas Cloud account, generate an API key, then send a prompt plus any reference files to the video generation endpoint and poll for the finished result. Passing media as hosted URLs is recommended over inline uploads, since the request body is capped at 64MB. Start building today.

Billing is pay-as-you-go, so you pay per call instead of buying a subscription or a seat license. Cost tracks what you actually render, which means resolution and clip length drive the total for a batch. Check the model page for the current per generation rate before planning large volume runs.

Yes. Every H3 result is delivered with sound in native stereo, so dialogue, effects, and ambience arrive in the same pass as the picture. Supply a reference audio clip and the model can carry that timbre onto a character, which removes a separate voice synthesis step from the pipeline.

Clips run from 5 to 15 seconds at 24 FPS. In 1440p mode the short side renders at 1440 pixels for ratios between 16:9 and 9:16, and outside that band the frame holds roughly 3.7M pixels in total, for example 2976 by 1248 at 21:9. Text to video and omni reference requests accept 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, while first and last frame requests inherit the aspect ratio of the input image.

Up to nine images, three video segments, and three audio clips can travel with one prompt, capped at twelve files in total. Video and audio each stay within 15 seconds combined, and audio must accompany an image or a video rather than arrive on its own. Prompts reach 7000 characters, which leaves room for shot by shot direction covering camera, performance, and sound.

Accepted inputs include H.264 and H.265 video, JPG, JPEG, PNG, WEBP, HEIC, and HEIF images, and WAV or MP3 audio, with AAC or MP3 for the audio track inside a video file. Per file ceilings are 50MB for video, 30MB for images, and 15MB for audio. Because the request body is limited to 64MB, hosted URLs remain the safer route for heavy assets.

Editing is one of its core modes. Send a source clip with instructions and the MiniMax H3 API can swap characters or objects, replace backgrounds and lighting, layer in visual effects, and rewrite dialogue while keeping untouched regions stable. Compound instructions are handled in one request, so several changes land together instead of across repeated round trips.

Most video models take one prompt plus one image and return a silent clip. H3 instead consumes a mixed set of references, reads character, motion, camera, and sound intent across all of them, then returns a clip with native audio. If your pipeline currently stitches a video model, a voice model, and an editor together, H3 collapses those stages into a single call.

Explore More Families

Seedance 2.5

Seedance 2.5 API is now available on Atlas Cloud! It gives developers ByteDance's newest video model. It generates up to 30 seconds of native video in a single pass from text, a single image, or as many as 50 multimodal references, with synchronized audio and in-frame multilingual text. On Atlas Cloud you reach it through one key, with subject consistency and improved physics keeping long shots coherent.

View Family

MiniMax H3

The MiniMax H3 API opens MiniMax's general purpose multimodal video model, which reads text, images, video and audio as one context instead of one task at a time. Clips run 5 to 15 seconds at 24 FPS across aspect ratios from 21:9 to 9:16, and one prompt can swap characters, replace backgrounds, rewrite dialogue or clone a voice from a reference clip. Atlas Cloud serves it all through one OpenAI-compatible endpoint. Start building today.

View Family

Seedream 5.0 Pro

Seedream 5.0 Pro API gives developers ByteDance's controllable image editing model on Atlas Cloud. It places edits precisely with anchors and coordinates, separates images into editable layers, fuses multiple references, and matches exact colors and materials, with multilingual text at 2K and 3K. On Atlas Cloud you reach it through one key!

View Family

Seedance 2.0

The Seedance 2.0 API gives you production access to ByteDance's multimodal video model — quad-modal inputs (text, image, video, audio) and an industry-leading "Universal Reference" system that locks composition, camera movement, and character actions across shots. Integrate director-level control with one API call, a flat $0.09/s, instant key, and no waitlist — backed by enterprise-grade uptime and compliance. Seedance 2.0 Native 4K is now live!

View Family

GPT Image 2

The GPT Image 2 API gives developers access to OpenAI's latest image model, the successor to GPT Image 1.5. It generates and edits images with accurate text rendering across Latin and CJK scripts, plus strong composition for posters, mockups, and infographics. On Atlas Cloud you reach it through one unified API alongside 300+ models, with free credits, 99.99% uptime, and no OpenAI organization verification required.

View Family

Gemini Omni Flash

The Gemini Omni API brings Google DeepMind's multimodal video generation and editing model, introduced at Google I/O 2026, to your stack. Gemini Omni fuses Gemini's reasoning engine with generative media, accepting any mix of text, images, video, and audio to produce consistent, knowledge-grounded output. Refine results through natural conversation, swapping objects, rewriting scenes, and shifting styles while physics, characters, and continuity stay intact. Atlas Cloud serves the full Gemini Omni Flash lineup, text-to-video, image-to-video with up to 7 reference images, and reference-to-video, through one unified API with transparent per-second pricing from $0.112 and no subscription. Start building today.

View Family

Grok Imagine

The Grok Imagine API covers xAI's image, video, and speech models, from Image 2.0 to Video 1.5 and xAI TTS v1. Render 1K or 2K stills across 14 aspect ratios, push a scene to 15 seconds of 1080p motion, steer shots with up to 7 reference images, or narrate them in 20 languages. Atlas Cloud runs every mode on one endpoint, priced pay-as-you-go from $0.02 per image and $0.05 per second. Start building today.

View Family

Google

Google's most powerful creative models are all available on Atlas Cloud. Veo 3.1 delivers cinematic video generation, Nano Banana 2 powers high-fidelity image creation, and Gemini brings multimodal intelligence to every workflow. Access the full Google model suite through one API key with Day-0 availability and pay-as-you-go pricing.

View Family

Seedance 2.0 Mini

The Seedance 2.0 Mini API is the lightest, lowest-cost tier of ByteDance's Seedance video line, built for teams where throughput and unit cost matter more than maximum polish. Use it for batch generation, rapid prototyping, and draft passes, all through one OpenAI-compatible key on Atlas Cloud.

View Family

ByteDance

From cinematic video generation to high-fidelity image creation, ByteDance's most powerful models are live on Atlas Cloud. Run Seedance and Seedream at scale with the lowest inference pricing and zero infrastructure overhead.

View Family

Alibaba

Atlas Cloud brings together Alibaba's full model lineup under one API: Qwen for language and image tasks, Wan for video generation up to 1080p. Access every model pay-as-you-go with no subscriptions. The Alibaba API is available via a single base URL using your existing OpenAI-compatible client.

View Family

OpenAI

Atlas Cloud gives you access to the full OpenAI API lineup, from GPT Image 2 for image generation to Sora 2 for video. Every model is available pay-as-you-go with no monthly commitment. Plug in with a single base URL swap using the OpenAI-compatible API.

View Family

One API for All Media AI.

Explore all models