
MiniMax H3 Developer Text-to-Video API by MiniMax
MiniMax H3-Developer self-hosted text-to-video: generate a video (with audio) from a text prompt. Supports 480P/768P/2K, 16:9/9:16/1:1 aspect ratios.
MiniMax H3 Developer Text-to-Video is developed by MiniMax. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.
MiniMax H3-Developer Text-to-Video
MiniMax H3-Developer Text-to-Video turns a single text prompt into a smooth, cinematic clip with sound. Generate native 480P or 768P video, or select an ESR tier to enhance the output to a 1080P, 1440P, or 4K short edge. Describe the scene and the action, and the model generates the video together with a synchronized soundtrack (ambient sound, effects, and music) synthesized from what you described.
Why Choose This?
-
Text-driven generation Create a video from nothing but a prompt.
-
Native audio A matching soundtrack is generated alongside the video — no separate audio input needed.
-
Enhanced resolution output Generate native 480P or 768P video, or use ESR for a 1080P, 1440P, or 4K short edge.
-
Flexible aspect ratios Choose 16:9, 9:16, or 1:1 to match your target platform.
-
Selectable duration Produce 5–15s clips.
Parameters
| Parameter | Required | Description |
|---|---|---|
| prompt | Yes | Text description of the video content, action, and scene audio |
| resolution | No | Video resolution. Available options: 768P (default), 480P, 1080p-esr, 1440p-esr, 4k-esr. ESR tiers generate a native 768P source and enhance the output to the requested short edge. |
| ratio | No | Aspect ratio of the generated video. Available options: 16:9 (default), 9:16, 1:1 |
| duration | No | Duration of the generated video in seconds. Integer between 5 and 15 (default 5) |
| prompt_expansion | No | Enable AI prompt expansion via H3 Context-IR. When true, your prompt is first expanded into a rich, structured description (shot, soundscape, music) before generation, and an additional per-token Context-IR fee applies. Default false: the prompt is used as-is with no extra charge. |
How to Use
- Write your prompt — describe the subject, motion, camera movement, and the sound of the scene.
- Choose an aspect ratio — match your target platform (16:9, 9:16, or 1:1).
- Set resolution and duration — balance quality against generation speed.
- (Optional) Tune seed and steps — fix the seed for reproducibility, raise steps for quality.
- Run — submit and download your video.
Best Use Cases
- Concept & Ideation — Visualize an idea straight from a written description.
- Marketing & Ads — Produce short promotional clips without footage or images.
- Social Content — Generate vertical (9:16) or square (1:1) clips for social feeds.
- Storytelling — Turn a scene description into an animated shot with matching sound.
Pro Tips
- Be specific about the subject, movement, camera angle, and lighting.
- Describe the scene's audio (traffic, music, crowd) to shape the generated soundtrack.
- Keep one clear main action per prompt for the most coherent motion.
- Pick the aspect ratio up front to avoid unwanted cropping later.
Notes
- The prompt is required.
- ESR tiers add an enhancement stage after native generation and may take longer than native resolutions.
- Generation is asynchronous — submit, then poll for the finished video.
Related Models
- MiniMax H3-Sol Image-to-Video — Animate a first-frame image into a video.
- MiniMax H3-Sol Reference-to-Video — Keep a subject consistent from reference materials.


















