
MiniMax H3 Text-to-Video API by MiniMax
MiniMax H3 text-to-video: generate a cinematic video from a text prompt. Supports 2K, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.
MiniMax H3 Text-to-Video
MiniMax H3 Text-to-Video is a state-of-the-art AI video generation model that creates cinematic, high-fidelity videos directly from a text prompt. With crisp detail up to 2K, smooth natural motion, and flexible aspect ratios, it turns a single description into a polished clip.
Why Choose This?
-
High resolution output Generate videos in 2K quality.
-
Cinematic motion Fluid camera work and lifelike movement from a plain-text description.
-
Flexible aspect ratios 16:9, 9:16, 1:1, or adaptive to fit any platform.
-
Selectable duration Produce 5s or 10s clips.
Parameters
| Parameter | Required | Description |
|---|---|---|
| prompt | Yes | Text description of the scene, subject, and action |
| resolution | No | Output quality: 2K (default) |
| duration | No | Video length in seconds: 8 (default), 5-10 |
| ratio | No | Aspect ratio: adaptive (default), 9:16, 1:1, 4:3 |
How to Use
- Write your prompt — describe the scene, subject, lighting, and motion in detail.
- Set resolution — higher for quality, lower for faster generation.
- Choose an aspect ratio — match your target platform, or use adaptive.
- Adjust duration — pick 5s or 10s.
- Run — submit and download your video.
Pricing
Billed per second of generated video, by resolution:
| Resolution | Cost per second |
|---|---|
| 2K | $0.14 |
| 768p | $0.10 |
Best Use Cases
- Social Media Content — Short-form clips for TikTok, Reels, and Stories.
- Concept Visualization — Bring ideas to life without filming.
- Marketing Videos — Produce promotional content from text.
- Storytelling — Create narrative scenes for creative projects.
Pro Tips
- Be specific about camera angle, lighting, mood, and motion.
- Use "adaptive" ratio to let the model choose the best framing for the scene.
- Higher resolution (2K) suits hero shots; 768P is great for quick iteration.
- Describe environmental effects (wind, smoke, golden-hour light) for richer results.
Notes
- A prompt is required.
- Supported durations are 5s and 10s.
- Generation is asynchronous — submit, then poll for the finished video.
Related Models
- MiniMax H3 Image-to-Video — Animate a first-frame image (optionally with a last frame).
- MiniMax H3 Reference-to-Video — Keep a subject consistent from a reference image.


















