
MiniMax H3 Developer Image-to-Video API by MiniMax
MiniMax H3-Developer self-hosted image-to-video: animate a first-frame image (optionally a last frame) driven by a text prompt, with generated audio. Supports 480P/768P/2K.
MiniMax H3 Developer Image-to-Video is developed by MiniMax. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.
MiniMax H3-Developer Image-to-Video
MiniMax H3-Developer Image-to-Video brings a static image to life — with sound. Provide a first-frame image — and, optionally, a last frame — plus a prompt describing the motion. Generate native 480P or 768P video, or select an ESR tier to enhance the output to a 1080P, 1440P, or 4K short edge, complete with a synchronized soundtrack generated from the described scene.
Why Choose This?
-
Image-driven generation Animate any image with natural, controllable motion.
-
Native audio A matching soundtrack (ambient sound, effects, and music) is generated alongside the video — no separate audio input needed.
-
First & last frame control Set the opening frame, and optionally pin the closing frame for a precise transition.
-
Enhanced resolution output Generate native 480P or 768P video, or use ESR for a 1080P, 1440P, or 4K short edge.
-
Selectable duration Produce 5–15s clips.
Parameters
| Parameter | Required | Description |
|---|---|---|
| prompt | No | Text description of the desired motion and action |
| image | Yes | First frame of the video (public URL) |
| end_image | No | Last frame of the video (public URL) |
| resolution | No | Video resolution. Available options: 768P (default), 480P, 1080p-esr, 1440p-esr, 4k-esr. ESR tiers generate a native 768P source and enhance the output to the requested short edge. |
| duration | No | Duration of the generated video in seconds. Integer between 5 and 15 (default 5) |
| prompt_expansion | No | Enable AI prompt expansion via H3 Context-IR. When true, your prompt is first expanded into a rich, structured description (shot, soundscape, music) before generation, and an additional per-token Context-IR fee applies. Default false: the prompt is used as-is with no extra charge. |
How to Use
- Upload your first-frame image — the video will start from this image.
- (Optional) Upload a last-frame image — the video will end on this image.
- Write your prompt — describe the motion, camera movement, action, and the sound of the scene.
- Set resolution and duration — balance quality against generation speed.
- (Optional) Tune seed and steps — fix the seed for reproducibility, raise steps for quality.
- Run — submit and download your video.
Best Use Cases
- Photo Animation — Bring portraits, landscapes, and product images to life.
- Start/End Transitions — Morph smoothly from one image to another.
- Marketing & Ads — Turn product photos into dynamic promotional videos.
- Storytelling — Animate illustrations and artwork for narratives, with matching sound.
Pro Tips
- Be specific about movement direction, speed, and camera angles.
- Describe the scene's audio (traffic, music, crowd) to shape the generated soundtrack.
- When using a last frame, keep the two images' aspect ratios close for a clean transition.
- Use high-quality source images for better video results.
Notes
- The first-frame image is required; the prompt is optional but strongly recommended.
- The last-frame image is optional; when provided, the clip ends on it.
- ESR tiers add an enhancement stage after native generation and may take longer than native resolutions.
- Ensure image URLs are publicly accessible.
- Generation is asynchronous — submit, then poll for the finished video.
Related Models
- MiniMax H3-Sol Text-to-Video — Generate video directly from a text prompt.
- MiniMax H3-Sol Reference-to-Video — Keep a subject consistent from reference materials.


















