Seedance 2.0 Mini & Fast API at Lowest Prices Worldwide — up to 68% off official pricing
Home
Explore
MiniMax
MiniMax H3
minimax/h3-developer/reference-to-video
Atlas Cloud GeneratorUnlock your potential as a director.Go Create
MiniMax H3-Developer Reference-to-Video
image-to-video
DEV

MiniMax H3 Developer Reference-to-Video API by MiniMax

minimax/h3-developer/reference-to-video
Reference-to-video

MiniMax H3-Developer self-hosted reference-to-video: generate a video that keeps the subject from one or more reference images/videos, driven by a text prompt, with generated audio. Supports 480P768P/2K.

MiniMax H3 Developer Reference-to-Video is developed by MiniMax. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.

MiniMax H3-Developer Reference-to-Video

MiniMax H3-Developer Reference-to-Video keeps a subject consistent — with sound. Provide one or more reference materials (images and/or videos of a person, animal, or object) plus a prompt. Generate native 480P or 768P video, or select an ESR tier to enhance the output to a 1080P, 1440P, or 4K short edge while preserving the referenced subject's identity, complete with a synchronized soundtrack generated from the described scene.

Why Choose This?

  • Subject-consistent generation Keep the referenced subject's identity across the whole shot.

  • Multi-modal references Drive generation with images and/or videos in a single request.

  • Native audio A matching soundtrack is generated alongside the video — no separate audio input needed.

  • Enhanced resolution output Generate native 480P or 768P video, or use ESR for a 1080P, 1440P, or 4K short edge.

  • Selectable duration Produce 5–15s clips.

Parameters

ParameterRequiredDescription
promptNoText description of the desired video while keeping the referenced subject's identity
refersYesArray of reference materials. Each item is an object { "url": string, "type": "image" | "video" | "audio" }. type is inferred from the URL extension when omitted. At least one image or video is required; audio alone is rejected
resolutionNoVideo resolution. Available options: 768P (default), 480P, 1080p-esr, 1440p-esr, 4k-esr. ESR tiers generate a native 768P source and enhance the output to the requested short edge.
durationNoDuration of the generated video in seconds. Integer between 5 and 15 (default 5)
prompt_expansionNoEnable AI prompt expansion via H3 Context-IR. When true, your prompt is first expanded into a rich, structured description (shot, soundscape, music) before generation, and an additional per-token Context-IR fee applies. Default false: the prompt is used as-is with no extra charge.

How to Use

  1. Add your reference materials — pass one or more refers entries (image and/or video URLs).
  2. Write your prompt — describe the desired scene, action, and sound while keeping the subject.
  3. Set resolution and duration — balance quality against generation speed.
  4. (Optional) Tune seed and steps — fix the seed for reproducibility, raise steps for quality.
  5. Run — submit and download your video.

Best Use Cases

  • Character Consistency — Keep the same character across multiple shots.
  • Product Showcases — Feature a specific product while varying the scene.
  • Brand & IP — Reuse a mascot or signature subject in new content.
  • Storytelling — Build a narrative around a consistent subject, with matching sound.

Pro Tips

  • Use clear, high-quality reference images or videos of the subject.
  • Provide at least one image or video reference; audio-only requests are rejected.
  • Describe the scene's audio (traffic, music, crowd) to shape the generated soundtrack.
  • Keep the prompt focused on the scene and action; the subject's identity comes from the references.

Notes

  • At least one image or video reference is required.
  • The prompt is optional but strongly recommended.
  • ESR tiers add an enhancement stage after native generation and may take longer than native resolutions.
  • Ensure all reference URLs are publicly accessible.
  • Generation is asynchronous — submit, then poll for the finished video.

Explore Similar Models

One API for All Media AI.

Explore all models