
MiniMax H3 Max text-to-video: generate a cinematic video from a text prompt. Supports 480P、768P, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.

MiniMax H3 Max image-to-video: animate a first-frame image (optionally with a last frame) driven by a text prompt. Supports 480P、768P, 5-15s.

MiniMax H3 Fast text-to-video: generate a cinematic video from a text prompt. Supports 480P, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.

MiniMax H3 Fast image-to-video: animate a first-frame image (optionally with a last frame) driven by a text prompt. Supports 480P, 5-15s.

MiniMax H3 Fast reference-to-video: generate a video that keeps the subject from a reference image, driven by a text prompt. Supports 480P, 5-15s.

A natively multimodal Google DeepMind model that continues an existing clip with a seamlessly matched 3-to-10-second extension, chainable to grow a single shot up to a total of 40 seconds of coherent video with native audio.

A natively multimodal Google DeepMind model that applies a text-instructed edit to an existing video - adding, removing, replacing, or restyling elements with native audio - while preserving everything the prompt does not mention.

A natively multimodal Google DeepMind model that generates cinematic, natively sound-enabled videos from a text prompt plus up to 10 reference images and 3 reference video clips, keeping a character, product, or art direction consistent across generations.

