MiniMax H3 Developer Now Live — 60% Off, From $0.02 per Second
Seedance 2.5 Image-to-Video
image-to-video

Seedance 2.5 Image-to-Video

Generate videos from a first-frame image (and optional last-frame) with native audio.

Seedance 2.5 Image-to-Video
Seedream v5.0 Pro Edit
Nano Banana 2 Reference to Image
Kling Video O3 4K Image-to-Video
Wan-2.7 Text-to-video
CATEGORY
Discount (155)
Model function
Series
48 of 169 models
New
MiniMax H3 Max Text-to-Video
NEW
text-to-video

MiniMax H3 Max Text-to-Video

MiniMax H3 Max text-to-video: generate a cinematic video from a text prompt. Supports 480P、768P, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.

From$0.05/SEC
$0.048/SEC
-5%
MiniMax H3 Max Image-to-Video
NEW
image-to-video

MiniMax H3 Max Image-to-Video

MiniMax H3 Max image-to-video: animate a first-frame image (optionally with a last frame) driven by a text prompt. Supports 480P、768P, 5-15s.

From$0.05/SEC
$0.048/SEC
-5%
MiniMax H3-Developer Text-to-Video
NEW
text-to-video
DEV

MiniMax H3-Developer Text-to-Video

MiniMax H3-Developer self-hosted text-to-video: generate a video (with audio) from a text prompt. Supports 768P/1080P, 16:9/9:16/1:1 aspect ratios, tunable seed and inference steps.

VIDEO
From$0.05/SEC
$0.02/SEC
-60%
MiniMax H3-Developer Image-to-Video
NEW
image-to-video
DEV

MiniMax H3-Developer Image-to-Video

MiniMax H3-Developer self-hosted image-to-video: animate a first-frame image (optionally a last frame) driven by a text prompt, with generated audio. Supports 768P/1080P.

VIDEO
From$0.05/SEC
$0.02/SEC
-60%
MiniMax H3-Developer Reference-to-Video
NEW
image-to-video
DEV

MiniMax H3-Developer Reference-to-Video

MiniMax H3-Developer self-hosted reference-to-video: generate a video that keeps the subject from one or more reference images/videos, driven by a text prompt, with generated audio. Supports 768P/1080P.

VIDEO
From$0.05/SEC
$0.02/SEC
-60%
Wan-3.0-Prime Text-to-video
NEW
HOT
text-to-video

Wan-3.0-Prime Text-to-video

All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.

From$0.068/SEC
$0.061/SEC
-10%
Wan-3.0-Prime Image-to-video
NEW
HOT
image-to-video

Wan-3.0-Prime Image-to-video

Animate a first frame (optionally with a last frame) into a coherent clip, with native audio and smart-duration up to 30s.

From$0.068/SEC
$0.061/SEC
-10%
Wan-3.0 Text-to-video
NEW
HOT
text-to-video

Wan-3.0 Text-to-video

All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.

From$0.05/SEC
$0.04/SEC
-20%
Wan-3.0 Image-to-video
NEW
HOT
image-to-video

Wan-3.0 Image-to-video

Animate a first frame (optionally with a last frame) into a coherent clip, with native audio and smart-duration up to 30s.

From$0.05/SEC
$0.04/SEC
-20%
Seedance 2.5 Reference-to-Video
NEW
image-to-video

Seedance 2.5 Reference-to-Video

Multimodal video generation from reference images, videos, and audio. Supports video editing and extension.

From$0.167/SEC
$0.134/SEC
1080P -36%
Seedance 2.5 Image-to-Video
NEW
image-to-video

Seedance 2.5 Image-to-Video

Generate videos from a first-frame image (and optional last-frame) with native audio.

From$0.167/SEC
$0.134/SEC
1080P -36%
Seedance 2.5 Text-to-Video
NEW
text-to-video

Seedance 2.5 Text-to-Video

Generate videos from text prompts with native audio and optional web search.

From$0.167/SEC
$0.134/SEC
1080P -36%
MiniMax H3 Text-to-Video
NEW
text-to-video

MiniMax H3 Text-to-Video

MiniMax H3 text-to-video: generate a cinematic video from a text prompt. Supports 2K, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.

From
$0.08/SEC
MiniMax H3 Image-to-Video
NEW
image-to-video

MiniMax H3 Image-to-Video

MiniMax H3 image-to-video: animate a first-frame image (optionally with a last frame) driven by a text prompt. Supports 2K, 5-15s.

From
$0.08/SEC
MiniMax H3 Reference-to-Video
NEW
image-to-video

MiniMax H3 Reference-to-Video

MiniMax H3 reference-to-video: generate a video that keeps the subject from a reference image, driven by a text prompt. Supports 2K, 5-15s.

From
$0.08/SEC
Youchuan V8.2 Image-to-Video
NEW
image-to-video

Youchuan V8.2 Image-to-Video

Youchuan V8.2 animates an input image into four 5-second videos at 480p or 720p.

From
$0.086/SEC
Seedance 2.0 Mini Reference-to-Video
NEW
image-to-video

Seedance 2.0 Mini Reference-to-Video

Lightweight, economical multimodal video generation from reference images, videos, and audio with native audio.

From$0.056/SEC
$0.039/SEC
-30%
Seedance 2.0 Mini Image-to-Video
NEW
image-to-video

Seedance 2.0 Mini Image-to-Video

Lightweight, economical video generation from a first-frame image (and optional last-frame) with native audio.

From$0.056/SEC
$0.039/SEC
-30%
Seedance 2.0 Mini Text-to-Video
NEW
text-to-video

Seedance 2.0 Mini Text-to-Video

Lightweight, economical video generation from text prompts with native audio.

From$0.056/SEC
$0.039/SEC
-30%
HappyHorse-1.1 Text-to-video
NEW
text-to-video

HappyHorse-1.1 Text-to-video

Generates videos from text prompts with HappyHorse 1.1, supporting 480P, 720P, or 1080P output, flexible aspect ratios, and durations from 3 to 15 seconds.

From
$0.07/SEC
HappyHorse-1.1 Image-to-video
NEW
image-to-video

HappyHorse-1.1 Image-to-video

Animates a first-frame image into video with optional prompt guidance, 480P, 720P, or 1080P output, and durations from 3 to 15 seconds.

From
$0.07/SEC
Gemini Omni Flash Image-to-Video
NEW
image-to-video

Gemini Omni Flash Image-to-Video

A natively multimodal Google DeepMind model that animates a still image into a cinematic, sound-enabled video guided by a text prompt while preserving the source subject and composition.

From
$0.13/SEC
Gemini Omni Flash Text-to-Video
NEW
text-to-video

Gemini Omni Flash Text-to-Video

A natively multimodal Google DeepMind model that generates cinematic videos with synchronized native audio from a text prompt alone, grounded in real-world physics for controllable, high-speed video generation.

From
$0.125/SEC
Kling V3.0 Turbo Image-to-Video
NEW
image-to-video
TURBO

Kling V3.0 Turbo Image-to-Video

Kling V3.0 Turbo Image-to-Video transforms static images into dynamic cinematic videos using MVL technology. Supports first/last frame control and audio generation.

From$0.112/SEC
$0.095/SEC
-15%
Kling V3.0 Turbo Text-to-Video
NEW
text-to-video
TURBO

Kling V3.0 Turbo Text-to-Video

Kling V3.0 Turbo Text-to-Video generates dynamic cinematic videos from text prompts using MVL technology. Supports first/last frame control and audio generation.

From$0.112/SEC
$0.095/SEC
-15%
Kling Video O3 4K Image-to-Video
NEW
image-to-video

Kling Video O3 4K Image-to-Video

Kling Omni Video O3 (4K) Image-to-Video transforms static images into dynamic cinematic videos using MVL technology. Supports first/last frame control and audio generation.

From$0.42/SEC
$0.357/SEC
-15%
Kling Video O3 4K Text-to-Video
NEW
text-to-video

Kling Video O3 4K Text-to-Video

Kling Omni Video O3 (4K) is Kuaishou advanced unified multi-modal video model with MVL (Multi-modal Visual Language) technology. Generates high-quality videos from text prompts with natural motion and audio generation support.

From$0.42/SEC
$0.357/SEC
-15%
Youchuan V8.1 Image-to-Video
NEW
image-to-video

Youchuan V8.1 Image-to-Video

Youchuan V8.1 animates an input image into four 5-second videos at 480p or 720p.

From
$0.086/SEC
Grok Imagine Video v1.5 Reference-to-Video
NEW
image-to-video

Grok Imagine Video v1.5 Reference-to-Video

xAI Grok Imagine Video v1.5 generates video guided by 1-7 reference images plus an optional reference voice, with native synchronized audio. Up to 15s at 480p or 720p.

From
$0.08/SEC
Grok Imagine Video v1.5 Text-to-Video
NEW
text-to-video

Grok Imagine Video v1.5 Text-to-Video

xAI Grok Imagine Video v1.5 generates video with native synchronized audio from a text prompt alone. Up to 15s at 480p, 720p, or 1080p.

From
$0.08/SEC
Grok Imagine Video v1.5 Image-to-Video
NEW
image-to-video

Grok Imagine Video v1.5 Image-to-Video

xAI Grok Imagine Video v1.5 animates a starting frame image with natural-language motion prompts at 480p/720p/1080P.

From
$0.08/SEC
Gemini Omni Flash Image-to-Video Developer
NEW
image-to-video
DEV

Gemini Omni Flash Image-to-Video Developer

Gemini Omni Flash is Google's multimodal video generation model. This image-to-video variant creates subject-consistent videos from up to 7 reference images combined with a text prompt, preserving visual identity across the full generated video.

From
$0.112/SEC
Gemini Omni Flash Text-to-Video Developer
NEW
text-to-video
DEV

Gemini Omni Flash Text-to-Video Developer

Gemini Omni Flash is Google's multimodal video generation model. This text-to-video variant generates high-quality cinematic videos from text prompts with support for multiple resolutions, aspect ratios, and controllable duration.

From
$0.112/SEC
HappyHorse-1.0 Text-to-video
NEW
text-to-video

HappyHorse-1.0 Text-to-video

Generates videos from text prompts with HappyHorse 1.0, supporting 720P or 1080P output, flexible aspect ratios, and durations from 3 to 15 seconds.

From
$0.14/SEC
HappyHorse-1.0 Image-to-video
NEW
image-to-video

HappyHorse-1.0 Image-to-video

Animates a first-frame image into video with optional prompt guidance, 720P or 1080P output, and durations from 3 to 15 seconds.

From
$0.14/SEC
Seedance 2.0 Text-to-Video
NEW
text-to-video

Seedance 2.0 Text-to-Video

Generate videos from text prompts with native audio and optional web search.

From
$0.112/SEC
Seedance 2.0 Image-to-Video
NEW
image-to-video

Seedance 2.0 Image-to-Video

Generate videos from a first-frame image (and optional last-frame) with native audio.

From
$0.112/SEC
Seedance 2.0 Reference-to-Video
NEW
image-to-video

Seedance 2.0 Reference-to-Video

Multimodal video generation from reference images, videos, and audio. Supports video editing and extension.

From
$0.112/SEC
Seedance 2.0 Fast Text-to-Video
NEW
text-to-video

Seedance 2.0 Fast Text-to-Video

Fast video generation from text prompts with native audio.

From$0.09/SEC
$0.072/SEC
-20%
Seedance 2.0 Fast Image-to-Video
NEW
image-to-video

Seedance 2.0 Fast Image-to-Video

Fast video generation from first-frame image (and optional last-frame) with native audio.

From$0.09/SEC
$0.072/SEC
-20%
Seedance 2.0 Fast Reference-to-Video
NEW
image-to-video

Seedance 2.0 Fast Reference-to-Video

Fast multimodal video generation from reference images, videos, and audio. Supports video editing and extension.

From$0.09/SEC
$0.072/SEC
-20%
Wan-2.7 Text-to-video
NEW
HOT
text-to-video

Wan-2.7 Text-to-video

Generates videos from text prompts with multi-shot narrative, audio generation, and sound-image synchronization.

From
$0.1/SEC
Wan-2.7 Image-to-video
NEW
HOT
image-to-video

Wan-2.7 Image-to-video

Animates images into videos with first-frame, first-and-last-frame, video continuation, and audio-driven modes.

From
$0.1/SEC
Veo 3.1 Lite Text-to-video
NEW
text-to-video

Veo 3.1 Lite Text-to-video

High-efficiency Veo 3.1 Lite text-to-video: create video with synchronized audio from text prompts. Targets high-volume applications with strong price efficiency; 720p/1080p and flexible duration options. Does not support 4K outputs or Extension.

From
$0.05/SEC
Veo 3.1 Lite Start-End Frame to Video
NEW
image-to-video

Veo 3.1 Lite Start-End Frame to Video

Veo 3.1 Lite start-end frame to video: generate motion between a first and last frame with audio. Lightweight, developer-oriented option with 8s duration and 720p/1080p. Does not support 4K outputs or Extension.

From
$0.05/SEC
Veo 3.1 Lite Image-to-video
NEW
image-to-video

Veo 3.1 Lite Image-to-video

High-efficiency Veo 3.1 Lite image-to-video: animate an input image into video with synchronized audio. Cost-effective for scalable workflows; supports 720p/1080p and common aspect ratios. Does not support 4K outputs or Extension.

From
$0.05/SEC
Vidu Q3-Mix Reference to Video
NEW
image-to-video

Vidu Q3-Mix Reference to Video

Vidu Q3-Mix Reference-to-Video generates videos from 1-4 reference images with consistent subjects. Offers strong visual quality with intelligent scene transitions, smooth dynamic effects, and audio support up to 1080p.

From$0.125/SEC
$0.106/SEC
-15%
Vidu Q3 Reference to Video
NEW
image-to-video

Vidu Q3 Reference to Video

Vidu Q3 Reference-to-Video generates videos from 1-4 reference images with consistent subjects. Features intelligent camera switching with better consistency across multiple camera positions, audio support, and resolutions up to 1080p.

From$0.05/SEC
$0.042/SEC
-15%