MiniMax H3 Developer Now Live — 60% Off, From $0.02 per Second
Seedance 2.5 Image-to-Video
image-to-video

Seedance 2.5 Image-to-Video

Generate videos from a first-frame image (and optional last-frame) with native audio.

Seedance 2.5 Image-to-Video
Seedream v5.0 Pro Edit
Nano Banana 2 Reference to Image
Kling Video O3 4K Image-to-Video
Wan-2.7 Text-to-video
CATEGORY
Discount (148)
Model function
Series
48 of 340 models
New
MiniMax H3-Developer Text-to-Video
NEW
text-to-video
DEV

MiniMax H3-Developer Text-to-Video

MiniMax H3-Developer self-hosted text-to-video: generate a video (with audio) from a text prompt. Supports 768P/1080P, 16:9/9:16/1:1 aspect ratios, tunable seed and inference steps.

VIDEO
From$0.05/SEC
$0.02/SEC
-60%
MiniMax H3-Developer Image-to-Video
NEW
image-to-video
DEV

MiniMax H3-Developer Image-to-Video

MiniMax H3-Developer self-hosted image-to-video: animate a first-frame image (optionally a last frame) driven by a text prompt, with generated audio. Supports 768P/1080P.

VIDEO
From$0.05/SEC
$0.02/SEC
-60%
MiniMax H3-Developer Reference-to-Video
NEW
image-to-video
DEV

MiniMax H3-Developer Reference-to-Video

MiniMax H3-Developer self-hosted reference-to-video: generate a video that keeps the subject from one or more reference images/videos, driven by a text prompt, with generated audio. Supports 768P/1080P.

VIDEO
From$0.05/SEC
$0.02/SEC
-60%
Wan-3.0-Prime Text-to-video
NEW
HOT
text-to-video

Wan-3.0-Prime Text-to-video

All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.

From$0.068/SEC
$0.061/SEC
-10%
Wan-3.0-Prime Image-to-video
NEW
HOT
image-to-video

Wan-3.0-Prime Image-to-video

Animate a first frame (optionally with a last frame) into a coherent clip, with native audio and smart-duration up to 30s.

From$0.068/SEC
$0.061/SEC
-10%
Wan-3.0-Prime Reference-to-video
NEW
video-to-video

Wan-3.0-Prime Reference-to-video

All-in-One Reference: keep subjects consistent from any mix of reference images, videos, and audio; pixel-level identity/voice/space alignment.

From$0.068/SEC
$0.061/SEC
-10%
Wan-3.0 Text-to-video
NEW
HOT
text-to-video

Wan-3.0 Text-to-video

All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.

From$0.05/SEC
$0.04/SEC
-20%
Wan-3.0 Image-to-video
NEW
HOT
image-to-video

Wan-3.0 Image-to-video

Animate a first frame (optionally with a last frame) into a coherent clip, with native audio and smart-duration up to 30s.

From$0.05/SEC
$0.04/SEC
-20%
Wan-3.0 Reference-to-video
NEW
video-to-video

Wan-3.0 Reference-to-video

All-in-One Reference: keep subjects consistent from any mix of reference images, videos, and audio; pixel-level identity/voice/space alignment.

From$0.05/SEC
$0.04/SEC
-20%
Seedance 2.5 Reference-to-Video
NEW
image-to-video

Seedance 2.5 Reference-to-Video

Multimodal video generation from reference images, videos, and audio. Supports video editing and extension.

From$0.167/SEC
$0.134/SEC
1080P -36%
Seedance 2.5 Image-to-Video
NEW
image-to-video

Seedance 2.5 Image-to-Video

Generate videos from a first-frame image (and optional last-frame) with native audio.

From$0.167/SEC
$0.134/SEC
1080P -36%
Seedance 2.5 Text-to-Video
NEW
text-to-video

Seedance 2.5 Text-to-Video

Generate videos from text prompts with native audio and optional web search.

From$0.167/SEC
$0.134/SEC
1080P -36%
MiniMax Music 3.0
NEW
text-to-audio

MiniMax Music 3.0

MiniMax Music 3.0 is MiniMax's 11.1B-parameter open-weights music model that turns a musical description and optional lyrics into a complete, fully arranged and mixed song of up to five minutes - vocals, instrumentation and production included - in a single generation, with section-tag control over the arrangement and vocal or instrumental output.

From
$0.15/gen
MiniMax Lyrics Generation
NEW
text-to-audio

MiniMax Lyrics Generation

MiniMax Lyrics Generation is a dedicated lyric-writing model that turns a one-line theme into a complete, professionally structured set of song lyrics - title, style tags, and sections marked with [Verse]/[Chorus] structure tags - and can also edit, continue, or restructure existing lyrics, with output directly usable as the lyrics input of MiniMax's music models.

From
$0.01/gen
Grok Imagine Image 2.0 Text-to-Image
NEW
text-to-image

Grok Imagine Image 2.0 Text-to-Image

xAI Grok Imagine Image 2.0 generates polished visuals from natural-language prompts at 1K or 2K resolution, with 14 aspect ratios and selectable low/medium quality tiers.

From
$0.04/PIC
Grok Imagine Image 2.0 Edit
NEW
image-to-image

Grok Imagine Image 2.0 Edit

xAI Grok Imagine Image 2.0 edits up to three reference images with natural-language instructions at 1K or 2K resolution, with selectable low/medium quality tiers.

From
$0.04/PIC
Qwen Image 3.0 Pro Text-to-Image
NEW
text-to-image
PRO

Qwen Image 3.0 Pro Text-to-Image

Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen strength in complex text rendering and precise prompt adherence

From
$0.04/PIC
Qwen Image 3.0 Pro Edit
NEW
image-to-image
PRO

Qwen Image 3.0 Pro Edit

Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes

From
$0.04/PIC
Seedream v5.0 Pro Layer Decomposition
NEW
image-to-image
PRO

Seedream v5.0 Pro Layer Decomposition

ByteDance flagship image layer decomposition. Splits a single input image into an editable stack: one base image plus up to 16 transparent PNG layers, each returned with stacking order (z_index), bounding box coordinates, name, and description for downstream drag/scale/recompose editing.

From
$0.022/PIC
Suno chirp-v4-5-all
NEW
text-to-audio

Suno chirp-v4-5-all

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

From
$0.132/gen
Suno chirp-v4-5-plus
NEW
text-to-audio

Suno chirp-v4-5-plus

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

From
$0.132/gen
Qwen Image 3.0 Text-to-Image
NEW
text-to-image

Qwen Image 3.0 Text-to-Image

Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen strength in complex text rendering and precise prompt adherence

From
$0.03/PIC
Qwen Image 3.0 Edit
NEW
image-to-image

Qwen Image 3.0 Edit

Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes

From
$0.03/PIC
Suno chirp-auk
NEW
text-to-audio

Suno chirp-auk

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

From
$0.132/gen
Suno chirp-fenix
NEW
text-to-audio

Suno chirp-fenix

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

From
$0.132/gen
Suno chirp-v3-5
NEW
text-to-audio

Suno chirp-v3-5

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

From
$0.132/gen
Suno chirp-v4
NEW
text-to-audio

Suno chirp-v4

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

From
$0.132/gen
Suno chirp-v5
NEW
text-to-audio

Suno chirp-v5

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

From
$0.132/gen
MiniMax H3 Text-to-Video
NEW
text-to-video

MiniMax H3 Text-to-Video

MiniMax H3 text-to-video: generate a cinematic video from a text prompt. Supports 2K, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.

From
$0.08/SEC
MiniMax H3 Image-to-Video
NEW
image-to-video

MiniMax H3 Image-to-Video

MiniMax H3 image-to-video: animate a first-frame image (optionally with a last frame) driven by a text prompt. Supports 2K, 5-15s.

From
$0.08/SEC
MiniMax H3 Reference-to-Video
NEW
image-to-video

MiniMax H3 Reference-to-Video

MiniMax H3 reference-to-video: generate a video that keeps the subject from a reference image, driven by a text prompt. Supports 2K, 5-15s.

From
$0.08/SEC
Reve 2.1 Remix
NEW
image-to-image

Reve 2.1 Remix

Reve 2.1 Remix composes one to six reference images with a natural-language prompt into a single coherent image at native 4K, blending subject, style, and background while keeping references consistent.

From
$0.24/PIC
Reve 2.1 Edit
NEW
image-to-image

Reve 2.1 Edit

Reve 2.1 Edit applies precise, instruction-driven, element-level edits to a single input image at native 4K, changing targeted regions while preserving the rest of the scene.

From
$0.24/PIC
Reve 2.1 Text-to-Image
NEW
text-to-image

Reve 2.1 Text-to-Image

Reve 2.1 is a layout-first text-to-image model that turns natural-language prompts into sharp, production-ready images at native 4K, with best-in-class typography and high prompt adherence.

From
$0.24/PIC
Youchuan V8.2 Image-to-Video
NEW
image-to-video

Youchuan V8.2 Image-to-Video

Youchuan V8.2 animates an input image into four 5-second videos at 480p or 720p.

From
$0.086/SEC
Youchuan V8.2 Remove Background
NEW
image-to-image

Youchuan V8.2 Remove Background

Youchuan automatically removes the background from an input image, returning one transparent-background result.

From
$0.086/PIC
Youchuan V8.2 Style Transfer
NEW
image-to-image

Youchuan V8.2 Style Transfer

Youchuan retexture changes the artistic style of an input image while preserving its composition, returning four restyled results.

From
$0.129/PIC
Youchuan V8.2 Blend
NEW
image-to-image

Youchuan V8.2 Blend

Youchuan V8.2 blends two to five input images into four fused results, with an optional guiding prompt and native 2K HD.

From
$0.086/PIC
Youchuan V8.2 Image-to-Image
NEW
image-to-image

Youchuan V8.2 Image-to-Image

Youchuan V8.2 re-imagines an input image guided by a text prompt, returning four variations. Supports native 2K HD, style reference, and aspect-ratio / stylize / chaos / weird controls.

From
$0.086/PIC
Youchuan V8.2 Text-to-Image
NEW
text-to-image

Youchuan V8.2 Text-to-Image

Youchuan V8.2 generates four images from a text prompt, with optional native 2K HD, a style reference, and aspect-ratio / stylize / chaos / weird controls.

From
$0.086/PIC
Seedream v5.0 Pro Edit
NEW
HOT
image-to-image
PRO

Seedream v5.0 Pro Edit

ByteDance flagship next-generation image editing model. Supports up to 10 reference images while preserving identity, lighting, and color tones for professional-quality modifications.

From
$0.045/PIC
Seedream v5.0 Pro Text-to-Image
NEW
HOT
text-to-image
PRO

Seedream v5.0 Pro Text-to-Image

ByteDance flagship next-generation image generation model with stronger prompt adherence, refined typography, and photorealistic detail. Single-image output at 1.5K and 2K tiers with JPEG and PNG support.

From
$0.045/PIC
Nano Banana 2 Lite Edit Developer
NEW
image-to-image
DEV

Nano Banana 2 Lite Edit Developer

Google's fastest and most cost-efficient Nano Banana image model for editing, applying natural-language edits and multi-image composition to up to 14 reference images with low latency.

From$0.04/PIC
$0.028/PIC
-30%
Nano Banana 2 Lite Text-to-Image Developer
NEW
text-to-image
DEV

Nano Banana 2 Lite Text-to-Image Developer

Google's fastest and most cost-efficient Nano Banana image model, turning natural-language text prompts into high-quality 1k images in as little as 4 seconds for rapid, high-volume generation.

From$0.04/PIC
$0.028/PIC
-30%
Nano Banana 2 Lite Edit
NEW
image-to-image

Nano Banana 2 Lite Edit

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

From
$0.04/PIC
Nano Banana 2 Lite Text-to-image
NEW
text-to-image

Nano Banana 2 Lite Text-to-image

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

From
$0.04/PIC
Seed Audio 1.0
NEW
text-to-audio

Seed Audio 1.0

Doubao‑Audio‑Generate‑1.0 is Doubao Voice’s next‑generation audio‑generation engine. The industry‑first commercial tool creates film‑grade audio with just one prompt. It eliminates cumbersome audio‑engineering work. Creators generate publish‑ready radio dramas, podcasts and branded audio easily, shifting from a simple voice‑generator to an AI audio director. It serves audiobooks, serialized episodes and commercial audio for high‑quality narrative‑driven production.

AUDIO-GENERATION
From
$0.143/min
Seedance 2.0 Mini Reference-to-Video
NEW
image-to-video

Seedance 2.0 Mini Reference-to-Video

Lightweight, economical multimodal video generation from reference images, videos, and audio with native audio.

From$0.056/SEC
$0.039/SEC
-30%