Seedance 2.5 Now Live — First on Atlas Cloud
Vidu API: Consistent Subjects from Reference Images

Vidu API: Consistent Subjects from Reference Images

The Vidu API gives developers access to video models from Shengshu AI and Tsinghua University, built on the U-ViT architecture that unifies Diffusion and Transformer design. Feed in one to four reference images and Vidu keeps subjects consistent across shots, with intelligent camera switching and smooth, cinematic motion. Atlas Cloud adds Day-0 model access and one OpenAI-compatible key for the whole lineup. Start building today.

Explore the Leading Vidu

Atlas Cloud provides you with the latest industry-leading creative models.

Every Vidu API Endpoint Compared: Q1, Q2, Q3, and 2.0 at a Glance

Scan durations, resolutions, reference image limits, and audio options across all 25 Vidu API video generation endpoints to match the right model to your workload.

ModalityDescription
Vidu Q3-Pro T2V API (Text to Video)Describe a scene in plain text and Q3-Pro renders it as video up to 16 seconds long, with five aspect ratios plus general and anime styles. Super resolution output reaches 1440p, and audio with background music generates in the same pass. It suits cinematic ads and narrative shots where fidelity matters most.
Vidu Q3-Turbo T2V API (Text to Video)Need volume more than polish? This turbo tier converts text prompts into clips of 1 to 16 seconds below Pro tier pricing, while keeping the same 1440p super resolution ceiling, style presets, and optional audio track. Teams use it for rapid storyboarding and testing creative directions.
Vidu Q3-Pro I2V API (Image to Video)Upload one still image, describe the motion you want, and the model animates it with smooth, cinematic movement for up to 16 seconds. Output scales from 540p to 1440p super resolution with generated audio and background music. Product showcases and photo animation pipelines are its natural home.
Vidu Q3-Turbo I2V API (Image to Video)When turnaround matters, the lowest priced Q3 tier brings photos to life from a single reference image and a motion prompt. It reaches 1440p super resolution and attaches audio automatically unless disabled. Social clips and bulk catalog animation run economically on it.
Vidu Q3-Pro SE2V API (Start-End to Video)Give it a first frame and a last frame, and Q3-Pro interpolates a coherent transition between them lasting up to 16 seconds. The two frames must keep their aspect ratios within 0.8 to 1.25 of each other, and audio plus background music are generated on request. Storyboards, morph effects, and scene bridges all fit here.
Vidu Q3-Turbo SE2V API (Start-End to Video)Fast frame bridging is the turbo variant's specialty. It connects a start image and an end image into one continuous shot, supports 540p through 1440p super resolution output, and toggles audio with a single flag. Iterate transitions quickly here before committing to a Pro render.
Vidu Q3 R2V API (Reference to Video)Consistency across camera angles sets this endpoint apart. Feed it 1 to 4 reference images and it keeps subjects stable through intelligent camera switching, producing clips of 3 to 16 seconds at up to 1440p super resolution with audio. Character driven ads and episodic content benefit most.
Vidu Q3-Mix R2V API (Reference to Video)Blend up to 4 reference images into one scene with intelligent transitions and smooth dynamic effects. The Q3-Mix endpoint holds subject identity while settings shift, renders clips of 1 to 16 seconds in five aspect ratios, and adds generated audio by default. Multi-shot brand stories and mashup creative are built on it.
Vidu Q2 T2V API (Text to Video)Text goes in, and a finished clip of 1 to 8 seconds comes out at up to 1080p. General and anime styles, five aspect ratios, and automatic audio with background music are all controlled through simple parameters. It makes a solid default for everyday prompt to video work.
Vidu Q2-Pro I2V API (Image to Video)Animating a single photo with prompts of up to 5000 characters, this Pro endpoint outputs 540p to 1080p clips running as long as 10 seconds. Movement amplitude can be pinned to small, medium, or large for precise motion control. Explainers and marketing stills convert well here.
Vidu Q2-Pro-Fast I2V API (Image to Video)The Fast variant keeps 720p and 1080p output and the same price as the standard Pro tier while shortening generation queues for image animation jobs. Durations stretch from 1 to 10 seconds with optional audio and background music. Pick it when pipelines demand quick iteration.
Vidu Q2-Turbo I2V API (Image to Video)At the lowest price in the entire Vidu family, this endpoint animates one image into clips of up to 10 seconds at 540p, 720p, or 1080p. Audio and background music generate automatically unless switched off. High volume UGC tools and batch jobs are where it earns its keep.
Vidu Q2-Pro SE2V API (Start-End to Video)Two keyframes become one fluid shot. Q2-Pro fills the gap between a start image and an end image over 1 to 8 seconds at up to 1080p, with generated audio and background music available. Keep the frame aspect ratios within 0.8 to 1.25 of each other for a valid request.
Vidu Q2-Pro-Fast SE2V API (Start-End to Video)Prefer speed for frame interpolation? This variant renders start to end transitions at 720p or 1080p in clips up to 8 seconds, with audio and background music toggles included. It shares standard Pro pricing while tightening iteration loops for editors.
Vidu Q2-Turbo SE2V API (Start-End to Video)The most economical start-end option in the family lives here. Supply first and last frames and receive a 1 to 8 second clip at up to 1080p, with optional audio in the same call. Prototyping transitions before a Pro render is a common workflow.
Vidu Q2 R2V API (Reference to Video)Up to 7 distinct subjects, each defined by 1 to 3 reference images, can appear together in a single generated scene. Clips run 1 to 10 seconds at up to 1080p, and an audio type switch isolates speech or sound effects. Multi-character storytelling is the headline use case.
Vidu Q2-Pro R2V API (Reference to Video)Casting several consistent characters into premium footage is what the Pro reference tier does best. It accepts up to 7 subjects with 3 images each, prompts up to 1500 characters, and outputs clips of up to 10 seconds at 1080p with speech only or sound effect only audio modes. Serialized content and virtual character work fit naturally.
Vidu Q1 T2V API (Text to Video)The first generation text endpoint outputs fixed 5 second clips locked at 1080p, with 16:9, 9:16, and 1:1 aspect ratios plus general and anime styles. Audio and background music remain switchable. Choose it when predictable, uniform clip specs simplify a pipeline.
Vidu Q1 I2V API (Image to Video)Simplicity defines this endpoint: one image, one prompt of up to 5000 characters, one 5 second 1080p clip. Movement amplitude and audio toggles are the only knobs to tune. Integrations that standardized on Q1 output keep running unchanged.
Vidu Q1 SE2V API (Start-End to Video)Interpolating between two uploaded frames, Q1 produces a 5 second 1080p transition with adjustable motion amplitude and optional audio plus background music. Frame pairs must keep their aspect ratios within 0.8 to 1.25 of each other. It remains a stable choice for fixed length transition effects.
Vidu Q1 R2V API (Reference to Video)Structured casting is the pattern here: up to 7 subjects, each backed by 1 to 3 images, render into a fixed 5 second 1080p clip. Audio narrows to speech only or sound effects only when needed. Multi-character scenes with locked output specs suit it well.
Vidu 2.0 I2V API (Image to Video)Would a short silent clip do the job? Vidu 2.0 turns one image into a 4 or 8 second video with adjustable movement amplitude and seed control for reproducible results. Thumbnails, previews, and lightweight animation tasks are typical fits.
Vidu 2.0 SE2V API (Start-End to Video)Exactly two images, a prompt of up to 1500 characters, and a choice of 4 or 8 second duration define this compact endpoint. Motion intensity runs from small to large under one movement amplitude parameter. It handles quick morphs and simple frame pair transitions.
Vidu 2.0 R2V API (Reference to Video)From 1 to 3 reference images, this endpoint composes videos in 16:9, 9:16, or 1:1 while preserving the referenced subjects. Seed and movement amplitude round out a deliberately small parameter set. Consistent character clips for social formats are the common case.
Vidu R2V Q1 API (Multi-Image Reference to Video)Broader flat input distinguishes this reference endpoint, which takes a single list of 1 to 7 images in 16:9, 9:16, or 1:1. Prompts cap at 1500 characters and each image can reach 50MB. It works well when subjects arrive as loose image sets rather than structured casts.

Inside the Vidu API: Consistency, Sound, and Cinematic Control

From four image reference consistency to native audio and sixteen second 1080p clips, the Vidu API on Atlas Cloud puts Shengshu's entire Q series behind one pay-as-you-go endpoint.

Subject Consistency with the Vidu API

The Vidu API accepts one to four reference images and holds characters, products, and scenes visually consistent across every frame. That reliability makes serialized storytelling and brand safe product videos practical at scale.

Start and End Frame Interpolation

Upload a start frame and an end frame, and the model builds the smooth transition between them on its own. Motion designers use this for morphs, reveals, and precisely timed scene changes.

Native Audio in a Single Pass

Why render silent clips and score them later? Every Q3 endpoint can generate synchronized sound effects and background music in the same pass, so dialogue scenes and ads arrive ready to publish.

Sixteen Seconds, Up to 1440p

If a scene needs room to breathe, run clips up to 16 seconds at 540p to 1080p, with super resolution upscaling to 1440p. Fewer stitched shots means smoother stories and lighter post production.

Tunable Motion and Style

A movement amplitude setting shifts output from subtle cinematic drift to large dramatic action, while text to video endpoints switch between general and anime styles. Directors get shot level control without rewriting prompts.

One Key Across the Vidu API Lineup

Every Vidu model generation from Q1 to Q3 runs behind one OpenAI compatible key with transparent pay-as-you-go pricing on Atlas Cloud. Draft on Turbo, then rerun the same call on Pro for final renders.

Vidu vs Other Models - One Prompt

The same prompt, generated by Vidu and other leading video models

Prompt

A lone stunt driver drifts a matte black muscle car through a neon-lit night market in the rain. Open on a low tracking shot skimming wet asphalt as the car slides around a corner, sparks flying off a guardrail. Whip pan to vendors yanking their stalls back, paper lanterns swinging wildly. Cut to a drone shot spiraling overhead as the car threads between two buses with inches to spare, then a hard cut inside the cockpit: hands snapping the wheel, eyes flicking to a mirror filled with flashing police lights. End on a slow-motion 180 degree spin into a steaming alley, engine roar and rain hiss filling the soundtrack. Photorealistic, cinematic lighting, 10 seconds, 16:9 aspect ratio.

Vidu Q3

Seedance 2.0

Vidu Q2 Pro

Prompt

Anime style: a young swordswoman faces a colossal storm dragon above a shattered floating shrine. Start with a fast dolly-in on her eyes as lightning splits the sky, then a whip pan following her sprint across a crumbling stone bridge, debris lifting around her feet. The camera orbits as she leaps, her blade igniting with blue flame, while the dragon's jaws crash down where she stood a frame earlier. Pull back to an epic wide shot as her strike cleaves a lightning bolt in half and the shockwave ripples through the clouds. Crisp 2D cel shading, dynamic speed lines, thunder and a swelling choir in the score. 8 seconds, 16:9 aspect ratio.

Vidu Q3 Pro

Seedance 2.0

Vidu Q2 Pro

Where the Vidu API Fits Your Production Pipeline

From character consistent brand films to start and end frame motion design, the Vidu API brings text-to-video, image-to-video, and reference-to-video workflows together behind one pay-as-you-go endpoint on Atlas Cloud.

Consistent Characters Across Every Shot

Feed 1 to 4 reference images into Vidu Q3-Mix or Q3 reference-to-video and subjects stay consistent across scene transitions. Short drama teams and IP creators keep the same face from episode to episode.

Cinematic Ads Through the Vidu API

Product launches and brand spots come to life as the Vidu API renders smooth dynamic effects with optional audio at up to 1080p. Agencies iterate on ad concepts without booking a single shoot day.

Start and End Frame Motion Design

Need a precise transition between two keyframes? Q3-Pro and Q2 start-end-to-video models interpolate smooth motion from your first frame to your last, powering logo reveals, scene morphs, and product transformations.

Script to Screen with the Vidu API

Writers turn a text prompt into 1080p footage with multiple visual styles through the Vidu API text-to-video endpoints. Optional audio generation gives storyboards and previz clips a finished feel in one pass.

Still Images That Move

When a static product photo or illustration needs motion, image-to-video models from Q1 through Q3-Turbo animate it with smooth, natural movement. E-commerce teams and social editors publish scroll-stopping clips daily.

Animation Pipelines Built on Vidu

Vidu's U-ViT architecture was built for long-form, highly consistent output, a natural fit for animation design and episodic series. Studios storyboard, animate, and revise entire sequences via API calls instead of manual keyframing.

Where the Vidu API Stands Among Kling, Veo, and Sora

Weigh clip length, resolution ceilings, audio output, and per second rates side by side, and you can see exactly what the Vidu API delivers for every dollar in your video pipeline.

ModelInput ModesMax Clip LengthTop ResolutionNative Audio
Vidu Q3-MixText plus 1 to 4 reference images16s1080p native, 1440p with super resolution
Vidu Q3 (Pro & Turbo)Text, image, start-end frames, reference images16sUp to 1080p, 1440p upscaled
Vidu Q2 (Pro & Turbo)Text, image, start-end frames, or references10s1080p
Kling 2.5 TurboText or a single image with prompt10s1080p (Pro tier)-
Google Veo 3.1Text, image, up to 3 reference images8s per clip, extendable1080p, 4K in preview
OpenAI Sora 2Text prompt12s720p (1080p on Sora 2 Pro)

How to Use Vidu on Atlas Cloud

Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.

Create an Atlas Cloud Account

Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.

Why Use Vidu on Atlas Cloud

Combining the advanced Vidu models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.

Performance & flexibility

Low Latency:
GPU-optimized inference for real-time reasoning.

Unified API:
Run Vidu, GPT, Gemini, and DeepSeek with one integration.

Transparent Pricing:
Predictable per-token billing with serverless options.

Enterprise & Scale

Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.

Reliability:
99.99% uptime, RBAC, and compliance-ready logging.

Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.

Vidu API Questions, Answered

Vidu is a video generation model family developed by Shengshu AI in collaboration with Tsinghua University, built on the U-ViT architecture that merges Diffusion and Transformer designs. The Vidu API on Atlas Cloud gives developers programmatic access to the full family, from Vidu 2.0 through the latest Q3 series, for cinematic video generation from text, images, and reference inputs.

Four modes are available: text-to-video, image-to-video, start-end-to-video for interpolating between a first and last frame, and reference-to-video for keeping subjects consistent across shots. Each mode ships in multiple tiers, so you can pick Turbo endpoints for speed and cost or Pro and Q3-Mix endpoints for maximum visual quality.

Billing is metered per second of generated video with no subscription required. Q2-Turbo endpoints start at $0.026 per second, Q3-Turbo runs $0.034 per second, Q3-Pro and Q3 reference-to-video cost $0.042 per second, and Q3-Mix reference-to-video is $0.106 per second. You only pay for what you generate under transparent per-call pricing.

Q3 endpoints support clips up to 16 seconds with resolution options of 540p, 720p, and 1080p, plus super resolution outputs up to 1440p. Aspect ratios cover 16:9, 9:16, 4:3, 3:4, and 1:1, so the same endpoint can target landscape, vertical, or square formats.

Yes. Q3 models include an audio toggle that is enabled by default, and text-to-video endpoints add a separate background music switch. Turn both off when you need silent footage for your own sound design.

Upload one to four reference images of a character, product, or scene, and the model preserves that subject's identity while generating new motion and camera angles. Q3 reference-to-video adds intelligent camera switching with better consistency across multiple camera positions, which suits episodic and multi-shot storytelling.

It does. Text-to-video endpoints expose a style parameter with general and anime options, and the family is known for stylized character animation. Pair the anime style with reference images to keep the same character across an entire series.

Create an Atlas Cloud API key, pick a Vidu endpoint, and send a REST request with your prompt and parameters. Every model has a playground and a published parameter schema, so you can validate settings before writing code. Pricing is pay-as-you-go per call, and new Vidu releases arrive with Day-0 access. Start building today.

Explore More Families

Seedance 2.5

The Seedance 2.5 API delivers ByteDance's newest video generation model, the successor to Seedance 2.0 built on a unified multimodal architecture. It renders up to 30 seconds of footage in one pass, keeps subjects consistent under believable physics, draws text and multilingual subtitles directly in frame. Atlas Cloud brings Day-0 access on the same unified endpoints that already serve Seedance 2.0 and 1.5. Start building today.

View Family

MiniMax H3

The MiniMax H3 API opens MiniMax's general purpose multimodal video model, which reads text, images, video and audio as one context instead of one task at a time. Clips run 5 to 15 seconds at 24 FPS across aspect ratios from 21:9 to 9:16, and one prompt can swap characters, replace backgrounds, rewrite dialogue or clone a voice from a reference clip. Atlas Cloud serves it all through one OpenAI-compatible endpoint. Start building today.

View Family

Seedream 5.0 Pro

Seedream 5.0 Pro API gives developers ByteDance's controllable image editing model on Atlas Cloud. It places edits precisely with anchors and coordinates, separates images into editable layers, fuses multiple references, and matches exact colors and materials, with multilingual text at 2K and 3K. On Atlas Cloud you reach it through one key!

View Family

Seedance 2.0

The Seedance 2.0 API gives you production access to ByteDance's multimodal video model — quad-modal inputs (text, image, video, audio) and an industry-leading "Universal Reference" system that locks composition, camera movement, and character actions across shots. Integrate director-level control with one API call, a flat $0.09/s, instant key, and no waitlist — backed by enterprise-grade uptime and compliance. Seedance 2.0 Native 4K is now live!

View Family

GPT Image 2

The GPT Image 2 API gives developers access to OpenAI's latest image model, the successor to GPT Image 1.5. It generates and edits images with accurate text rendering across Latin and CJK scripts, plus strong composition for posters, mockups, and infographics. On Atlas Cloud you reach it through one unified API alongside 300+ models, with free credits, 99.99% uptime, and no OpenAI organization verification required.

View Family

Gemini Omni Flash

The Gemini Omni API brings Google DeepMind's multimodal video generation and editing model, introduced at Google I/O 2026, to your stack. Gemini Omni fuses Gemini's reasoning engine with generative media, accepting any mix of text, images, video, and audio to produce consistent, knowledge-grounded output. Refine results through natural conversation, swapping objects, rewriting scenes, and shifting styles while physics, characters, and continuity stay intact. Atlas Cloud serves the full Gemini Omni Flash lineup, text-to-video, image-to-video with up to 7 reference images, and reference-to-video, through one unified API with transparent per-second pricing from $0.112 and no subscription. Start building today.

View Family

Grok Imagine

The Grok Imagine API gives developers xAI's image, video, and audio generation in one suite. It produces up to 2K images with multilingual text rendering, plus video up to 15 seconds with native, synchronized audio and reference-based editing. On Atlas Cloud one key runs every Grok Imagine mode, so you move between image, video, and audio without separate setups, from $0.02 per image and $0.05 per second.

View Family

Google

Google's most powerful creative models are all available on Atlas Cloud. Veo 3.1 delivers cinematic video generation, Nano Banana 2 powers high-fidelity image creation, and Gemini brings multimodal intelligence to every workflow. Access the full Google model suite through one API key with Day-0 availability and pay-as-you-go pricing.

View Family

Seedance 2.0 Mini

The Seedance 2.0 Mini API is the lightest, lowest-cost tier of ByteDance's Seedance video line, built for teams where throughput and unit cost matter more than maximum polish. Use it for batch generation, rapid prototyping, and draft passes, all through one OpenAI-compatible key on Atlas Cloud.

View Family

ByteDance

From cinematic video generation to high-fidelity image creation, ByteDance's most powerful models are live on Atlas Cloud. Run Seedance and Seedream at scale with the lowest inference pricing and zero infrastructure overhead.

View Family

Alibaba

Atlas Cloud brings together Alibaba's full model lineup under one API: Qwen for language and image tasks, Wan for video generation up to 1080p. Access every model pay-as-you-go with no subscriptions. The Alibaba API is available via a single base URL using your existing OpenAI-compatible client.

View Family

OpenAI

Atlas Cloud gives you access to the full OpenAI API lineup, from GPT Image 2 for image generation to Sora 2 for video. Every model is available pay-as-you-go with no monthly commitment. Plug in with a single base URL swap using the OpenAI-compatible API.

View Family

One API for All Media AI.

Explore all models