
The Vidu API gives developers access to video models from Shengshu AI and Tsinghua University, built on the U-ViT architecture that unifies Diffusion and Transformer design. Feed in one to four reference images and Vidu keeps subjects consistent across shots, with intelligent camera switching and smooth, cinematic motion. Atlas Cloud adds Day-0 model access and one OpenAI-compatible key for the whole lineup. Start building today.
Atlas Cloud provides you with the latest industry-leading creative models.
Scan durations, resolutions, reference image limits, and audio options across all 25 Vidu API video generation endpoints to match the right model to your workload.
| Modality | Description |
|---|---|
| Vidu Q3-Pro T2V API (Text to Video) | Describe a scene in plain text and Q3-Pro renders it as video up to 16 seconds long, with five aspect ratios plus general and anime styles. Super resolution output reaches 1440p, and audio with background music generates in the same pass. It suits cinematic ads and narrative shots where fidelity matters most. |
| Vidu Q3-Turbo T2V API (Text to Video) | Need volume more than polish? This turbo tier converts text prompts into clips of 1 to 16 seconds below Pro tier pricing, while keeping the same 1440p super resolution ceiling, style presets, and optional audio track. Teams use it for rapid storyboarding and testing creative directions. |
| Vidu Q3-Pro I2V API (Image to Video) | Upload one still image, describe the motion you want, and the model animates it with smooth, cinematic movement for up to 16 seconds. Output scales from 540p to 1440p super resolution with generated audio and background music. Product showcases and photo animation pipelines are its natural home. |
| Vidu Q3-Turbo I2V API (Image to Video) | When turnaround matters, the lowest priced Q3 tier brings photos to life from a single reference image and a motion prompt. It reaches 1440p super resolution and attaches audio automatically unless disabled. Social clips and bulk catalog animation run economically on it. |
| Vidu Q3-Pro SE2V API (Start-End to Video) | Give it a first frame and a last frame, and Q3-Pro interpolates a coherent transition between them lasting up to 16 seconds. The two frames must keep their aspect ratios within 0.8 to 1.25 of each other, and audio plus background music are generated on request. Storyboards, morph effects, and scene bridges all fit here. |
| Vidu Q3-Turbo SE2V API (Start-End to Video) | Fast frame bridging is the turbo variant's specialty. It connects a start image and an end image into one continuous shot, supports 540p through 1440p super resolution output, and toggles audio with a single flag. Iterate transitions quickly here before committing to a Pro render. |
| Vidu Q3 R2V API (Reference to Video) | Consistency across camera angles sets this endpoint apart. Feed it 1 to 4 reference images and it keeps subjects stable through intelligent camera switching, producing clips of 3 to 16 seconds at up to 1440p super resolution with audio. Character driven ads and episodic content benefit most. |
| Vidu Q3-Mix R2V API (Reference to Video) | Blend up to 4 reference images into one scene with intelligent transitions and smooth dynamic effects. The Q3-Mix endpoint holds subject identity while settings shift, renders clips of 1 to 16 seconds in five aspect ratios, and adds generated audio by default. Multi-shot brand stories and mashup creative are built on it. |
| Vidu Q2 T2V API (Text to Video) | Text goes in, and a finished clip of 1 to 8 seconds comes out at up to 1080p. General and anime styles, five aspect ratios, and automatic audio with background music are all controlled through simple parameters. It makes a solid default for everyday prompt to video work. |
| Vidu Q2-Pro I2V API (Image to Video) | Animating a single photo with prompts of up to 5000 characters, this Pro endpoint outputs 540p to 1080p clips running as long as 10 seconds. Movement amplitude can be pinned to small, medium, or large for precise motion control. Explainers and marketing stills convert well here. |
| Vidu Q2-Pro-Fast I2V API (Image to Video) | The Fast variant keeps 720p and 1080p output and the same price as the standard Pro tier while shortening generation queues for image animation jobs. Durations stretch from 1 to 10 seconds with optional audio and background music. Pick it when pipelines demand quick iteration. |
| Vidu Q2-Turbo I2V API (Image to Video) | At the lowest price in the entire Vidu family, this endpoint animates one image into clips of up to 10 seconds at 540p, 720p, or 1080p. Audio and background music generate automatically unless switched off. High volume UGC tools and batch jobs are where it earns its keep. |
| Vidu Q2-Pro SE2V API (Start-End to Video) | Two keyframes become one fluid shot. Q2-Pro fills the gap between a start image and an end image over 1 to 8 seconds at up to 1080p, with generated audio and background music available. Keep the frame aspect ratios within 0.8 to 1.25 of each other for a valid request. |
| Vidu Q2-Pro-Fast SE2V API (Start-End to Video) | Prefer speed for frame interpolation? This variant renders start to end transitions at 720p or 1080p in clips up to 8 seconds, with audio and background music toggles included. It shares standard Pro pricing while tightening iteration loops for editors. |
| Vidu Q2-Turbo SE2V API (Start-End to Video) | The most economical start-end option in the family lives here. Supply first and last frames and receive a 1 to 8 second clip at up to 1080p, with optional audio in the same call. Prototyping transitions before a Pro render is a common workflow. |
| Vidu Q2 R2V API (Reference to Video) | Up to 7 distinct subjects, each defined by 1 to 3 reference images, can appear together in a single generated scene. Clips run 1 to 10 seconds at up to 1080p, and an audio type switch isolates speech or sound effects. Multi-character storytelling is the headline use case. |
| Vidu Q2-Pro R2V API (Reference to Video) | Casting several consistent characters into premium footage is what the Pro reference tier does best. It accepts up to 7 subjects with 3 images each, prompts up to 1500 characters, and outputs clips of up to 10 seconds at 1080p with speech only or sound effect only audio modes. Serialized content and virtual character work fit naturally. |
| Vidu Q1 T2V API (Text to Video) | The first generation text endpoint outputs fixed 5 second clips locked at 1080p, with 16:9, 9:16, and 1:1 aspect ratios plus general and anime styles. Audio and background music remain switchable. Choose it when predictable, uniform clip specs simplify a pipeline. |
| Vidu Q1 I2V API (Image to Video) | Simplicity defines this endpoint: one image, one prompt of up to 5000 characters, one 5 second 1080p clip. Movement amplitude and audio toggles are the only knobs to tune. Integrations that standardized on Q1 output keep running unchanged. |
| Vidu Q1 SE2V API (Start-End to Video) | Interpolating between two uploaded frames, Q1 produces a 5 second 1080p transition with adjustable motion amplitude and optional audio plus background music. Frame pairs must keep their aspect ratios within 0.8 to 1.25 of each other. It remains a stable choice for fixed length transition effects. |
| Vidu Q1 R2V API (Reference to Video) | Structured casting is the pattern here: up to 7 subjects, each backed by 1 to 3 images, render into a fixed 5 second 1080p clip. Audio narrows to speech only or sound effects only when needed. Multi-character scenes with locked output specs suit it well. |
| Vidu 2.0 I2V API (Image to Video) | Would a short silent clip do the job? Vidu 2.0 turns one image into a 4 or 8 second video with adjustable movement amplitude and seed control for reproducible results. Thumbnails, previews, and lightweight animation tasks are typical fits. |
| Vidu 2.0 SE2V API (Start-End to Video) | Exactly two images, a prompt of up to 1500 characters, and a choice of 4 or 8 second duration define this compact endpoint. Motion intensity runs from small to large under one movement amplitude parameter. It handles quick morphs and simple frame pair transitions. |
| Vidu 2.0 R2V API (Reference to Video) | From 1 to 3 reference images, this endpoint composes videos in 16:9, 9:16, or 1:1 while preserving the referenced subjects. Seed and movement amplitude round out a deliberately small parameter set. Consistent character clips for social formats are the common case. |
| Vidu R2V Q1 API (Multi-Image Reference to Video) | Broader flat input distinguishes this reference endpoint, which takes a single list of 1 to 7 images in 16:9, 9:16, or 1:1. Prompts cap at 1500 characters and each image can reach 50MB. It works well when subjects arrive as loose image sets rather than structured casts. |
From four image reference consistency to native audio and sixteen second 1080p clips, the Vidu API on Atlas Cloud puts Shengshu's entire Q series behind one pay-as-you-go endpoint.
The Vidu API accepts one to four reference images and holds characters, products, and scenes visually consistent across every frame. That reliability makes serialized storytelling and brand safe product videos practical at scale.
Upload a start frame and an end frame, and the model builds the smooth transition between them on its own. Motion designers use this for morphs, reveals, and precisely timed scene changes.
Why render silent clips and score them later? Every Q3 endpoint can generate synchronized sound effects and background music in the same pass, so dialogue scenes and ads arrive ready to publish.
If a scene needs room to breathe, run clips up to 16 seconds at 540p to 1080p, with super resolution upscaling to 1440p. Fewer stitched shots means smoother stories and lighter post production.
A movement amplitude setting shifts output from subtle cinematic drift to large dramatic action, while text to video endpoints switch between general and anime styles. Directors get shot level control without rewriting prompts.
Every Vidu model generation from Q1 to Q3 runs behind one OpenAI compatible key with transparent pay-as-you-go pricing on Atlas Cloud. Draft on Turbo, then rerun the same call on Pro for final renders.
The same prompt, generated by Vidu and other leading video models
A lone stunt driver drifts a matte black muscle car through a neon-lit night market in the rain. Open on a low tracking shot skimming wet asphalt as the car slides around a corner, sparks flying off a guardrail. Whip pan to vendors yanking their stalls back, paper lanterns swinging wildly. Cut to a drone shot spiraling overhead as the car threads between two buses with inches to spare, then a hard cut inside the cockpit: hands snapping the wheel, eyes flicking to a mirror filled with flashing police lights. End on a slow-motion 180 degree spin into a steaming alley, engine roar and rain hiss filling the soundtrack. Photorealistic, cinematic lighting, 10 seconds, 16:9 aspect ratio.
Vidu Q3
Seedance 2.0
Vidu Q2 Pro
Anime style: a young swordswoman faces a colossal storm dragon above a shattered floating shrine. Start with a fast dolly-in on her eyes as lightning splits the sky, then a whip pan following her sprint across a crumbling stone bridge, debris lifting around her feet. The camera orbits as she leaps, her blade igniting with blue flame, while the dragon's jaws crash down where she stood a frame earlier. Pull back to an epic wide shot as her strike cleaves a lightning bolt in half and the shockwave ripples through the clouds. Crisp 2D cel shading, dynamic speed lines, thunder and a swelling choir in the score. 8 seconds, 16:9 aspect ratio.
Vidu Q3 Pro
Seedance 2.0
Vidu Q2 Pro
From character consistent brand films to start and end frame motion design, the Vidu API brings text-to-video, image-to-video, and reference-to-video workflows together behind one pay-as-you-go endpoint on Atlas Cloud.
Feed 1 to 4 reference images into Vidu Q3-Mix or Q3 reference-to-video and subjects stay consistent across scene transitions. Short drama teams and IP creators keep the same face from episode to episode.
Product launches and brand spots come to life as the Vidu API renders smooth dynamic effects with optional audio at up to 1080p. Agencies iterate on ad concepts without booking a single shoot day.
Need a precise transition between two keyframes? Q3-Pro and Q2 start-end-to-video models interpolate smooth motion from your first frame to your last, powering logo reveals, scene morphs, and product transformations.
Writers turn a text prompt into 1080p footage with multiple visual styles through the Vidu API text-to-video endpoints. Optional audio generation gives storyboards and previz clips a finished feel in one pass.
When a static product photo or illustration needs motion, image-to-video models from Q1 through Q3-Turbo animate it with smooth, natural movement. E-commerce teams and social editors publish scroll-stopping clips daily.
Vidu's U-ViT architecture was built for long-form, highly consistent output, a natural fit for animation design and episodic series. Studios storyboard, animate, and revise entire sequences via API calls instead of manual keyframing.
Weigh clip length, resolution ceilings, audio output, and per second rates side by side, and you can see exactly what the Vidu API delivers for every dollar in your video pipeline.
| Model | Input Modes | Max Clip Length | Top Resolution | Native Audio |
|---|---|---|---|---|
| Vidu Q3-Mix | Text plus 1 to 4 reference images | 16s | 1080p native, 1440p with super resolution | √ |
| Vidu Q3 (Pro & Turbo) | Text, image, start-end frames, reference images | 16s | Up to 1080p, 1440p upscaled | √ |
| Vidu Q2 (Pro & Turbo) | Text, image, start-end frames, or references | 10s | 1080p | √ |
| Kling 2.5 Turbo | Text or a single image with prompt | 10s | 1080p (Pro tier) | - |
| Google Veo 3.1 | Text, image, up to 3 reference images | 8s per clip, extendable | 1080p, 4K in preview | √ |
| OpenAI Sora 2 | Text prompt | 12s | 720p (1080p on Sora 2 Pro) | √ |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced Vidu models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run Vidu, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
Vidu is a video generation model family developed by Shengshu AI in collaboration with Tsinghua University, built on the U-ViT architecture that merges Diffusion and Transformer designs. The Vidu API on Atlas Cloud gives developers programmatic access to the full family, from Vidu 2.0 through the latest Q3 series, for cinematic video generation from text, images, and reference inputs.
Four modes are available: text-to-video, image-to-video, start-end-to-video for interpolating between a first and last frame, and reference-to-video for keeping subjects consistent across shots. Each mode ships in multiple tiers, so you can pick Turbo endpoints for speed and cost or Pro and Q3-Mix endpoints for maximum visual quality.
Billing is metered per second of generated video with no subscription required. Q2-Turbo endpoints start at $0.026 per second, Q3-Turbo runs $0.034 per second, Q3-Pro and Q3 reference-to-video cost $0.042 per second, and Q3-Mix reference-to-video is $0.106 per second. You only pay for what you generate under transparent per-call pricing.
Q3 endpoints support clips up to 16 seconds with resolution options of 540p, 720p, and 1080p, plus super resolution outputs up to 1440p. Aspect ratios cover 16:9, 9:16, 4:3, 3:4, and 1:1, so the same endpoint can target landscape, vertical, or square formats.
Yes. Q3 models include an audio toggle that is enabled by default, and text-to-video endpoints add a separate background music switch. Turn both off when you need silent footage for your own sound design.
Upload one to four reference images of a character, product, or scene, and the model preserves that subject's identity while generating new motion and camera angles. Q3 reference-to-video adds intelligent camera switching with better consistency across multiple camera positions, which suits episodic and multi-shot storytelling.
It does. Text-to-video endpoints expose a style parameter with general and anime options, and the family is known for stylized character animation. Pair the anime style with reference images to keep the same character across an entire series.
Create an Atlas Cloud API key, pick a Vidu endpoint, and send a REST request with your prompt and parameters. Every model has a playground and a published parameter schema, so you can validate settings before writing code. Pricing is pay-as-you-go per call, and new Vidu releases arrive with Day-0 access. Start building today.
Guides, tutorials, and product updates to help you get the most out of Atlas Cloud.