
Seedance 2.5 Image-to-Video API by ByteDance
Generate videos from a first-frame image (and optional last-frame) with native audio.
Seedance 2.5 Image-to-Video은(는) ByteDance에서 개발한 모델입니다. Atlas Cloud(운영: Atlas Cloud AI LLC)는 해당 모델에 대한 액세스를 제공할 뿐이며 이를 소유하지 않습니다. 모든 상표는 각 소유자에게 귀속됩니다.
1. Introduction
Seedance 2.5 is ByteDance's next-generation multimodal video generation model, officially unveiled by Volcano Engine president Tan Dai at the 2026 Volcano Engine FORCE conference in Beijing on June 23, 2026. This README covers the following API model identifiers:
bytedance/seedance-2.5/text-to-videobytedance/seedance-2.5/image-to-videobytedance/seedance-2.5/reference-to-video
Succeeding the Seedance 2.0 family, Seedance 2.5 advances generative video along four axes announced at launch: native single-pass generation of clips up to 30 seconds (double the 15-second ceiling of Seedance 2.0, with substantially improved shot-to-shot camera continuity), joint conditioning on up to 50 all-modality reference assets (up from 12), precise consistency-preserving video editing and extension, and native multilingual generation spanning 10+ languages with stronger instruction following. ByteDance positions the 30-second single-pass length and the 50-asset reference capacity as industry firsts. The model completed a global enterprise beta following the announcement, with public availability rolling out through Volcano Engine, BytePlus, and the Dreamina platform in July 2026.
2. Key Features & Innovations
-
30-Second Single-Pass Generation: Produces a complete video of up to 30 seconds in one native generation pass — no stitching of shorter segments — with markedly improved camera and shot continuity across the full clip. This doubles the Seedance 2.0 family's 15-second output ceiling and enables genuine short-narrative work in a single request.
-
50 All-Modality Reference Assets: The reference-to-video variant conditions jointly on up to 50 reference materials — up to 30 reference images, 10 reference videos, and 10 reference audios in a single request, with a combined audio/video reference budget of 30 seconds. That expands Seedance 2.0's capacity (9 images and 3 audio/video clips, 15 seconds total) on every axis. Subjects, styles, motion cues, and audio timing can all be anchored to user-supplied assets at once.
-
Audio-Only References: New in this release, a single BGM track, voice track, or sound-effect track can serve as the sole reference — directly guiding visual pacing, beat matching, and lip synchronization without any accompanying image or video input.
-
Consistency-Preserving Localized Editing and Extension: Introduces precise video editing that keeps the overall frame intact while changing only targeted local elements, alongside high-fidelity temporal extension of existing footage — extending the family's world-model approach from pure generation into controllable editing workflows.
-
Three Task-Focused Variants:
text-to-videogenerates from a prompt alone;image-to-videoanimates a first frame (optionally pinning a last frame for precise start/end control);reference-to-videocomposes new footage from large multimodal reference sets. All variants share the same generation core, prompt understanding, and audio pipeline. -
Flexible Duration and Output Controls: Requests specify any duration from 4 to 30 seconds (or delegate the choice to the model), native 480p, 720p, or 1080p generation, optional FlashVSR
-srand Atlas Video Enhance-esrdelivery modes up to 4K, aspect ratios from vertical 9:16 to widescreen 16:9, MP4 or MOV containers, plus seed, watermark, and audio-generation toggles for reproducible, pipeline-ready results. -
Synchronized Audio Generation in 10+ Languages: Carries forward the Seedance 2.0 family's native audio track generation — soundscapes, effects, and speech aligned to the visuals — now sustained across the longer 30-second output window, with native speech generation in more than 10 languages and stronger instruction following for dialogue and delivery.
Atlas Video Enhance - Atlas-exclusive technology
Atlas Video Enhance is Atlas-exclusive video enhancement technology for generated and existing video. It treats resolution, texture quality, temporal stability, and motion continuity as one coordinated enhancement problem instead of applying simple scaling or sharpening.
Resolution provenance: Seedance 2.5 natively generates only 480p, 720p, or 1080p. Every -sr and -esr option first generates the nearest native source, then upscales or enhances that source. It is not a native higher-resolution Seedance render. Example: 4k-esr starts from native 1080p, not 480p or 720p. 1080p, 1080p-sr, and 1080p-esr are separate products with separate pricing: native 1080p is a 1920x1080 Seedance render, 1080p-sr FlashVSR-upscales a 720p source, and 1080p-esr Atlas-enhances a 720p source.
- Structure-aware super-resolution: Reconstructs contours, materials, depth relationships, faces, text, fabric, and natural textures without relying on edge halos or an over-sharpened look.
- Temporal-coherence constraints: Tracks motion and texture correspondence across frames to reduce flicker, texture drift, and unstable edges.
- Perceptual detail reconstruction: Suppresses blur, noise, banding, ringing, and compression blocks while preserving natural lighting and texture fidelity.
- Intelligent motion interpolation: Models motion, occlusion, and camera movement to create continuous intermediate frames, including smooth 60fps delivery.
The coordinated pipeline performs source analysis, spatiotemporal feature modeling, detail reconstruction and cleanup, motion-continuity enhancement, and delivery-grade encoding.
| Resolution option | Output |
|---|---|
720p-sr | Nearest native source is 480p, then FlashVSR to a 720-pixel short edge. |
1080p-sr | Nearest native source is 720p, then FlashVSR to a 1080-pixel short edge. |
1440p-sr | Nearest native source is 1080p, then FlashVSR to a 1440-pixel short edge. |
720p-esr | Nearest native source is 480p, then Atlas-enhances to a 720-pixel short edge. |
1080p-esr | Nearest native source is 720p, then Atlas-enhances to a 1080-pixel short edge. |
1440p-esr | Nearest native source is 1080p, then Atlas-enhances to a 1440-pixel short edge. |
4k-esr | Nearest native source is 1080p, then Atlas-enhances to a 2160-pixel short edge (3840x2160 for 16:9). This is enhanced 4K delivery, not native 4K generation. |
1080p-esr & 60fps | Nearest native source is 720p, Atlas-enhances to 1080p, and applies intelligent motion interpolation for 60fps delivery. Interpolation is ESR-only. |
All enhanced modes preserve the requested aspect ratio. Enhancement can reconstruct perceptual detail and improve temporal consistency, but it does not change the fact that the original Seedance generation was produced at the source resolution shown above.
3. Model Architecture & Technical Details
ByteDance has not yet published a technical report for Seedance 2.5; the details below reflect the official launch presentation and the model's documented lineage.
Seedance 2.5 builds on the Seedance 2.0 generation stack — a diffusion-transformer video-audio architecture coupled with a physics-informed world model for spatial and temporal consistency — and extends it in two directions. First, long-sequence generation: the model sustains subject identity, scene layout, and camera logic across a 30-second single-pass rollout, where prior versions required multi-segment stitching beyond 15 seconds. Second, large-scale multi-reference conditioning: the reference pathway scales to 50 simultaneous all-modality inputs, letting the model resolve identity, style, motion, and audio cues from a far richer conditioning set. The launch presentation additionally demonstrated consistency-preserving localized editing, indicating region-targeted denoising control over existing footage while the world model maintains global scene coherence.
At announcement the model was in a global enterprise beta ahead of its July 2026 public rollout, and platform API documentation was still being published; specifications exposed by this gateway (durations, resolutions, reference caps, output formats) follow the launch specification.
4. Performance Highlights
Formal third-party benchmarks for Seedance 2.5 have not yet been published. The launch announcement quantified its advances relative to the shipping Seedance 2.0 family:
| Capability | Seedance 2.0 | Seedance 2.5 |
|---|---|---|
| Native single-pass duration | up to 15 s | up to 30 s |
| Reference assets per request | 12 (9 images + 3 A/V clips) | up to 50 (30 images / 10 videos / 10 audios) |
| Combined reference A/V duration | 15 s | 30 s |
| Audio-only reference | — | Supported (BGM / voice / SFX track) |
| Native speech languages | — | 10+ |
| Localized (regional) editing | — | Supported, consistency-preserving |
| Shot / camera continuity | strong | substantially improved |
ByteDance states the 30-second single-pass duration exceeds the 15–20 second ceilings of competing generators, and that the 50-asset joint reference input is the largest in the industry. These are vendor claims from the launch event; independent side-by-side evaluations are expected once the model reaches general availability.
5. Intended Use & Applications
-
Short Narrative Films: A full 30-second scene — establishing shot through resolution — generated in one pass with coherent camera movement, eliminating the seam-matching work that multi-segment stitching requires.
-
Character- and Brand-Consistent Content: Anchor a character, product, or visual identity across large reference sets (up to 30 images plus video and audio cues) so serialized episodes, campaigns, and virtual-influencer content stay on-model shot after shot.
-
Music-Driven and Dialogue Videos: Supply reference audio tracks to drive timing, rhythm, and vocal delivery over the extended 30-second window — music videos, performance clips, and speaking characters with synchronized sound.
-
E-commerce and Product Showcases: Turn product stills and demonstration clips into dynamic showcase videos, keeping materials, logos, and product geometry faithful to the supplied references.
-
Social Media Content at Scale: Vertical 9:16 through widescreen 16:9 output at platform-friendly resolutions and durations, with seed control and MP4/MOV containers for automated content pipelines.
-
Iterative Editing Workflows: Use localized editing to revise a specific element of generated footage — wardrobe, prop, background detail — without regenerating or destabilizing the rest of the scene.


















