Seedance 2.0 Mini & Fast API dengan harga terendah di dunia — diskon hingga 68% dari harga resmi
Beranda
Jelajahi
ByteDance
Seedream 5.0 Pro
bytedance/seedream-v5.0-pro/layer-decomposition
Seedream v5.0 Pro Layer Decomposition
gambar-ke-gambar
PRO

Seedream v5.0 Pro Layer Decomposition API by ByteDance

bytedance/seedream-v5.0-pro/layer-decomposition
Layer-decomposition

ByteDance flagship image layer decomposition. Splits a single input image into an editable stack: one base image plus up to 16 transparent PNG layers, each returned with stacking order (z_index), bounding box coordinates, name, and description for downstream drag/scale/recompose editing.

Seedream v5.0 Pro Layer Decomposition dikembangkan oleh ByteDance. Atlas Cloud (dioperasikan oleh Atlas Cloud AI LLC) menyediakan akses ke model ini dan tidak memilikinya. Semua merek dagang adalah milik pemiliknya masing-masing.

1. Introduction

Seedream 5.0 Pro Layer Decomposition is the intelligent layer-separation variant of ByteDance's flagship Seedream 5.0 Pro model, introduced with the Pro release on July 8, 2026 under the theme "Beyond Generation, It Understands Design." It covers the API model identifier:

  • bytedance/seedream-v5.0-pro/layer-decomposition

Supply a single image and the model decomposes it into an editable stack: one base image plus up to 16 independent transparent layers, each returned with its stacking order, bounding-box coordinates, a machine-generated name, and a semantic description. Regions of the background previously hidden behind extracted subjects are seamlessly inpainted, so every layer — and the base beneath it — is a complete, standalone design asset ready for dragging, scaling, and recomposition (ByteDance Seed).

This capability is the layer-separation pillar of Seedream 5.0 Pro's interactive precision editing: grounded in the model's understanding of spatial positions and regional semantics, it turns a flat bitmap — a poster, a banner, a product shot — back into the kind of layered document a designer would have built by hand.


2. Key Features & Innovations

  • Intelligent Full-Image Decomposition: With no prompt at all, the model automatically identifies every major element — text blocks, subjects, decorations, background — and splits each into its own layer. Complex posters decompose into ten or more independent layers in a single pass.

  • Three Targeting Modes: Omit the prompt for automatic full decomposition; describe target elements in natural language ("split out the parrot and the headline text"); or pinpoint elements exactly with <bbox>x1 y1 x2 y2</bbox> coordinate tags using normalized [0, 1000] coordinates — a natural fit for click- or box-select canvas interactions.

  • Occlusion Inpainting: Background areas obscured by extracted subjects are restored seamlessly in the output base image, so removing a foreground layer never leaves a hole.

  • Transparent, Recomposable Assets: Every layer is delivered as a transparent PNG that preserves its own aspect ratio from the original image; the base image follows the input's aspect ratio. Stacking the layers back in z-order reconstructs the full picture.

  • Structured Layer Metadata: Each layer carries z_index (bottom-up stacking order), a bounding_box in both absolute base-image pixels and normalized [0, 1000] coordinates, plus a model-generated name and richer description — everything a layer panel or canvas editor needs, aligned one-to-one with the output image list.

  • Design-Grade Text Separation: Inherits the Seedream 5.0 family's class-leading multilingual and CJK text rendering, extracting headlines, taglines, and labels as clean standalone text layers — the foundation for localization and copy-swap workflows.

  • Predictable Output Contract: At most 17 images per request (1 base + 16 layers), ordered base-first then ascending z-order, with all-or-nothing reliability — a request never returns a partially decomposed stack.


3. Model Architecture & Technical Details

Layer decomposition runs on the same multimodal transformer architecture as Seedream 5.0 Pro's generation and editing modes — advanced vision encoders feeding a diffusion-based decoder through a reasoning-driven planning stage. For decomposition, the planning stage performs regional-semantic analysis of the input: segmenting elements, resolving their stacking relationships, and scheduling both the per-layer extractions and the inpainting of newly exposed base regions, all synthesized in one pass.

Input is a single image (URL or Base64) in JPEG, PNG, WEBP, BMP, TIFF, or GIF format, up to 30MB, with total pixels between 512×512 and 6000×6000 and aspect ratio between 1/16 and 16. Output resolution is selected by tier keyword — 1K, 1.5K, 2K, or auto (the default, which follows the input's own size within the 1K–2K envelope). The base image renders as JPEG or PNG per the output_format setting, while layers are always transparent PNG. Bounding-box coordinates are expressed in the output base image's coordinate system with the origin at the top left, making layer placement a direct pixel mapping.


4. Performance Highlights

Layer separation is a new capability class without an established public benchmark; observed behavior frames its practical performance:

ScenarioInputResult
Automatic full decomposition1:1 poster, 2K tier1 base + 8 layers (6 text blocks, subject, foliage) in ~2 minutes
Prompted 3-element split1:1 poster, 1K tier1 base + 3 layers in ~1 minute
Complex poster (family showcase)mixed text + subject + decorations10+ independent layers with occluded background restored

Element identification quality draws directly on the Pro model's strengths: text blocks are extracted with the family's leading CJK typography fidelity, subject boundaries follow regional semantics rather than rectangular crops, and generated layer names and descriptions are accurate enough to serve as layer-panel labels without human editing.


5. Intended Use & Applications

  • Design Asset Extraction: Turn finished posters, banners, and marketing visuals back into editable layered assets — recover the text, subject, and background as independent pieces from a flat export.

  • Localization & Copy Revision: Extract text elements as standalone layers, then swap, translate, or restyle them while the artwork beneath remains untouched.

  • Interactive Canvas Editors: Power click-to-select and box-select editing experiences — the normalized <bbox> prompt syntax and per-layer bounding-box metadata map directly onto canvas coordinates for drag, scale, and reorder workflows.

  • E-Commerce Creative Repurposing: Decompose product creatives to isolate the product, price tags, and promotional text, then recompose them across formats, campaigns, and marketplaces.

  • Subject Isolation with Clean Plates: Extract a subject and simultaneously receive an inpainted background plate — both halves usable independently for compositing.

  • Motion Graphics Preparation: Layered output with z-order and transparency feeds directly into parallax, reveal, and micro-animation pipelines that require separated foreground and background elements.

Jelajahi Model Serupa

Satu API untuk semua AI multimedia.

Jelajahi semua model