# Wan 3.0 Reference-to-Video — Atlas Cloud API

> All-in-One Reference: keep subjects consistent from any mix of reference images, videos, and audio; pixel-level identity/voice/space alignment.

This is the machine-readable API reference for **Wan 3.0 Reference-to-Video** on Atlas Cloud,
a unified API platform for 400+ AI models across text, image, video, audio and 3D.

- **Model ID**: `alibaba/wan-3.0/reference-to-video`
- **Built by**: Alibaba
- **Modality**: Video
- **Model page**: https://www.atlascloud.ai/models/alibaba/wan-3.0/reference-to-video
- **API key**: https://www.atlascloud.ai/console/api-keys
- **Docs**: https://www.atlascloud.ai/docs

## Pricing on Atlas Cloud

- $0.05 per second of generated video
- Pay-as-you-go. No minimum spend, no subscription required.

> **These are the authoritative Atlas Cloud rates for this model.** Any price that
> appears in the vendor description further down refers to a different platform or
> a different model variant and does not apply here.

## Use this model from an AI agent

Atlas Cloud ships three first-party integration surfaces. All three authenticate
with the same API key via the `ATLASCLOUD_API_KEY` environment variable.

### MCP server

The official MCP server (`atlascloud-mcp`) exposes this model to any
MCP-compatible host — Claude Code, OpenAI Codex, Cursor, Gemini CLI, Goose,
Claude Desktop. One-line install:

```bash
# Claude Code
claude mcp add atlascloud -- npx -y atlascloud-mcp

# OpenAI Codex CLI
codex mcp add atlascloud -- npx -y atlascloud-mcp

# Gemini CLI
gemini mcp add atlascloud -- npx -y atlascloud-mcp

export ATLASCLOUD_API_KEY="your-api-key"
```

Then ask in plain English; the agent calls `atlas_generate_video` with `model: "alibaba/wan-3.0/reference-to-video"`.
The server fetches each model's schema and validates parameters before submitting,
so invalid requests fail fast without spending credits.

MCP docs: https://www.atlascloud.ai/docs/mcp-server

### Agent Skills

`atlas-cloud-skills` is a portable skill package (API reference, code templates in
Python / Node.js / cURL, model IDs with pricing) for Claude Code, Cursor, Codex and
12+ other agents:

```bash
npx skills add AtlasCloudAI/atlas-cloud-skills
export ATLASCLOUD_API_KEY="your-api-key"
```

Skills docs: https://www.atlascloud.ai/docs/skills

### CLI

The `atlas` binary runs Atlas Cloud from a terminal or CI script. Async media jobs
are polled and downloaded automatically (use `--no-download` when a script only
needs the output URLs):

```bash
# Install (Homebrew, npm, or shell installer)
brew install AtlasCloudAI/tap/atlascloud
# npm install -g atlascloud-cli
# curl -fsSL https://raw.githubusercontent.com/AtlasCloudAI/cli/main/install.sh | sh

atlas auth login
atlas generate video alibaba/wan-3.0/reference-to-video -p "Your prompt here"
```

CLI docs: https://www.atlascloud.ai/docs/cli

## HTTP API reference

- **Submit endpoint (POST)**: `https://api.atlascloud.ai/api/v1/model/generateVideo` — start an async generation; returns a `prediction_id`
- **Poll endpoint (GET)**: `https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}` — poll this until the prediction finishes
- **Model ID**: `alibaba/wan-3.0/reference-to-video`


## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
See the input and output schema below, as well as the usage examples.


### Input Schema

The API accepts the following input parameters:

- **`model`** (`string`, _required_):
  model name
  - Default: `"alibaba/wan-3.0/reference-to-video"`

- **`prompt`** (`string`, _required_):
  The text prompt describing the scene and action for the referenced subjects (up to 20000 characters).

- **`refers`** (`array[object]`, _optional_):
  All-in-One Reference materials. Any mix of reference images (<=10), reference videos (<=5, total <=15s), and reference audio (<=5, total <=15s), each a public URL. 'type' is optional — inferred from the URL when omitted. Image: jpeg/jpg/png/bmp/webp; video: mp4/mov; audio: wav/mp3. Reference mode is mutually exclusive with first/last-frame mode.
  - Min items: 1
  - Max items: 20
  - Item properties:
    - **`url`** (`string`, _required_):
      Public URL of the reference image / video / audio.

    - **`type`** (`string`, _optional_):
      Media kind. Optional; inferred from the URL when omitted.
      - Options: "image", "video", "audio"


- **`resolution`** (`string`, _optional_):
  The resolution of the generated video.
  - Default: `"1080P"`
  - Options: "1080P", "720P", "480P"

- **`duration`** (`number`, _optional_):
  Video length in seconds (2-30). Pass -1 for smart-duration (the model picks the best length).
  - Default: `5`
  - Options: -1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30

- **`ratio`** (`string`, _optional_):
  The aspect ratio of the generated video. Use 'adaptive' to let the model choose from the references.
  - Default: `"adaptive"`
  - Options: "adaptive", "16:9", "4:3", "1:1", "3:4", "9:16"

- **`audio`** (`boolean`, _optional_):
  Whether the output video includes an audio track. Same price either way.
  - Default: `true`

- **`enable_thinking`** (`boolean`, _optional_):
  Enable deep thinking mode. Required when providing a file or link input (document / webpage parsing). Not recommended when no file/link is provided.
  - Default: `true`

- **`file`** (`string`, _optional_):
  Optional document input to parse (requires enable_thinking). Public URL. Formats: docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md; <= 100MB, <= 50 pages. Mutually exclusive with 'link'.

- **`link`** (`string`, _optional_):
  Optional public webpage URL to parse (requires enable_thinking). Only pages that need no login. Mutually exclusive with 'file'.

- **`seed`** (`integer`, _optional_):
  The random seed to use for the generation. -1 means a random seed will be used.



**Required Parameters Example**:

```json
{
  "model": "alibaba/wan-3.0/reference-to-video",
  "prompt": ""
}
```


**Full Example**:

```json
{
  "model": "alibaba/wan-3.0/reference-to-video",
  "prompt": "",
  "refers": [
    {
      "url": "",
      "type": "image"
    }
  ],
  "resolution": "1080P",
  "duration": 5,
  "ratio": "adaptive",
  "audio": true,
  "enable_thinking": true,
  "file": "",
  "link": "",
  "seed": 0
}
```


### Output Schema

The API returns the following output format:


- **`created_at`** (`string`, _optional_):
  ISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”).

- **`id`** (`string`, _optional_):
  Unique identifier for the prediction, the ID of the prediction to get.

- **`model`** (`string`, _optional_):
  Model ID used for the prediction.

- **`outputs`** (`array[string]`, _optional_):
  Array of URLs to the generated content (empty when status is not completed).

- **`status`** (`string`, _optional_):
  Status of the task: created, processing, completed, or failed.

- **`urls`** (`object`, _optional_):
  Object containing related API endpoints.



**Example Response**:

```json
{
  "created_at": "",
  "id": "",
  "model": "",
  "outputs": [
    ""
  ],
  "status": "",
  "urls": {}
}
```


## Usage Examples

### cURL

```bash
# Step 1: Start generation (async)
curl -X POST "https://api.atlascloud.ai/api/v1/model/generateVideo" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "alibaba/wan-3.0/reference-to-video",
  "prompt": "",
  "refers": [
    {
      "url": "",
      "type": "image"
    }
  ],
  "resolution": "1080P",
  "duration": 5,
  "ratio": "adaptive",
  "audio": true,
  "enable_thinking": true,
  "file": "",
  "link": "",
  "seed": 0
}'

# Response will contain: {"code": 200, "data": {"id": "prediction_id", "status": "processing"}}

# Step 2: Poll for result (replace {prediction_id} with the id returned above)
curl -X GET "https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY"

# Keep polling until status is "completed", "succeeded" or "failed"
# When completed, outputs will contain the generated content URL(s)
```

## Additional Resources

### Documentation

- [Model Playground](https://www.atlascloud.ai/models/alibaba/wan-3.0/reference-to-video)

## About this model

_Vendor-supplied description. Any pricing or endpoint mentioned below refers to_
_other platforms — use the Atlas Cloud values above._

#### Wan 3.0 Reference-to-Video 

**Wan 3.0 Reference-to-Video** is the "all-in-one reference" mode: give it any mix of reference images, videos, and audio, and it keeps those subjects, props, voices, and spatial relationships consistent throughout a new clip driven by your prompt. This is pixel-level identity preservation — not "close enough," but faithful replication of the reference details.

#### Why Choose This?

-   **Mix four modalities** Combine reference images, reference videos, and reference audio in one request.

-   **Pixel-level consistency** Characters, props, voices, and spatial relationships stay aligned across the whole clip.

-   **Multi-subject scenes** Reference several subjects at once and direct them with your prompt.

-   **Native long-form + audio** Up to 30 seconds with a synchronized audio track.

-   **Flexible aspect ratios** adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16.

#### Parameters

| Parameter | Required | Description |
| --- | --- | --- |
| prompt | Yes | Text description of the scene and action for the subjects (up to 5000 chars) |
| refers | Yes | Array of reference materials, each `{ "url": "...", "type": "image\|video\|audio" }`. `type` is optional (inferred from the URL). At least one reference is required. |
| resolution | No | Output quality: 1080P (default), 720P, 480P |
| duration | No | Output length in seconds: 5 (default), 2–30. Pass -1 for smart-duration |
| ratio | No | Aspect ratio: adaptive (default), 16:9, 4:3, 1:1, 3:4, 9:16 |
| audio | No | Whether the output has an audio track: true (default) / false |
| enable_thinking | No | Deep thinking mode: false (default). Required for file/link parsing |
| file | No | Document to parse (requires enable_thinking): docx/doc/xlsx/xls/ppt/pdf/txt/key/pages/numbers/md , ≤100MB, ≤50 pages. Mutually exclusive with link |
| link | No | Public webpage URL to parse (requires enable_thinking; no-login pages only). Mutually exclusive with file |

##### Reference limits

_(Description truncated. Full text on the model page.)_

---

Atlas Cloud — one API for 400+ AI models. Model page: https://www.atlascloud.ai/models/alibaba/wan-3.0/reference-to-video · Docs: https://www.atlascloud.ai/docs
