# Wan 2.7 Reference-to-Video — Atlas Cloud API

> Generates character-driven videos from reference images and videos, with multi-subject and voice-cloning support.

This is the machine-readable API reference for **Wan 2.7 Reference-to-Video** on Atlas Cloud,
a unified API platform for 400+ AI models across text, image, video, audio and 3D.

- **Model ID**: `alibaba/wan-2.7/reference-to-video`
- **Built by**: Alibaba
- **Modality**: Video
- **Model page**: https://www.atlascloud.ai/models/alibaba/wan-2.7/reference-to-video
- **API key**: https://www.atlascloud.ai/console/api-keys
- **Docs**: https://www.atlascloud.ai/docs

## Pricing on Atlas Cloud

- $0.1 per second of generated video
- Pay-as-you-go. No minimum spend, no subscription required.

> **These are the authoritative Atlas Cloud rates for this model.** Any price that
> appears in the vendor description further down refers to a different platform or
> a different model variant and does not apply here.

## Use this model from an AI agent

Atlas Cloud ships three first-party integration surfaces. All three authenticate
with the same API key via the `ATLASCLOUD_API_KEY` environment variable.

### MCP server

The official MCP server (`atlascloud-mcp`) exposes this model to any
MCP-compatible host — Claude Code, OpenAI Codex, Cursor, Gemini CLI, Goose,
Claude Desktop. One-line install:

```bash
# Claude Code
claude mcp add atlascloud -- npx -y atlascloud-mcp

# OpenAI Codex CLI
codex mcp add atlascloud -- npx -y atlascloud-mcp

# Gemini CLI
gemini mcp add atlascloud -- npx -y atlascloud-mcp

export ATLASCLOUD_API_KEY="your-api-key"
```

Then ask in plain English; the agent calls `atlas_generate_video` with `model: "alibaba/wan-2.7/reference-to-video"`.
The server fetches each model's schema and validates parameters before submitting,
so invalid requests fail fast without spending credits.

MCP docs: https://www.atlascloud.ai/docs/mcp-server

### Agent Skills

`atlas-cloud-skills` is a portable skill package (API reference, code templates in
Python / Node.js / cURL, model IDs with pricing) for Claude Code, Cursor, Codex and
12+ other agents:

```bash
npx skills add AtlasCloudAI/atlas-cloud-skills
export ATLASCLOUD_API_KEY="your-api-key"
```

Skills docs: https://www.atlascloud.ai/docs/skills

### CLI

The `atlas` binary runs Atlas Cloud from a terminal or CI script. Async media jobs
are polled and downloaded automatically (use `--no-download` when a script only
needs the output URLs):

```bash
# Install (Homebrew, npm, or shell installer)
brew install AtlasCloudAI/tap/atlascloud
# npm install -g atlascloud-cli
# curl -fsSL https://raw.githubusercontent.com/AtlasCloudAI/cli/main/install.sh | sh

atlas auth login
atlas generate video alibaba/wan-2.7/reference-to-video -p "Your prompt here"
```

CLI docs: https://www.atlascloud.ai/docs/cli

## HTTP API reference

- **Submit endpoint (POST)**: `https://api.atlascloud.ai/api/v1/model/generateVideo` — start an async generation; returns a `prediction_id`
- **Poll endpoint (GET)**: `https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}` — poll this until the prediction finishes
- **Model ID**: `alibaba/wan-2.7/reference-to-video`


## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
See the input and output schema below, as well as the usage examples.


### Input Schema

The API accepts the following input parameters:

- **`model`** (`string`, _required_):
  Model name.
  - Default: `"alibaba/wan-2.7/reference-to-video"`
  - Options: "alibaba/wan-2.7/reference-to-video"

- **`prompt`** (`string`, _required_):
  Text prompt describing the desired video. Use labels like "character1" and "character2" to map reference materials to characters. Maximum length is 5000 characters.

- **`negative_prompt`** (`string`, _optional_):
  Text describing elements to exclude from the video. Maximum length is 500 characters.

- **`images`** (`array[string]`, _optional_):
  Reference image URLs. Each image represents one character or subject (person, animal, object). Up to 1 image per subject. Supported formats: JPEG, JPG, PNG (transparent channel not supported), BMP, WEBP. Width and height must each be between 240 and 8000 pixels. Aspect ratio must be between 1:8 and 8:1. Maximum file size: 20 MB per image.

- **`videos`** (`array[string]`, _optional_):
  Reference video URLs. Each video represents one character or subject and can also carry its voice. Up to 3 videos. Supported formats: mp4, mov. Duration: 1–30 seconds. Maximum file size: 100 MB per video.

- **`audio`** (`string`, _optional_):
  Audio URL for voice cloning. The model uses this audio as the voice for the character in the reference material. Supported formats: wav, mp3. Duration: 1–10 seconds. Maximum file size: 15 MB. Same role as `reference_voice`; `audio` is preferred for consistency with other Wan video models.

- **`resolution`** (`string`, _optional_):
  Output video resolution. Higher resolution increases cost.
  - Default: `"1080P"`
  - Options: "720P", "1080P"

- **`ratio`** (`string`, _optional_):
  Aspect ratio of the generated video.
  - Default: `"16:9"`
  - Options: "16:9", "9:16", "1:1", "4:3", "3:4"

- **`duration`** (`integer`, _optional_):
  Video duration in seconds. Longer duration increases cost.
  - Default: `5`
  - Min: 2
  - Max: 10

- **`prompt_extend`** (`boolean`, _optional_):
  Whether to use AI to enhance the prompt for better video quality. Increases generation time.
  - Default: `false`

- **`seed`** (`integer`, _optional_):
  Random seed for video generation. Range: 0 to 2147483647. Use -1 for a random seed.
  - Default: `-1`
  - Min: -1
  - Max: 2147483647



**Required Parameters Example**:

```json
{
  "model": "alibaba/wan-2.7/reference-to-video",
  "prompt": ""
}
```


**Full Example**:

```json
{
  "model": "alibaba/wan-2.7/reference-to-video",
  "prompt": "",
  "negative_prompt": "",
  "images": [
    ""
  ],
  "videos": [
    ""
  ],
  "audio": "",
  "resolution": "1080P",
  "ratio": "16:9",
  "duration": 5,
  "prompt_extend": false,
  "seed": -1
}
```


### Output Schema

The API returns the following output format:


- **`code`** (`integer`, _optional_):

- **`message`** (`string`, _optional_):

- **`data`** (`string`, _optional_):



**Example Response**:

```json
{
  "code": 0,
  "message": "",
  "data": null
}
```


## Usage Examples

### cURL

```bash
# Step 1: Start generation (async)
curl -X POST "https://api.atlascloud.ai/api/v1/model/generateVideo" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "alibaba/wan-2.7/reference-to-video",
  "prompt": "",
  "negative_prompt": "",
  "images": [
    ""
  ],
  "videos": [
    ""
  ],
  "audio": "",
  "resolution": "1080P",
  "ratio": "16:9",
  "duration": 5,
  "prompt_extend": false,
  "seed": -1
}'

# Response will contain: {"code": 200, "data": {"id": "prediction_id", "status": "processing"}}

# Step 2: Poll for result (replace {prediction_id} with the id returned above)
curl -X GET "https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY"

# Keep polling until status is "completed", "succeeded" or "failed"
# When completed, outputs will contain the generated content URL(s)
```

## Additional Resources

### Documentation

- [Model Playground](https://www.atlascloud.ai/models/alibaba/wan-2.7/reference-to-video)

## About this model

_Vendor-supplied description. Any pricing or endpoint mentioned below refers to_
_other platforms — use the Atlas Cloud values above._

#### Alibaba WAN 2.7 Reference-to-Video

**Alibaba WAN 2.7 Reference-to-Video** generates character-driven videos from reference images and videos, supporting multi-subject scenes and voice cloning.

#### What makes it stand out?

- **Character consistency:** Provide reference images or videos of characters, and the model preserves their appearance across the generated video.
- **Multi-subject scenes:** Include up to 5 reference materials (images + videos combined) to create scenes with multiple characters interacting.
- **Voice cloning:** Attach a voice reference audio to transfer a character's voice into the generated video.
- **Flexible framing:** Five aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4) at 720P or 1080P.

#### Designed For

- Creators building character-driven stories that need consistent character identity across clips.
- Teams producing multi-character interaction videos from a set of reference assets.
- Anyone who wants to animate a character from a photo or short clip with a specific voice.

#### How to Use

1. Write a prompt using labels like "character1" and "character2" to reference each subject.
2. Provide reference materials in order: the first image or video maps to character1, the second to character2, and so on.
3. Each reference should contain only one subject (person, animal, or object).
4. Optionally add a `reference_voice` audio URL to give a character a specific voice.
5. Set resolution, ratio, and duration for the output video.

#### Super Resolution

Use `1080P-SR` or `1440P-SR` when you want sharper detail than the native output. The request first renders a 720P source video, then applies FlashVSR super-resolution to the requested target tier. `1080P-SR` is intended as a lower-cost HD option, while `1440P-SR` is intended for larger screens or publishing workflows.

_(Description truncated. Full text on the model page.)_

---

Atlas Cloud — one API for 400+ AI models. Model page: https://www.atlascloud.ai/models/alibaba/wan-2.7/reference-to-video · Docs: https://www.atlascloud.ai/docs
