# Grok Imagine Video v1.5 Reference-to-Video — Atlas Cloud API

> xAI Grok Imagine Video v1.5 generates video guided by 1-7 reference images plus an optional reference voice, with native synchronized audio. Up to 15s at 480p or 720p.

This is the machine-readable API reference for **Grok Imagine Video v1.5 Reference-to-Video** on Atlas Cloud,
a unified API platform for 400+ AI models across text, image, video, audio and 3D.

- **Model ID**: `xai/grok-imagine-video-v1.5/reference-to-video`
- **Built by**: xAI
- **Modality**: Video
- **Model page**: https://www.atlascloud.ai/models/xai/grok-imagine-video-v1.5/reference-to-video
- **API key**: https://www.atlascloud.ai/console/api-keys
- **Docs**: https://www.atlascloud.ai/docs

## Pricing on Atlas Cloud

- $0.08 per second of generated video
- Pay-as-you-go. No minimum spend, no subscription required.

> **These are the authoritative Atlas Cloud rates for this model.** Any price that
> appears in the vendor description further down refers to a different platform or
> a different model variant and does not apply here.

## Use this model from an AI agent

Atlas Cloud ships three first-party integration surfaces. All three authenticate
with the same API key via the `ATLASCLOUD_API_KEY` environment variable.

### MCP server

The official MCP server (`atlascloud-mcp`) exposes this model to any
MCP-compatible host — Claude Code, OpenAI Codex, Cursor, Gemini CLI, Goose,
Claude Desktop. One-line install:

```bash
# Claude Code
claude mcp add atlascloud -- npx -y atlascloud-mcp

# OpenAI Codex CLI
codex mcp add atlascloud -- npx -y atlascloud-mcp

# Gemini CLI
gemini mcp add atlascloud -- npx -y atlascloud-mcp

export ATLASCLOUD_API_KEY="your-api-key"
```

Then ask in plain English; the agent calls `atlas_generate_video` with `model: "xai/grok-imagine-video-v1.5/reference-to-video"`.
The server fetches each model's schema and validates parameters before submitting,
so invalid requests fail fast without spending credits.

MCP docs: https://www.atlascloud.ai/docs/mcp-server

### Agent Skills

`atlas-cloud-skills` is a portable skill package (API reference, code templates in
Python / Node.js / cURL, model IDs with pricing) for Claude Code, Cursor, Codex and
12+ other agents:

```bash
npx skills add AtlasCloudAI/atlas-cloud-skills
export ATLASCLOUD_API_KEY="your-api-key"
```

Skills docs: https://www.atlascloud.ai/docs/skills

### CLI

The `atlas` binary runs Atlas Cloud from a terminal or CI script. Async media jobs
are polled and downloaded automatically (use `--no-download` when a script only
needs the output URLs):

```bash
# Install (Homebrew, npm, or shell installer)
brew install AtlasCloudAI/tap/atlascloud
# npm install -g atlascloud-cli
# curl -fsSL https://raw.githubusercontent.com/AtlasCloudAI/cli/main/install.sh | sh

atlas auth login
atlas generate video xai/grok-imagine-video-v1.5/reference-to-video -p "Your prompt here"
```

CLI docs: https://www.atlascloud.ai/docs/cli

## HTTP API reference

- **Submit endpoint (POST)**: `https://api.atlascloud.ai/api/v1/model/generateVideo` — start an async generation; returns a `prediction_id`
- **Poll endpoint (GET)**: `https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}` — poll this until the prediction finishes
- **Model ID**: `xai/grok-imagine-video-v1.5/reference-to-video`


## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
See the input and output schema below, as well as the usage examples.


### Input Schema

The API accepts the following input parameters:

- **`model`** (`string`, _required_):
  Model name.
  - Default: `"xai/grok-imagine-video-v1.5/reference-to-video"`

- **`prompt`** (`string`, _required_):
  Natural-language video description. Reference <IMAGE_N> tags to attach specific reference images to scene elements.

- **`image_urls`** (`array[string]`, _required_):
  1–7 reference images contributing people, objects, or styles to the generated video. Each is a public HTTPS URL or base64 data URI. Referenced in the prompt as <IMAGE_0> … <IMAGE_N>.
  - Min items: 1
  - Max items: 7

- **`voice_ids`** (`array[string]`, _optional_):
  Optional xAI preset voices supplying the speaking voice(s) for generated dialogue (up to 3). Preset identifiers only — custom or uploaded audio is not supported. Choices: altair (elegant, refined, and effortlessly premium); ara (warm and friendly); atlas (confident, commanding, and reassuring); carina (soft, empathetic, and soothing); castor (charismatic, down-to-earth, and easygoing); celeste (compassionate, confident, and reassuring); cosmo (bright, curious, and easy to follow); eve (energetic and upbeat); helios (upbeat, energetic, and endlessly versatile); helix (bold, dynamic, and adrenaline-fueled); iris (friendly, upbeat, and naturally charming); kepler (inventive, forward-thinking, and charismatic); leo (authoritative and strong); lumen (warm, articulate, and engaging); luna (gentle, patient, and deeply nurturing); lux (grounded, calm, and quietly wise); naksh (warm, thoughtful, and wise); orion (rich, cinematic, and resonant); perseus (strong, confident, and trustworthy); rex (confident and clear); rigel (precise, professional, and calmly confident); sal (smooth and balanced); sirius (quick-witted, clever, and playful); ursa (friendly, warm, and steadfast); zagan (powerful, dramatic, and unmistakable); zenith (sharp, focused, and driven).
  - Min items: 1
  - Max items: 3

- **`duration`** (`integer`, _optional_):
  Length of generated video in seconds. Range: 1–15.
  - Default: `8`
  - Min: 1
  - Max: 15

- **`resolution`** (`string`, _optional_):
  Output resolution. Reference-to-video is capped at 720p.
  - Default: `"720p"`
  - Options: "480p", "720p"

- **`aspect_ratio`** (`string`, _optional_):
  Output aspect ratio.
  - Default: `"16:9"`
  - Options: "1:1", "16:9", "9:16", "4:3", "3:4", "3:2", "2:3"



**Required Parameters Example**:

```json
{
  "model": "xai/grok-imagine-video-v1.5/reference-to-video",
  "prompt": "",
  "image_urls": [
    ""
  ]
}
```


**Full Example**:

```json
{
  "model": "xai/grok-imagine-video-v1.5/reference-to-video",
  "prompt": "",
  "image_urls": [
    ""
  ],
  "voice_ids": [
    ""
  ],
  "duration": 8,
  "resolution": "720p",
  "aspect_ratio": "16:9"
}
```


### Output Schema

The API returns the following output format:


- **`id`** (`string`, _optional_):
  Unique identifier for the prediction.

- **`urls`** (`object`, _optional_):
  Object containing related API endpoints.

- **`model`** (`string`, _optional_):
  Model ID used for the prediction.

- **`status`** (`string`, _optional_):
  Status of the task: created, processing, completed, or failed.

- **`outputs`** (`array[string]`, _optional_):
  Array of URLs to the generated video (empty when status is not completed).

- **`created_at`** (`string`, _optional_):
  ISO timestamp of when the request was created.



**Example Response**:

```json
{
  "id": "",
  "urls": {},
  "model": "",
  "status": "",
  "outputs": [
    ""
  ],
  "created_at": ""
}
```


## Usage Examples

### cURL

```bash
# Step 1: Start generation (async)
curl -X POST "https://api.atlascloud.ai/api/v1/model/generateVideo" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "xai/grok-imagine-video-v1.5/reference-to-video",
  "prompt": "",
  "image_urls": [
    ""
  ],
  "voice_ids": [
    ""
  ],
  "duration": 8,
  "resolution": "720p",
  "aspect_ratio": "16:9"
}'

# Response will contain: {"code": 200, "data": {"id": "prediction_id", "status": "processing"}}

# Step 2: Poll for result (replace {prediction_id} with the id returned above)
curl -X GET "https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY"

# Keep polling until status is "completed", "succeeded" or "failed"
# When completed, outputs will contain the generated content URL(s)
```

## Additional Resources

### Documentation

- [Model Playground](https://www.atlascloud.ai/models/xai/grok-imagine-video-v1.5/reference-to-video)

## About this model

_Vendor-supplied description. Any pricing or endpoint mentioned below refers to_
_other platforms — use the Atlas Cloud values above._

#### 1. Introduction

**Grok Imagine Video V1.5** is a frontier-tier video generation model developed by xAI that produces short clips of up to 15 seconds with natively generated, synchronized audio — including dialogue, lip-sync, sound effects, and ambient music — in a single inference pass.

This README applies to the following API model identifier:

- `xai/grok-imagine-video-v1.5/reference-to-video`

Reference-to-video is the identity-preserving mode: you supply 1–7 reference images that contribute specific people, products, or visual styles. The prompt then directs how those references behave, using `<IMAGE_N>` tags to bind each reference to its role in the scene.

Built on xAI's Aurora engine — an autoregressive mixture-of-experts (MoE) network that jointly models text, image, video, and audio tokens — the model produces tightly coupled audiovisual output rather than dubbing audio in post-processing.

---

#### 2. Key Features

- **1–7 Reference Images**: Multiple references can be combined in a single generation — a person from one image, a product from another, a setting or style from a third. Each is addressed independently in the prompt.

- **Tag-Based Binding**: References are bound to scene roles explicitly through `<IMAGE_0>` … `<IMAGE_N>` tags in the prompt, so the model is told which reference plays which part rather than inferring it.

- **Native Synchronized Audio Generation**: Audio is generated jointly with video tokens in a single inference pass, producing event-aligned sound effects and natural lip-sync.

- **Granular Duration Control (1–15 seconds)**: Clips can be requested at any integer second from 1 to 15.

- **Broad Format Support**: Outputs H.264 MP4 at 24 FPS across seven aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3), at 480p or 720p.

---

#### 3. Using References

Reference images are passed in `image_urls` and are addressed **in array order**: the first entry is `<IMAGE_0>`, the second `<IMAGE_1>`, and so on.

_(Description truncated. Full text on the model page.)_

---

Atlas Cloud — one API for 400+ AI models. Model page: https://www.atlascloud.ai/models/xai/grok-imagine-video-v1.5/reference-to-video · Docs: https://www.atlascloud.ai/docs
