# Nvidia Cosmos 3 Super Image-to-Video — Atlas Cloud API

> Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.

This is the machine-readable API reference for **Nvidia Cosmos 3 Super Image-to-Video** on Atlas Cloud,
a unified API platform for 400+ AI models across text, image, video, audio and 3D.

- **Model ID**: `nvidia/cosmos-3-super/image-to-video`
- **Modality**: Video
- **Model page**: https://www.atlascloud.ai/models/nvidia/cosmos-3-super/image-to-video
- **API key**: https://www.atlascloud.ai/console/api-keys
- **Docs**: https://www.atlascloud.ai/docs

## Pricing on Atlas Cloud

- $0.055 per second of generated video
- Pay-as-you-go. No minimum spend, no subscription required.

> **These are the authoritative Atlas Cloud rates for this model.** Any price that
> appears in the vendor description further down refers to a different platform or
> a different model variant and does not apply here.

## Use this model from an AI agent

Atlas Cloud ships three first-party integration surfaces. All three authenticate
with the same API key via the `ATLASCLOUD_API_KEY` environment variable.

### MCP server

The official MCP server (`atlascloud-mcp`) exposes this model to any
MCP-compatible host — Claude Code, OpenAI Codex, Cursor, Gemini CLI, Goose,
Claude Desktop. One-line install:

```bash
# Claude Code
claude mcp add atlascloud -- npx -y atlascloud-mcp

# OpenAI Codex CLI
codex mcp add atlascloud -- npx -y atlascloud-mcp

# Gemini CLI
gemini mcp add atlascloud -- npx -y atlascloud-mcp

export ATLASCLOUD_API_KEY="your-api-key"
```

Then ask in plain English; the agent calls `atlas_generate_video` with `model: "nvidia/cosmos-3-super/image-to-video"`.
The server fetches each model's schema and validates parameters before submitting,
so invalid requests fail fast without spending credits.

MCP docs: https://www.atlascloud.ai/docs/mcp-server

### Agent Skills

`atlas-cloud-skills` is a portable skill package (API reference, code templates in
Python / Node.js / cURL, model IDs with pricing) for Claude Code, Cursor, Codex and
12+ other agents:

```bash
npx skills add AtlasCloudAI/atlas-cloud-skills
export ATLASCLOUD_API_KEY="your-api-key"
```

Skills docs: https://www.atlascloud.ai/docs/skills

### CLI

The `atlas` binary runs Atlas Cloud from a terminal or CI script. Async media jobs
are polled and downloaded automatically (use `--no-download` when a script only
needs the output URLs):

```bash
# Install (Homebrew, npm, or shell installer)
brew install AtlasCloudAI/tap/atlascloud
# npm install -g atlascloud-cli
# curl -fsSL https://raw.githubusercontent.com/AtlasCloudAI/cli/main/install.sh | sh

atlas auth login
atlas generate video nvidia/cosmos-3-super/image-to-video -p "Your prompt here"
```

CLI docs: https://www.atlascloud.ai/docs/cli

## HTTP API reference

- **Submit endpoint (POST)**: `https://api.atlascloud.ai/api/v1/model/generateVideo` — start an async generation; returns a `prediction_id`
- **Poll endpoint (GET)**: `https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}` — poll this until the prediction finishes
- **Model ID**: `nvidia/cosmos-3-super/image-to-video`


## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
See the input and output schema below, as well as the usage examples.


### Input Schema

The API accepts the following input parameters:

- **`model`** (`string`, _required_):
  model name
  - Default: `"nvidia/cosmos-3-super/image-to-video"`

- **`prompt`** (`string`, _required_):
  Text description of the motion and scene to generate. Maximum 2,500 characters.

- **`image_url`** (`string`, _required_):
  The reference image to animate. Accepts a public URL or Base64-encoded image. Supported formats: png, jpeg, jpg, webp.

- **`negative_prompt`** (`string`, _optional_):
  Text describing content to exclude from the generated video.
  - Default: `""`

- **`image_size`** (`string`, _optional_):
  Output image resolution. Use a preset size or specify custom width and height via an ImageSize object.
  - Default: `"landscape_16_9"`

- **`duration`** (`integer`, _optional_):
  The duration of the generated media in seconds (1-7).
  - Default: `5`
  - Options: 1, 2, 3, 4, 5, 6, 7

- **`num_inference_steps`** (`integer`, _optional_):
  Number of denoising steps. Higher values improve quality but increase generation time.
  - Default: `28`
  - Min: 1
  - Max: 50

- **`guidance_scale`** (`number`, _optional_):
  Classifier-free guidance scale. Higher values make the output follow the prompt more closely.
  - Default: `6`
  - Min: 1
  - Max: 20

- **`seed`** (`integer`, _optional_):
  Random seed for reproducibility. Use the same seed and prompt to reproduce results.
  - Min: 0



**Required Parameters Example**:

```json
{
  "model": "nvidia/cosmos-3-super/image-to-video",
  "prompt": "",
  "image_url": ""
}
```


**Full Example**:

```json
{
  "model": "nvidia/cosmos-3-super/image-to-video",
  "prompt": "",
  "image_url": "",
  "negative_prompt": "",
  "image_size": "landscape_16_9",
  "duration": 5,
  "num_inference_steps": 28,
  "guidance_scale": 6,
  "seed": 0
}
```


### Output Schema

The API returns the following output format:


- **`model`** (`string`, _optional_):
  model name
  - Default: `"nvidia/cosmos-3-super/image-to-video"`

- **`created_at`** (`string`, _optional_):
  ISO timestamp of when the request was created (e.g., "2023-04-01T12:34:56.789Z").

- **`id`** (`string`, _optional_):
  Unique identifier for the prediction.

- **`outputs`** (`array[string]`, _optional_):
  Array of URLs to the generated videos. Null when status is not completed.

- **`status`** (`string`, _optional_):
  Status of the task: created, processing, completed, or failed.

- **`urls`** (`object`, _optional_):
  Object containing related API endpoints.



**Example Response**:

```json
{
  "model": "nvidia/cosmos-3-super/image-to-video",
  "created_at": "",
  "id": "",
  "outputs": [
    ""
  ],
  "status": "",
  "urls": {}
}
```


## Usage Examples

### cURL

```bash
# Step 1: Start generation (async)
curl -X POST "https://api.atlascloud.ai/api/v1/model/generateVideo" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "nvidia/cosmos-3-super/image-to-video",
  "prompt": "",
  "image_url": "",
  "negative_prompt": "",
  "image_size": "landscape_16_9",
  "duration": 5,
  "num_inference_steps": 28,
  "guidance_scale": 6,
  "seed": 0
}'

# Response will contain: {"code": 200, "data": {"id": "prediction_id", "status": "processing"}}

# Step 2: Poll for result (replace {prediction_id} with the id returned above)
curl -X GET "https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY"

# Keep polling until status is "completed", "succeeded" or "failed"
# When completed, outputs will contain the generated content URL(s)
```

## Additional Resources

### Documentation

- [Model Playground](https://www.atlascloud.ai/models/nvidia/cosmos-3-super/image-to-video)

## About this model

_Vendor-supplied description. Any pricing or endpoint mentioned below refers to_
_other platforms — use the Atlas Cloud values above._

### NVIDIA Cosmos 3 Super Image-to-Video

**NVIDIA Cosmos 3 Super Image-to-Video** is NVIDIA's flagship image animation model built on the Cosmos 3 architecture. It transforms a single still image into a smooth, physically grounded video clip guided by a text prompt — delivering cinematic motion, accurate temporal coherence, and fine-grained control over frame count and pacing.

#### 🌟 Key Features

##### 🎬 Image-Driven Animation
Animates any still image into a fluid video sequence, preserving the original composition, subject identity, and visual style of the reference image.

##### 📝 Prompt-Guided Motion
Accepts a text prompt to direct the type, direction, and style of motion — from subtle ambient movement to dynamic camera trajectories and subject actions.

##### 🚫 Negative Prompting
Accepts a `negative_prompt` to explicitly suppress unwanted motion artifacts, distortions, or visual elements in the output video.

##### ⏱ Flexible Duration Control
Set the output duration in whole seconds (1–7). Each second renders at 24 fps — pay only for the duration you generate.

##### 🔁 Reproducible Results
Pin a seed value to reproduce or iterate on a specific generation, enabling systematic prompt and parameter refinement.

#### ⚙️ Parameters

_(Description truncated. Full text on the model page.)_

---

Atlas Cloud — one API for 400+ AI models. Model page: https://www.atlascloud.ai/models/nvidia/cosmos-3-super/image-to-video · Docs: https://www.atlascloud.ai/docs
