# Seed Audio 1.0 — Atlas Cloud API

> Doubao‑Audio‑Generate‑1.0 is Doubao Voice’s next‑generation audio‑generation engine. The industry‑first commercial tool creates film‑grade audio with just one prompt. It eliminates cumbersome audio‑engineering work. Creators generate publish‑ready radio dramas, podcasts and branded audio easily, shifting from a simple voice‑generator to an AI audio director. It serves audiobooks, serialized episodes and commercial audio for high‑quality narrative‑driven production.

This is the machine-readable API reference for **Seed Audio 1.0** on Atlas Cloud,
a unified API platform for 400+ AI models across text, image, video, audio and 3D.

- **Model ID**: `bytedance/seed-audio-1.0`
- **Built by**: ByteDance
- **Modality**: Audio
- **Model page**: https://www.atlascloud.ai/models/bytedance/seed-audio-1.0
- **API key**: https://www.atlascloud.ai/console/api-keys
- **Docs**: https://www.atlascloud.ai/docs

## Pricing on Atlas Cloud

- $0.143 per minute of audio
- Pay-as-you-go. No minimum spend, no subscription required.

> **These are the authoritative Atlas Cloud rates for this model.** Any price that
> appears in the vendor description further down refers to a different platform or
> a different model variant and does not apply here.

## Use this model from an AI agent

Atlas Cloud ships three first-party integration surfaces. All three authenticate
with the same API key via the `ATLASCLOUD_API_KEY` environment variable.

### MCP server

The official MCP server (`atlascloud-mcp`) exposes this model to any
MCP-compatible host — Claude Code, OpenAI Codex, Cursor, Gemini CLI, Goose,
Claude Desktop. One-line install:

```bash
# Claude Code
claude mcp add atlascloud -- npx -y atlascloud-mcp

# OpenAI Codex CLI
codex mcp add atlascloud -- npx -y atlascloud-mcp

# Gemini CLI
gemini mcp add atlascloud -- npx -y atlascloud-mcp

export ATLASCLOUD_API_KEY="your-api-key"
```

Then ask in plain English; the agent calls `atlas_generate_audio` with `model: "bytedance/seed-audio-1.0"`.
The server fetches each model's schema and validates parameters before submitting,
so invalid requests fail fast without spending credits.

MCP docs: https://www.atlascloud.ai/docs/mcp-server

### Agent Skills

`atlas-cloud-skills` is a portable skill package (API reference, code templates in
Python / Node.js / cURL, model IDs with pricing) for Claude Code, Cursor, Codex and
12+ other agents:

```bash
npx skills add AtlasCloudAI/atlas-cloud-skills
export ATLASCLOUD_API_KEY="your-api-key"
```

Skills docs: https://www.atlascloud.ai/docs/skills

### CLI

The `atlas` binary runs Atlas Cloud from a terminal or CI script. Async media jobs
are polled and downloaded automatically (use `--no-download` when a script only
needs the output URLs):

```bash
# Install (Homebrew, npm, or shell installer)
brew install AtlasCloudAI/tap/atlascloud
# npm install -g atlascloud-cli
# curl -fsSL https://raw.githubusercontent.com/AtlasCloudAI/cli/main/install.sh | sh

atlas auth login
atlas models get bytedance/seed-audio-1.0 --json
```

CLI docs: https://www.atlascloud.ai/docs/cli

## HTTP API reference

- **Submit endpoint (POST)**: `https://api.atlascloud.ai/api/v1/model/generateAudio` — start an async generation; returns a `prediction_id`
- **Poll endpoint (GET)**: `https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}` — poll this until the prediction finishes
- **Model ID**: `bytedance/seed-audio-1.0`


## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
See the input and output schema below, as well as the usage examples.


### Input Schema

The API accepts the following input parameters:

- **`model`** (`string`, _required_):
  Model name.
  - Default: `"bytedance/seed-audio-1.0"`
  - Options: "bytedance/seed-audio-1.0"

- **`text`** (`string`, _required_):
  Prompt or text to synthesize into audio. For audio-reference generation, cite the provider-specific audio reference placeholders supported by the upstream service.
  - Default: `"A phone vibrates first, then a calm male voice says: Welcome to Seed Audio."`

- **`references`** (`array[object]`, _optional_):
  Optional reference resources. Omit for text-only generation. Add up to 3 audio/speaker references or 1 image reference. Each item must choose exactly one source.
  - Max items: 3
  - Item properties:
    - **`speaker`** (`string`, _optional_):
      Use a preset Doubao TTS 2.0 voice ID or a cloned voice ID.
      - Options: "zh_female_vv_uranus_bigtts", "zh_female_xiaohe_uranus_bigtts", "zh_male_m191_uranus_bigtts", "zh_male_taocheng_uranus_bigtts", "zh_male_liufei_uranus_bigtts", "zh_female_sophie_uranus_bigtts", "zh_female_qingxinnvsheng_uranus_bigtts", "zh_female_cancan_uranus_bigtts", "zh_female_tianmeitaozi_uranus_bigtts", "zh_male_ruyayichen_uranus_bigtts"

    - **`audio_url`** (`string`, _optional_):
      Use a remote reference audio file URL.

    - **`audio_data`** (`string`, _optional_):
      Paste a Base64-encoded reference audio file.

    - **`image_url`** (`string`, _optional_):
      Use a remote reference image URL.

    - **`image_data`** (`string`, _optional_):
      Paste a Base64-encoded reference image file.


- **`format`** (`string`, _optional_):
  Output audio format.
  - Default: `"mp3"`
  - Options: "mp3", "wav", "pcm", "ogg_opus"

- **`sample_rate`** (`integer`, _optional_):
  Output sample rate.
  - Default: `24000`
  - Options: 8000, 16000, 24000, 32000, 44100, 48000

- **`pitch_rate`** (`integer`, _optional_):
  Pitch adjustment. Range [-12, 12], default 0.
  - Default: `0`
  - Min: -12
  - Max: 12

- **`speech_rate`** (`integer`, _optional_):
  Speech speed adjustment. Range [-50, 100], where 100 means 2.0x and -50 means 0.5x.
  - Default: `0`
  - Min: -50
  - Max: 100

- **`loudness_rate`** (`integer`, _optional_):
  Loudness adjustment. Range [-50, 100], where 100 means 2.0x loudness and -50 means 0.5x.
  - Default: `0`
  - Min: -50
  - Max: 100



**Required Parameters Example**:

```json
{
  "model": "bytedance/seed-audio-1.0",
  "text": "A phone vibrates first, then a calm male voice says: Welcome to Seed Audio."
}
```


**Full Example**:

```json
{
  "model": "bytedance/seed-audio-1.0",
  "text": "A phone vibrates first, then a calm male voice says: Welcome to Seed Audio.",
  "references": [
    {
      "speaker": "zh_female_vv_uranus_bigtts",
      "audio_url": "",
      "audio_data": "",
      "image_url": "",
      "image_data": ""
    }
  ],
  "format": "mp3",
  "sample_rate": 24000,
  "pitch_rate": 0,
  "speech_rate": 0,
  "loudness_rate": 0
}
```


### Output Schema

The API returns the following output format:


- **`created_at`** (`string`, _optional_):
  ISO timestamp of when the request was created.

- **`id`** (`string`, _optional_):
  Unique identifier for the prediction.

- **`model`** (`string`, _optional_):
  Model ID used for the prediction.

- **`outputs`** (`array[string]`, _optional_):
  Array of URLs to the generated audio.

- **`status`** (`string`, _optional_):
  Status of the task: created, processing, completed, or failed.

- **`urls`** (`object`, _optional_):
  Object containing related API endpoints.



**Example Response**:

```json
{
  "created_at": "",
  "id": "",
  "model": "",
  "outputs": [
    ""
  ],
  "status": "",
  "urls": {}
}
```


## Usage Examples

### cURL

```bash
# Step 1: Start generation (async)
curl -X POST "https://api.atlascloud.ai/api/v1/model/generateAudio" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "bytedance/seed-audio-1.0",
  "text": "A phone vibrates first, then a calm male voice says: Welcome to Seed Audio.",
  "references": [
    {
      "speaker": "zh_female_vv_uranus_bigtts",
      "audio_url": "",
      "audio_data": "",
      "image_url": "",
      "image_data": ""
    }
  ],
  "format": "mp3",
  "sample_rate": 24000,
  "pitch_rate": 0,
  "speech_rate": 0,
  "loudness_rate": 0
}'

# Response will contain: {"code": 200, "data": {"id": "prediction_id", "status": "processing"}}

# Step 2: Poll for result (replace {prediction_id} with the id returned above)
curl -X GET "https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY"

# Keep polling until status is "completed", "succeeded" or "failed"
# When completed, outputs will contain the generated content URL(s)
```

## Additional Resources

### Documentation

- [Model Playground](https://www.atlascloud.ai/models/bytedance/seed-audio-1.0)

## About this model

_Vendor-supplied description. Any pricing or endpoint mentioned below refers to_
_other platforms — use the Atlas Cloud values above._

### Seed Audio 1.0

Seed Audio 1.0 is ByteDance's audio generation model for producing speech from text prompts, with optional reference audio, speaker, or image inputs. It is exposed on AtlasCloud through the standard asynchronous audio generation API.

#### Highlights

- **Text-to-speech generation:** Convert text prompts into speech audio.
- **Reference audio control:** Provide up to three reference audios or a speaker ID to guide the voice, tone, or delivery. Refer to them in the prompt using the upstream placeholder tokens `@audio1`, `@audio2`, and `@audio3`.
- **Reference image control:** Provide one reference image to guide the generated audio style or character context.
- **Reference exclusivity:** Each reference item must contain exactly one of `speaker`, `audio_url`, `audio_data`, `image_url`, or `image_data`. The upstream API rejects mixed audio + image references in the same request.
- **Audio format control:** Generate `mp3`, `wav`, `pcm`, or `ogg_opus`.
- **Sample rate control:** Choose common output sample rates from `8000` to `48000`.
- **Speech controls:** Adjust pitch, speech speed, and loudness with optional rate parameters.

#### Parameters

| Parameter | Required | Description |
|-----------|----------|-------------|
| `model` | Yes | Use `bytedance/seed-audio-1.0`. |
| `text` | Yes | Text prompt to synthesize into speech. |
| `references` | No | Optional array of reference inputs. Use `audio_url`, `audio_data`, or `speaker` for audio/voice references; use `image_url` or `image_data` for image references. Do not mix image references with audio or speaker references. |
| `format` | No | Output audio format. Default: `mp3`. |
| `sample_rate` | No | Output sample rate. Default: `24000`. |
| `pitch_rate` | No | Pitch adjustment. Default: `0`. |
| `speech_rate` | No | Speech speed adjustment. Default: `0`. |
| `loudness_rate` | No | Loudness adjustment. Default: `0`. |

#### Example Request

_(Description truncated. Full text on the model page.)_

---

Atlas Cloud — one API for 400+ AI models. Model page: https://www.atlascloud.ai/models/bytedance/seed-audio-1.0 · Docs: https://www.atlascloud.ai/docs
