# xAI STT v1 — Atlas Cloud API

> xAI STT v1 is a production-grade speech-to-text model that transcribes audio into accurate, formatted text. It supports 24+ languages with automatic language detection, word-level timestamps, speaker diarization, multichannel transcription, and inverse text normalization.

This is the machine-readable API reference for **xAI STT v1** on Atlas Cloud,
a unified API platform for 400+ AI models across text, image, video, audio and 3D.

- **Model ID**: `xai/stt-v1`
- **Built by**: xAI
- **Modality**: Audio
- **Model page**: https://www.atlascloud.ai/models/xai/stt-v1
- **API key**: https://www.atlascloud.ai/console/api-keys
- **Docs**: https://www.atlascloud.ai/docs

## Pricing on Atlas Cloud

- $0.002 per minute of audio
- Pay-as-you-go. No minimum spend, no subscription required.

> **These are the authoritative Atlas Cloud rates for this model.** Any price that
> appears in the vendor description further down refers to a different platform or
> a different model variant and does not apply here.

## Use this model from an AI agent

Atlas Cloud ships three first-party integration surfaces. All three authenticate
with the same API key via the `ATLASCLOUD_API_KEY` environment variable.

### MCP server

The official MCP server (`atlascloud-mcp`) exposes this model to any
MCP-compatible host — Claude Code, OpenAI Codex, Cursor, Gemini CLI, Goose,
Claude Desktop. One-line install:

```bash
# Claude Code
claude mcp add atlascloud -- npx -y atlascloud-mcp

# OpenAI Codex CLI
codex mcp add atlascloud -- npx -y atlascloud-mcp

# Gemini CLI
gemini mcp add atlascloud -- npx -y atlascloud-mcp

export ATLASCLOUD_API_KEY="your-api-key"
```

Then ask in plain English; the agent calls `atlas_generate_audio` with `model: "xai/stt-v1"`.
The server fetches each model's schema and validates parameters before submitting,
so invalid requests fail fast without spending credits.

MCP docs: https://www.atlascloud.ai/docs/mcp-server

### Agent Skills

`atlas-cloud-skills` is a portable skill package (API reference, code templates in
Python / Node.js / cURL, model IDs with pricing) for Claude Code, Cursor, Codex and
12+ other agents:

```bash
npx skills add AtlasCloudAI/atlas-cloud-skills
export ATLASCLOUD_API_KEY="your-api-key"
```

Skills docs: https://www.atlascloud.ai/docs/skills

### CLI

The `atlas` binary runs Atlas Cloud from a terminal or CI script. Async media jobs
are polled and downloaded automatically (use `--no-download` when a script only
needs the output URLs):

```bash
# Install (Homebrew, npm, or shell installer)
brew install AtlasCloudAI/tap/atlascloud
# npm install -g atlascloud-cli
# curl -fsSL https://raw.githubusercontent.com/AtlasCloudAI/cli/main/install.sh | sh

atlas auth login
atlas models get xai/stt-v1 --json
```

CLI docs: https://www.atlascloud.ai/docs/cli

## HTTP API reference

- **Submit endpoint (POST)**: `https://api.atlascloud.ai/api/v1/model/generateAudio` — start an async generation; returns a `prediction_id`
- **Poll endpoint (GET)**: `https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}` — poll this until the prediction finishes
- **Model ID**: `xai/stt-v1`


## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
See the input and output schema below, as well as the usage examples.


### Input Schema

The API accepts the following input parameters:

- **`model`** (`string`, _required_):
  model name
  - Default: `"xai/stt-v1"`

- **`audio`** (`string`, _required_):
  The audio to transcribe. Accepts a publicly accessible URL of an audio file (max 500 MB) or a base64-encoded audio string. Container formats (e.g. mp3, wav, m4a, flac, ogg, opus, aac, mp4, mkv) are auto-detected; for raw/headerless audio (pcm, mulaw, alaw) set audio_format and sample_rate, support mono, stereo, or up to 8 channels (with multichannel=true).
  - Default: `""`

- **`text_normalization`** (`boolean`, _optional_):
  When true, enables Inverse Text Normalization — converts spoken-form numbers/currency into written form (e.g. "one hundred dollars" to "$100"). This controls text formatting of the transcript, not the audio format.
  - Default: `false`

- **`language`** (`string`, _optional_):
  ISO 639-1 language code (Filipino uses 'fil'). The model transcribes speech in any supported language regardless of this parameter — setting it, together with text_normalization=true, enables formatting of numbers, currencies, and units into their written form for that language. Leave unset for automatic language detection.
  - Default: `""`
  - Options: "ar", "cs", "da", "nl", "en", "fil", "fr", "de", "hi", "id", "it", "ja", "ko", "mk", "ms", "fa", "pl", "pt", "ro", "ru", "es", "sv", "th", "tr", "vi"

- **`keyterm`** (`array[string]`, _optional_):
  Key terms to bias transcription toward (e.g. product names). Up to 100 terms, each up to 50 characters.
  - Default: `[]`
  - Max items: 100

- **`diarize`** (`boolean`, _optional_):
  When true, enables speaker diarization. Each word in the result includes a speaker_id field identifying the detected speaker.
  - Default: `false`

- **`filler_words`** (`boolean`, _optional_):
  When true, filler words (e.g. "um", "uh") are included in the transcription; when false (default), they are automatically removed.
  - Default: `false`

- **`audio_format`** (`string`, _optional_):
  Format hint for raw/headerless audio. Auto-detected for container formats (such as mp3, wav, etc), set the value to 'auto'.
  - Default: `"auto"`
  - Options: "auto", "pcm", "mulaw", "alaw"

- **`sample_rate`** (`integer`, _optional_):
  Sample rate in Hz. Only required for raw audio (pcm, mulaw, alaw).
  - Options: 8000, 16000, 22050, 24000, 44100, 48000

- **`multichannel`** (`boolean`, _optional_):
  When true, transcribes each audio channel independently.
  - Default: `false`

- **`channels`** (`integer`, _optional_):
  Number of audio channels (2–8). Only required for multichannel raw audio. Auto-detected for container formats.
  - Min: 2
  - Max: 8

- **`enable_sync_mode`** (`boolean`, _optional_):
  If set to true, the function will wait for the result to be generated and uploaded before returning the response. It allows you to get the result directly in the response. This property is only available through the API.
  - Default: `false`



**Required Parameters Example**:

```json
{
  "model": "xai/stt-v1",
  "audio": ""
}
```


**Full Example**:

```json
{
  "model": "xai/stt-v1",
  "audio": "",
  "text_normalization": false,
  "language": "",
  "keyterm": [],
  "diarize": false,
  "filler_words": false,
  "audio_format": "auto",
  "sample_rate": 0,
  "multichannel": false,
  "channels": 2,
  "enable_sync_mode": false
}
```


### Output Schema

The API returns the following output format:


- **`code`** (`integer`, _optional_):
  HTTP status code of the response.

- **`message`** (`string`, _optional_):
  Human-readable message; non-empty on failure.

- **`data`** (`object`, _optional_):
  - Properties:
    - **`id`** (`string`, _optional_):
      Unique identifier for the prediction.

    - **`model`** (`string`, _optional_):
      Model ID used for the prediction.

    - **`outputs`** (`array[string]`, _optional_):
      Array of URLs to the generated content. Null when status is not completed.

    - **`urls`** (`object`, _optional_):
      Object containing related API endpoints.
      - Properties:
        - **`get`** (`string`, _optional_):
          URL to poll for the prediction result.


    - **`status`** (`string`, _optional_):
      Status of the task: created, processing, completed, timeout, or failed.

    - **`created_at`** (`string`, _optional_):
      ISO timestamp of when the request was created (e.g., "2023-04-01T12:34:56.789Z").

    - **`error`** (`string`, _optional_):
      Error message if the task failed, empty string otherwise.

    - **`error_code`** (`integer`, _optional_):
      Error code if the task failed.

    - **`executionTime`** (`number`, _optional_):
      Total execution time in milliseconds.

    - **`timings`** (`object`, _optional_):
      Detailed timing breakdown.
      - Properties:
        - **`inference`** (`number`, _optional_):
          Inference time in milliseconds.


    - **`stt_result`** (`object`, _optional_):
      The speech-to-text transcription result.




**Example Response**:

```json
{
  "code": 0,
  "message": "",
  "data": {
    "id": "",
    "model": "",
    "outputs": [
      ""
    ],
    "urls": {
      "get": ""
    },
    "status": "",
    "created_at": "",
    "error": "",
    "error_code": 0,
    "executionTime": 0,
    "timings": {
      "inference": 0
    },
    "stt_result": {}
  }
}
```


## Usage Examples

### cURL

```bash
# Step 1: Start generation (async)
curl -X POST "https://api.atlascloud.ai/api/v1/model/generateAudio" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "xai/stt-v1",
  "audio": "",
  "text_normalization": false,
  "language": "",
  "keyterm": [],
  "diarize": false,
  "filler_words": false,
  "audio_format": "auto",
  "sample_rate": 0,
  "multichannel": false,
  "channels": 2,
  "enable_sync_mode": false
}'

# Response will contain: {"code": 200, "data": {"id": "prediction_id", "status": "processing"}}

# Step 2: Poll for result (replace {prediction_id} with the id returned above)
curl -X GET "https://api.atlascloud.ai/api/v1/model/prediction/{prediction_id}" \
  -H "Authorization: Bearer $ATLASCLOUD_API_KEY"

# Keep polling until status is "completed", "succeeded" or "failed"
# When completed, outputs will contain the generated content URL(s)
```

## Additional Resources

### Documentation

- [Model Playground](https://www.atlascloud.ai/models/xai/stt-v1)

## About this model

_Vendor-supplied description. Any pricing or endpoint mentioned below refers to_
_other platforms — use the Atlas Cloud values above._

### xAI STT v1 — Speech to Text

**Developer:** xAI
**Model ID:** `xai/stt-v1`
**Release Date:** April 2026

#### Overview

xAI STT v1 is a production-grade speech-to-text model from xAI, the company behind Grok. It transcribes audio into accurate, formatted text either in a single batch API call or in real time over a WebSocket stream. The model supports 24+ languages with automatic language detection, word-level timestamps, speaker diarization, multichannel transcription, and Inverse Text Normalization that renders numbers, currencies, and units in their written form.

xAI STT v1 is built on the same audio infrastructure that powers Grok Voice, Tesla in-vehicle assistants, and Starlink customer support — infrastructure that has been battle-tested at scale across consumer and enterprise workloads. It is engineered for demanding real-world use cases such as call-center analytics, voice agents, meeting transcription, and media captioning, where named-entity accuracy (names, account numbers, dates) and low latency matter most.

In xAI's published benchmarks the model reports a **6.9% word error rate** on general audio and a **5.0% error rate** on phone-call entity recognition — ahead of ElevenLabs (12.0%), Deepgram (13.5%), and AssemblyAI (21.3%) on the same entity task — while matching the leading providers (≈2.4% error) on clean video/podcast audio.

#### Key Capabilities

_(Description truncated. Full text on the model page.)_

---

Atlas Cloud — one API for 400+ AI models. Model page: https://www.atlascloud.ai/models/xai/stt-v1 · Docs: https://www.atlascloud.ai/docs
