
Doubao‑Audio‑Generate‑1.0 is Doubao Voice’s next‑generation audio‑generation engine. The industry‑first commercial tool creates film‑grade audio with just one prompt. It eliminates cumbersome audio‑engineering work. Creators generate publish‑ready radio dramas, podcasts and branded audio easily, shifting from a simple voice‑generator to an AI audio director. It serves audiobooks, serialized episodes and commercial audio for high‑quality narrative‑driven production.

xAI TTS v1 is a high-fidelity text-to-speech model that converts text into natural, expressive speech with sub-second latency, supporting 20 languages and 80+ voices with fine-grained delivery control.

ElevenLabs v3 Text-to-Speech model. High-quality speech generation from text prompts.

Gemini text-to-speech (3.1 Flash): latest-generation expressive speech from text with 30 prebuilt voices and natural-language style control. Powered by gemini-3.1-flash-tts-preview.

MiniMax text-to-speech (HD): natural, high-fidelity speech from text with selectable preset voices, speed, volume and pitch.

MiniMax text-to-speech (Turbo): fast, low-latency speech from text with selectable preset voices, speed, volume and pitch.

Gemini text-to-speech: expressive, controllable speech from text with 30+ prebuilt voices (Kore, Puck, Charon, Aoede...). Powered by gemini-2.5-flash-preview-tts.

Gemini text-to-speech (Pro): studio-quality, expressive speech from text with 30 prebuilt voices and natural-language style control. Powered by gemini-2.5-pro-preview-tts.

Suno text-to-music (chirp): generate a full song from a text prompt (and optional lyrics). Async; returns 2 variations.

Suno text-to-music (chirp): generate a full song from a text prompt (and optional lyrics). Async; returns 2 variations.

Suno text-to-music (chirp): generate a full song from a text prompt (and optional lyrics). Async; returns 2 variations.

Suno text-to-music (chirp): generate a full song from a text prompt (and optional lyrics). Async; returns 2 variations.

Suno text-to-music (chirp): generate a full song from a text prompt (and optional lyrics). Async; returns 2 variations.

Suno text-to-music (chirp): generate a full song from a text prompt (and optional lyrics). Async; returns 2 variations.

Suno text-to-music (chirp): generate a full song from a text prompt (and optional lyrics). Async; returns 2 variations.

Suno text-to-music (chirp): generate a full song from a text prompt (and optional lyrics). Async; returns 2 variations.

MiniMax text-to-music (latest): generate a full vocal or instrumental song from a style prompt plus lyrics with [Verse]/[Chorus] structure tags. Synchronous single-call generation.
Join the Discord community for the latest model updates, prompts, and support.