
MiniMax Music 3.0 is MiniMax's 11.1B-parameter open-weights music model that turns a musical description and optional lyrics into a complete, fully arranged and mixed song of up to five minutes - vocals, instrumentation and production included - in a single generation, with section-tag control over the arrangement and vocal or instrumental output.

MiniMax Lyrics Generation is a dedicated lyric-writing model that turns a one-line theme into a complete, professionally structured set of song lyrics - title, style tags, and sections marked with [Verse]/[Chorus] structure tags - and can also edit, continue, or restructure existing lyrics, with output directly usable as the lyrics input of MiniMax's music models.

Doubao‑Audio‑Generate‑1.0 is Doubao Voice’s next‑generation audio‑generation engine. The industry‑first commercial tool creates film‑grade audio with just one prompt. It eliminates cumbersome audio‑engineering work. Creators generate publish‑ready radio dramas, podcasts and branded audio easily, shifting from a simple voice‑generator to an AI audio director. It serves audiobooks, serialized episodes and commercial audio for high‑quality narrative‑driven production.

MiniMax text-to-speech (Turbo): fast, low-latency speech from text with selectable preset voices, speed, volume and pitch.

MiniMax text-to-speech (HD): natural, high-fidelity speech from text with selectable preset voices, speed, volume and pitch.

xAI TTS v1 is a high-fidelity text-to-speech model that converts text into natural, expressive speech with sub-second latency, supporting 20 languages and 80+ voices with fine-grained delivery control.

MiniMax text-to-music (latest): generate a full vocal or instrumental song from a style prompt plus lyrics with [Verse]/[Chorus] structure tags. Synchronous single-call generation.

ElevenLabs v3 Text-to-Speech model. High-quality speech generation from text prompts.

Suno V6 text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics, with variety and target duration controls. Async; returns 2 tracks per generation.

Suno V6 Wild via APIMart: the experimental V6 variant with bolder, less conventional arrangements. Same modes and pricing as chirp-v6. Async; returns 2 tracks per generation.

Suno V6 Mini via APIMart: the lighter, faster V6 variant for drafts and short clips. Same modes and pricing as chirp-v6. Async; returns 2 tracks per generation.

Gemini text-to-speech (3.1 Flash): latest-generation expressive speech from text with 30 prebuilt voices and natural-language style control. Powered by gemini-3.1-flash-tts-preview.

Gemini text-to-speech: expressive, controllable speech from text with 30+ prebuilt voices (Kore, Puck, Charon, Aoede...). Powered by gemini-2.5-flash-preview-tts.

Gemini text-to-speech (Pro): studio-quality, expressive speech from text with 30 prebuilt voices and natural-language style control. Powered by gemini-2.5-pro-preview-tts.