рджреБрдирд┐рдпрд╛ рднрд░ рдореЗрдВ рд╕рдмрд╕реЗ рдХрдо рдХреАрдорддреЛрдВ рдкрд░ Seedance 2.0 Mini & Fast API тАФ рдЖрдзрд┐рдХрд╛рд░рд┐рдХ рдХреАрдордд рдкрд░ 68% рддрдХ рдХреА рдЫреВрдЯ
Atlas Cloud AI рдХреНрд░рд┐рдПрд╢рди рд╕реНрдЯреВрдбрд┐рдпреЛрдЕрдкрдиреЗ рднреАрддрд░ рдХреЗ рдирд┐рд░реНрджреЗрд╢рдХ рдХреЛ рдкрд╣рдЪрд╛рдиреЗрдВредрдмрдирд╛рдирд╛ рд╢реБрд░реВ рдХрд░реЗрдВ
Grok Imagine Video v1.5 Text-to-Video
рдЯреЗрдХреНрд╕реНрдЯ-рд╕реЗ-рд╡реАрдбрд┐рдпреЛ

Grok Imagine Video v1.5 Text-to-Video API by xAI

xai/grok-imagine-video-v1.5/text-to-video
Text-to-video

xAI Grok Imagine Video v1.5 generates video with native synchronized audio from a text prompt alone. Up to 15s at 480p, 720p, or 1080p.

рдореЙрдбрд▓ рдХреА рддреБрд▓рдирд╛ рдХрд░реЗрдВ

Grok Imagine Video v1.5 Text-to-Video рдХреЛ xAI рджреНрд╡рд╛рд░рд╛ рд╡рд┐рдХрд╕рд┐рдд рдХрд┐рдпрд╛ рдЧрдпрд╛ рд╣реИред Atlas Cloud (Atlas Cloud AI LLC рджреНрд╡рд╛рд░рд╛ рд╕рдВрдЪрд╛рд▓рд┐рдд) рдЗрд╕ рддрдХ рдкрд╣реБрдБрдЪ рдкреНрд░рджрд╛рди рдХрд░рддрд╛ рд╣реИ, рдЗрд╕рдХрд╛ рд╕реНрд╡рд╛рдореА рдирд╣реАрдВ рд╣реИред рд╕рднреА рдЯреНрд░реЗрдбрдорд╛рд░реНрдХ рдЙрдирдХреЗ рд╕рдВрдмрдВрдзрд┐рдд рд╕реНрд╡рд╛рдорд┐рдпреЛрдВ рдХреА рд╕рдВрдкрддреНрддрд┐ рд╣реИрдВред

1. Introduction

Grok Imagine Video V1.5 is a frontier-tier video generation model developed by xAI that produces short clips of up to 15 seconds with natively generated, synchronized audio тАФ including dialogue, lip-sync, sound effects, and ambient music тАФ in a single inference pass.

This README applies to the following API model identifier:

  • xai/grok-imagine-video-v1.5/text-to-video

Text-to-video is the prompt-only mode: no input image or reference material is required. The model composes subject, motion, camera work, and the full audio bed from the prompt alone, and renders natively at up to 1080p.

Built on xAI's Aurora engine тАФ an autoregressive mixture-of-experts (MoE) network that jointly models text, image, video, and audio tokens тАФ the model represents a departure from the diffusion-transformer paradigm used by Sora and Veo, enabling tightly coupled audiovisual generation with competitive cost and latency characteristics.


2. Key Features

  • Prompt-Only Generation: No source image or reference asset is needed. Scene, motion, camera behaviour, and audio are all derived from the text prompt, making this the fastest path from idea to finished clip.

  • Native Synchronized Audio Generation: Audio (dialogue, lip-sync, SFX, ambient sound, music) is generated jointly with video tokens in a single inference pass rather than dubbed in post-processing. This produces event-aligned sound effects and natural lip-sync without requiring separate audio pipelines.

  • Native 1080p Output: Text-to-video renders at 480p, 720p, or 1080p, with 1080p produced natively rather than upscaled from a lower-resolution pass.

  • Aurora Autoregressive MoE Architecture: Unlike diffusion-transformer competitors, V1.5 uses an autoregressive mixture-of-experts network trained to predict next tokens from interleaved multimodal data. This unified token-space approach is what enables single-pass audio-video coherence.

  • Granular Duration Control (1тАУ15 seconds): Clips can be requested at any integer second from 1 to 15, supporting precise targeting for short-form formats.

  • Broad Format Support: Outputs H.264 MP4 at 24 FPS across seven aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3).


3. Model Architecture & Technical Details

xai/grok-imagine-video-v1.5/text-to-video is built on xAI's Aurora engine, an autoregressive mixture-of-experts network that predicts next tokens across an interleaved sequence of text, image, video, and audio modalities. This is architecturally distinct from the diffusion-transformer designs used by OpenAI Sora and Google Veo, and is the mechanism by which V1.5 produces joint video+audio output in a single forward pass rather than chaining separate generative models.

Key infrastructure and lineage points:

  • Training infrastructure: Trained on xAI's Colossus 2 supercomputer, a ~2 GW, ~555,000 NVIDIA GPU facility тАФ the largest known single-site AI training cluster.
  • R&D lineage: The video pipeline incorporates technology from Hotshot, a video generation startup acquired by xAI in March 2025.
  • Aurora foundation: The underlying Aurora image model was first released on December 9, 2024, with video capability progressively layered on top through Imagine 0.9 (October 2025), Imagine 1.0 (February 2026), multi-image and extension support (March 2026), and the V1.5 line from May 2026.
  • Joint token modeling: Because audio and video tokens are produced in the same autoregressive stream, lip-sync and event-aligned SFX emerge from the model rather than from separate alignment models.

xAI has not published a technical report, parameter count, training-data disclosure, or formal model card for V1.5, so finer architectural details (expert count, context length, tokenizer design) are not publicly documented.


4. Prompting Guidance

Text-to-video has no visual anchor, so the prompt carries the entire burden of composition. Prompts that specify the following tend to produce markedly better results:

  • Subject and setting тАФ who or what is on screen, and where.
  • Motion тАФ what changes over the clip's duration, described as a progression rather than a static description.
  • Camera тАФ shot size and movement (wide establishing shot, slow push-in, handheld follow, static tripod).
  • Audio direction тАФ dialogue lines, ambient bed, and specific sound events. Because audio is generated jointly, naming the sound you want is usually enough.
  • Lighting and grade тАФ time of day, key light direction, and colour treatment.

Duration should match the amount of action described. A single beat of motion reads well at 3тАУ5 seconds; a multi-beat sequence with a camera move needs 8тАУ15.


5. Use Cases

  • Short-Form Social Video: Vertical (9:16) and square outputs at 1тАУ15 seconds map directly to TikTok, Instagram Reels, YouTube Shorts, and X clip formats, with native audio eliminating the need for separate sound design.

  • Concept and Pre-Visualization: Prompt-only generation makes this the fastest mode for testing a visual idea before committing reference assets or production time.

  • Marketing and Advertising Creative: Rapid generation of brand teasers and ad concepts supports high-volume creative iteration and A/B testing of motion concepts.

  • Establishing Shots and B-Roll: Landscapes, cityscapes, atmospheric inserts, and transitional footage generate well from description alone and cut cleanly into longer edits.

  • Stock-Style Footage: Generic scenes that would otherwise require a stock licence or a shoot.

  • Entertainment and Viral Content: Low cost and granular duration control support meme, parody, and viral content generation.

The model is less well-suited to long-form storytelling, work requiring an exact likeness or a specific product (use reference-to-video), and animating an existing image (use image-to-video).

рд╕рдорд╛рди рдореЙрдбрд▓ рджреЗрдЦреЗрдВ

рд╣рд░ рдореАрдбрд┐рдпрд╛ AI рдХреЗ рд▓рд┐рдП рдПрдХ рд╣реА APIред

рд╕рднреА рдореЙрдбрд▓ рдПрдХреНрд╕рдкреНрд▓реЛрд░ рдХрд░реЗрдВ