TYLKO DWA TYGODNIE | 20% ZNIŻKI na Seedream 5.0 Pro!
Strona główna
Eksploruj
xai/grok-imagine-video-v1.5/text-to-video
Grok Imagine Video v1.5 Text-to-Video
tekst-do-wideo

Grok Imagine Video v1.5 Text-to-Video API by xAI

xai/grok-imagine-video-v1.5/text-to-video
Text-to-video

xAI Grok Imagine Video v1.5 generates video with native synchronized audio from a text prompt alone. Up to 15s at 480p, 720p, or 1080p.

1. Introduction

Grok Imagine Video V1.5 is a frontier-tier video generation model developed by xAI that produces short clips of up to 15 seconds with natively generated, synchronized audio — including dialogue, lip-sync, sound effects, and ambient music — in a single inference pass.

This README applies to the following API model identifier:

  • xai/grok-imagine-video-v1.5/text-to-video

Text-to-video is the prompt-only mode: no input image or reference material is required. The model composes subject, motion, camera work, and the full audio bed from the prompt alone, and renders natively at up to 1080p.

Built on xAI's Aurora engine — an autoregressive mixture-of-experts (MoE) network that jointly models text, image, video, and audio tokens — the model represents a departure from the diffusion-transformer paradigm used by Sora and Veo, enabling tightly coupled audiovisual generation with competitive cost and latency characteristics.


2. Key Features

  • Prompt-Only Generation: No source image or reference asset is needed. Scene, motion, camera behaviour, and audio are all derived from the text prompt, making this the fastest path from idea to finished clip.

  • Native Synchronized Audio Generation: Audio (dialogue, lip-sync, SFX, ambient sound, music) is generated jointly with video tokens in a single inference pass rather than dubbed in post-processing. This produces event-aligned sound effects and natural lip-sync without requiring separate audio pipelines.

  • Native 1080p Output: Text-to-video renders at 480p, 720p, or 1080p, with 1080p produced natively rather than upscaled from a lower-resolution pass.

  • Aurora Autoregressive MoE Architecture: Unlike diffusion-transformer competitors, V1.5 uses an autoregressive mixture-of-experts network trained to predict next tokens from interleaved multimodal data. This unified token-space approach is what enables single-pass audio-video coherence.

  • Granular Duration Control (1–15 seconds): Clips can be requested at any integer second from 1 to 15, supporting precise targeting for short-form formats.

  • Broad Format Support: Outputs H.264 MP4 at 24 FPS across seven aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3).


3. Model Architecture & Technical Details

xai/grok-imagine-video-v1.5/text-to-video is built on xAI's Aurora engine, an autoregressive mixture-of-experts network that predicts next tokens across an interleaved sequence of text, image, video, and audio modalities. This is architecturally distinct from the diffusion-transformer designs used by OpenAI Sora and Google Veo, and is the mechanism by which V1.5 produces joint video+audio output in a single forward pass rather than chaining separate generative models.

Key infrastructure and lineage points:

  • Training infrastructure: Trained on xAI's Colossus 2 supercomputer, a ~2 GW, ~555,000 NVIDIA GPU facility — the largest known single-site AI training cluster.
  • R&D lineage: The video pipeline incorporates technology from Hotshot, a video generation startup acquired by xAI in March 2025.
  • Aurora foundation: The underlying Aurora image model was first released on December 9, 2024, with video capability progressively layered on top through Imagine 0.9 (October 2025), Imagine 1.0 (February 2026), multi-image and extension support (March 2026), and the V1.5 line from May 2026.
  • Joint token modeling: Because audio and video tokens are produced in the same autoregressive stream, lip-sync and event-aligned SFX emerge from the model rather than from separate alignment models.

xAI has not published a technical report, parameter count, training-data disclosure, or formal model card for V1.5, so finer architectural details (expert count, context length, tokenizer design) are not publicly documented.


4. Prompting Guidance

Text-to-video has no visual anchor, so the prompt carries the entire burden of composition. Prompts that specify the following tend to produce markedly better results:

  • Subject and setting — who or what is on screen, and where.
  • Motion — what changes over the clip's duration, described as a progression rather than a static description.
  • Camera — shot size and movement (wide establishing shot, slow push-in, handheld follow, static tripod).
  • Audio direction — dialogue lines, ambient bed, and specific sound events. Because audio is generated jointly, naming the sound you want is usually enough.
  • Lighting and grade — time of day, key light direction, and colour treatment.

Duration should match the amount of action described. A single beat of motion reads well at 3–5 seconds; a multi-beat sequence with a camera move needs 8–15.


5. Use Cases

  • Short-Form Social Video: Vertical (9:16) and square outputs at 1–15 seconds map directly to TikTok, Instagram Reels, YouTube Shorts, and X clip formats, with native audio eliminating the need for separate sound design.

  • Concept and Pre-Visualization: Prompt-only generation makes this the fastest mode for testing a visual idea before committing reference assets or production time.

  • Marketing and Advertising Creative: Rapid generation of brand teasers and ad concepts supports high-volume creative iteration and A/B testing of motion concepts.

  • Establishing Shots and B-Roll: Landscapes, cityscapes, atmospheric inserts, and transitional footage generate well from description alone and cut cleanly into longer edits.

  • Stock-Style Footage: Generic scenes that would otherwise require a stock licence or a shoot.

  • Entertainment and Viral Content: Low cost and granular duration control support meme, parody, and viral content generation.

The model is less well-suited to long-form storytelling, work requiring an exact likeness or a specific product (use reference-to-video), and animating an existing image (use image-to-video).

Eksploruj Podobne Modele

Jedno API do całej multimedialnej AI.

Przeglądaj wszystkie modele