
Grok Imagine Video 1.5 is xAI's video generation family for developers who need controlled motion from prompts and visual inputs. Animate a starting frame with natural language motion instructions, select 480p or 720p for reference workflows, and add an optional reference voice to guide the result. Access supported modes through one OpenAI-compatible Atlas Cloud key with transparent pay-as-you-go pricing. Start building today.
Grok Imagine Video 1.5 is developed by xAI. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.
Explore six endpoints that turn text, starting frames, or reference assets into video across developer and standard workflows.
| Modality | Description |
|---|---|
| Grok Imagine Video 1.5 Developer Reference-to-Video API | Turn 1 to 7 reference images and an optional reference voice into video with native synchronized audio. Outputs can run for up to 15 seconds at 480p or 720p. This endpoint suits reference-guided scenes and visual concepts that require multiple source images. |
| Grok Imagine Video 1.5 Reference-to-Video API | Guided by 1 to 7 reference images, this endpoint generates videos with native synchronized audio and can also accept an optional reference voice. It supports durations up to 15 seconds at 480p or 720p. Choose it for multi-reference video production and voice-guided sequences. |
| Grok Imagine Video 1.5 Developer Text-to-Video API | A text prompt is enough to generate video with native synchronized audio through this developer endpoint. Create clips up to 15 seconds long in 480p, 720p, or 1080p. It fits prompt-driven concepts, storyboards, and short-form video production. |
| Grok Imagine Video 1.5 Text-to-Video API | Create video directly from a text prompt while generating synchronized audio natively. Available output options include 480p, 720p, and 1080p, with a maximum duration of 15 seconds. Use it for scripted scenes, campaign concepts, or social video assets. |
| Grok Imagine Video 1.5 Developer Image-to-Video API | Starting from one image, this developer endpoint applies natural-language motion instructions to produce an animated video. Resolution options cover 480p, 720p, and 1080p. It is suited to animating artwork, product visuals, and prepared opening frames. |
| Grok Imagine Video 1.5 Image-to-Video API | Animate a starting-frame image by describing the desired movement in natural language. The resulting video can be generated at 480p, 720p, or 1080p. This workflow supports motion studies, visual storytelling, and image-led creative production. |
Grok Imagine Video 1.5 brings text, starting image, and one to seven reference image workflows together with native synchronized audio, up to 1080p output, and clips lasting up to 15 seconds in text or reference mode, while Atlas Cloud adds one API and transparent pay as you go pricing.
Start with a text prompt and generate clips up to 15 seconds at 480p, 720p, or 1080p. Native synchronized audio arrives with the video in the same generation. Shape the scene and action entirely through written direction when no source image exists. This route fits concept films, cinematic experiments, and content pipelines driven by prompts.
A single starting frame becomes a moving sequence through natural language motion prompts. Choose 480p, 720p, or 1080p output while the image establishes the opening composition. Direct subject movement and scene motion in the prompt, then let the model animate forward from that visual anchor. It suits product shots, character moments, and art directed stills that need motion.
Guide Grok Imagine Video 1.5 with one to seven reference images and an optional reference voice. Reference mode produces clips up to 15 seconds at 480p or 720p, with native synchronized audio. Use those inputs to steer the generated scene toward an established visual direction. This route suits campaigns and stories built from existing creative references.
Video and sound are generated together, so every clip can include native synchronized audio without a separate audio pass. Build scenes around visible actions whose timing can be reinforced by the accompanying audio track. For moments where synchronization matters, this combined generation path reduces the handoff between visual creation and sound production. It is especially useful for performance and atmosphere rich footage.
Stretch a complete visual beat across text or reference clips lasting up to 15 seconds instead of stopping at a brief loop. The available duration supports setup, movement, and a clear payoff within one generated sequence. Pair that runway with changing camera perspectives and continuous action to create a more developed moment. This length works well for short advertisements, narrative reveals, and social video scenes.
Move among generation from text, starting images, and references through one Atlas Cloud API, with transparent pay as you go pricing. Select the workflow that matches each asset rather than forcing every idea through one input type. The same integration gives developers access to all six listed Grok Imagine Video v1.5 endpoints, including standard and Developer routes. This simplifies experimentation and production planning across varied video jobs.
See how Grok Imagine Video 1.5 and two Atlas Cloud alternatives interpret identical prompts across cinematic action and stylized fantasy.
A 7-second single-take underwater Baroque fantasy film inside an abandoned, fully flooded opera house: a graceful freediver with flowing hair and a luminous pearl mask pursues a vivid coral-red octopus carrying an ornate brass key through drifting velvet curtains. Begin from inside a cracked stage floor with a dramatic low-angle upward shot as she rolls and dives headfirst through the fissure; transition into a tight lateral tracking move as she threads between toppled theater seats, her hair and costume responding naturally to currents while curtains billow and bubbles trail through layered arches; then orbit around her and crane upward as water pressure tears a crystal chandelier loose overhead. At the climax, she twists sideways at the last instant to avoid the falling chandelier, crystals colliding and scattering realistically, while the octopus deftly catches the tumbling key with its complex curling tentacles, pauses with solemn ceremony, and places it back into her open hand. Continuous fluid motion, clear beginning–chase–payoff, consistent character and pearl mask, physically accurate water resistance, buoyancy, fabric, hair, bubbles, tentacle deformation, impacts, and collision avoidance. Cold cyan shafts from shattered skylights cut through the dark submerged auditorium; coral red is the only highly saturated color, guiding the eye through deep layers of arches, curtains, suspended debris, and bubbles. Photorealistic cinematic underwater photography with a subtle oil-painted texture, elegant Baroque grandeur, volumetric caustics, high detail, no text, no interface, no cuts disguised as still frames. Immersive audio: muffled underwater rumbles, rushing currents, fabric flutter, bubbles, crystalline impacts, and a tense orchestral pulse resolving into one delicate harp note as the key is returned. 16:9 aspect ratio.
Generated with Grok Imagine Video v1.5 Text-to-Video on Atlas Cloud
Generated with Kling V3.0 Turbo Text-to-Video on Atlas Cloud
Generated with Grok Imagine Video v1.5 Developer Text-to-Video on Atlas Cloud
A 8-second absurdist action-comedy commercial set inside a meticulously detailed 1950s barbershop at dawn. A grumpy, shaggy Old English Sheepdog, drenched in thick white soap foam, bursts through the door and charges across the shop while shaking suds everywhere; a startled barber jumps onto a spinning chair and rides it in pursuit, snipping rapidly at the dog’s flying coat, every crisp scissor snap precisely synchronized with punchy jazz drumbeats. Begin with an extreme ground-level tracking shot racing beside wet paws as they skid across glossy black-and-white checkerboard tiles, sending realistic droplets and foam splashes toward the lens; whip-pan upward into a fast circular tracking shot that uses three large mirrors to continuously reveal and preserve the shifting positions of the dog and barber through complex, accurate reflections; then execute a rapid dolly-in as the final airborne clippings tumble, swirl, and land perfectly on the dog’s head, forming an outrageously tall pompadour. The dog stops, poses smugly in a one-second freeze-like hero beat, then suddenly sneezes with an explosive puff of foam and hair as the jazz ends on a sharp comedic rimshot. Warm tungsten practical lights slice through cool teal morning fog outside the windows; rich burgundy, cream, polished brass, and dark wood vintage palette; premium live-action advertising cinematography, seamless continuous high-speed choreography, consistent characters and spatial continuity, physically accurate wet fur, soap foam, loose hair, reflections, chair rotation, inertia, and collisions, tactile photorealistic detail, crisp motion clarity, no slow motion, no empty establishing shots, no screens, no software interfaces, no dashboards, no progress bars, no charts, no captions, no explanatory text, no logos, no watermarks, 16:9 aspect ratio.
Generated with Grok Imagine Video v1.5 Text-to-Video on Atlas Cloud
Generated with Kling V3.0 Turbo Text-to-Video on Atlas Cloud
Generated with Grok Imagine Video v1.5 Developer Text-to-Video on Atlas Cloud
Grok Imagine Video 1.5 turns prompts, source frames, and up to seven reference images into short videos with synchronized audio for campaigns, product showcases, character stories, social content, and rapid visual prototyping.
Animate a product still with natural language motion prompts and synchronized audio. Create short showcase clips for storefronts, campaign pages, launch assets, or paid social placements at up to 1080p.
Guide a scene with one to seven character references and an optional reference voice. This setup supports recurring cast appearances, dialogue moments, and short narrative clips with native synchronized audio.
Start from a text prompt to generate a complete video with native synchronized audio. Marketing teams can explore campaign concepts, mood pieces, and launch teasers without preparing a source image.
Turn a single starting frame into motion using natural language direction at up to 1080p. Produce concise campaign assets for feeds, launch pages, event announcements, product updates, and creator led posts.
Bring an approved storyboard frame to life with prompt directed movement while retaining it as the starting image. Directors and designers can create short previz shots with synchronized audio for review.
Combine product views, character portraits, wardrobe details, and scene references across as many as seven images. Brand teams can direct short promotional scenes with synchronized sound while grounding each request in supplied visual assets.
Compare Grok Imagine Video 1.5 with leading reference-to-video models on accepted inputs, clip length, resolution, synchronized audio, and standard per-second pricing.
| Model | Input Types | Clip Length | Max Resolution | Native Audio | Standard Price |
|---|---|---|---|---|---|
| Grok Imagine Video v1.5 | Text, Image, Reference Images, Voice | Up to 15 seconds | 1080p | √ | $0.08/second |
| Seedance 2.0 Reference-to-Video | Text, Image, Video, Audio | 4 to 15 seconds | 4K | √ | $0.112/second |
| MiniMax H3 Reference-to-Video | Text, Image, Video, Audio | 4 to 15 seconds | 2K | √ | $0.038/second |
| Wan-2.7 Reference-to-video | Text, Image, Video, Audio | 2 to 10 seconds | 1080p | √ | $0.10/second |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced Grok Imagine Video 1.5 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run Grok Imagine Video 1.5, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
Grok Imagine Video 1.5 is xAI's video generation model family for creating clips from text, a starting image, or multiple reference images. Through Atlas Cloud, developers can access separate endpoints for each generation mode.
Three modes are available: text-to-video, image-to-video, and reference-to-video. Choose text for creating a scene from scratch, image for animating a starting frame, or reference mode for guiding subjects, objects, and styles with multiple images.
Create an Atlas Cloud API key and send an authenticated request to the video generation endpoint with the appropriate model ID. Use the returned task ID to check the generation status and retrieve the completed video URL.
Text-to-video requires a natural-language prompt, while image-to-video adds one starting-frame image. Reference-to-video accepts a prompt and 1 to 7 reference images, with optional preset voice IDs for generated dialogue.
The available Atlas Cloud schemas support clips from 1 to 15 seconds. Text-to-video and image-to-video offer 480p, 720p, or 1080p output, while reference-to-video supports 480p or 720p.
Yes. Text-to-video generates native synchronized audio from the prompt, and reference-to-video can combine generated audio with an optional preset voice for dialogue.
Atlas Cloud lists a standard base price of $0.08 for each endpoint covered by this family page. Review the selected endpoint's current billing details before deployment because the supplied price record does not specify its billing unit.
Supply between 1 and 7 images through the reference-to-video endpoint. Identify specific assets in the prompt with tags such as <IMAGE_0> and <IMAGE_1> so the model can associate each reference with the intended subject or scene element.
The available Atlas Cloud schemas do not include a dedicated last-frame parameter. Image-to-video uses one image as the starting frame, while reference-to-video uses 1 to 7 images as visual guidance rather than guaranteed first and last frames.
Choose reference-to-video when several images need to guide the people, objects, or visual style in a new scene. Use image-to-video when one existing image should become the opening frame and its motion should follow a natural-language prompt.
Guides, tutorials, and product updates to help you get the most out of Atlas Cloud.