
Grok Imagine Video gives developers access to xAI’s video family for creating, animating, extending, and editing clips. Generate 1 to 15 second videos from prompts, guide people, objects, or styles with up to seven reference images, and work at 480p or 720p. Atlas Cloud brings these capabilities together through OpenAI-compatible endpoints with transparent pay-as-you-go pricing. Start building today.
Grok Imagine Video is developed by xAI. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.
Choose the Grok Imagine Video endpoint that matches your source assets, desired transformation, and production workflow.
| Modality | Description |
|---|---|
| Grok Imagine Video T2V API (Text To Video) | Turn natural language prompts into videos lasting 1 to 15 seconds at 480p or 720p. This endpoint suits concept visualization, social clips, and scenes created without source media. |
| Grok Imagine Video I2V API (Image To Video) | Starting with a supplied image, this endpoint applies natural language motion prompts to produce video at 480p or 720p. Use it to animate artwork, product imagery, character frames, or campaign visuals. |
| Grok Imagine Video R2V API (Reference To Video) | Guide video generation with 1 to 7 reference images representing people, objects, or visual styles. Outputs can run for up to 10 seconds at 480p or 720p, supporting reference led creative production. |
| Grok Imagine Video Extend API (Video Extension) | Continue an existing 2 to 15 second MP4 with a prompt directed extension lasting 2 to 10 seconds. The output matches the input format and is capped at 720p, making this endpoint useful for lengthening established shots. |
| Grok Imagine Video Edit API (Video Editing) | Provide an MP4 and natural language instructions to modify its visual content while retaining the source duration. Output is capped at 8.7 seconds, which fits targeted revisions and prompt based transformations of short footage. |
Grok Imagine Video turns text, starting frames, and up to seven references into short clips, then extends or edits existing MP4 footage through Atlas Cloud's unified, pay-as-you-go API.
Write a natural-language prompt and generate a 1 to 15 second clip at 480p or 720p. Seven aspect ratios, including 16:9, 1:1, and 9:16, let each shot fit its intended canvas. Control subject, action, camera movement, and atmosphere directly in the prompt. It is a flexible starting point for story beats, campaign concepts, and social video.
Turn a supplied starting-frame image into a 1 to 15 second video at 480p or 720p. The image establishes the opening composition while a natural-language motion prompt directs movement and scene development. Choose from seven aspect ratios when a different output shape is needed. This workflow suits product shots, illustrated scenes, and art-directed concepts that need motion without rebuilding the opening frame.
Guide a video with 1 to 7 reference images representing people, objects, or visual styles. Refer to individual inputs in the prompt, then generate up to 10 seconds at 480p or 720p across seven supported aspect ratios. Because references shape the scene without serving only as a starting frame, creators gain more control over recurring visual elements. Use it for character-led scenes, product storytelling, or style-aware campaigns.
Continue an existing 2 to 15 second MP4 with a prompt-driven extension lasting 2 to 10 seconds. Grok Imagine Video picks up from the final frame and returns the original clip plus the new segment, matching the input resolution up to 720p. Describe the next action, reveal, or camera move in plain language. This makes abrupt endings easier to turn into complete narrative beats.
Give Grok Imagine Video an MP4 and natural-language instructions to transform the footage while retaining its source duration. Inputs may run up to 8.7 seconds, and billing follows the input video length. Ask for a visual change, atmosphere shift, or scene-level revision without rebuilding the clip from scratch. It is well suited to short creative variations and controlled post-generation refinements.
Move among text-to-video, image-to-video, reference-guided generation, extension, and editing through one unified API. Select the endpoint that matches the available source material and submit each job through an asynchronous prediction flow. Keep authentication and integration patterns consistent across the workflow. Pay-as-you-go access supports experiments and repeatable production pipelines. Developers can build broader video tools with less integration overhead.
See how Grok Imagine Video and two leading alternatives interpret identical prompts through motion, continuity, camera control, lighting, and visual storytelling.
An 8–10-second hyper-realistic outdoor adventure film set in an emerald gorge immediately after a torrential rainstorm. A kayaker in a cobalt-blue waterproof suit and orange-red life vest charges through violent whitewater from the first frame to the last. Open with an ultra-low waterline tracking shot racing beside the kayak as the paddle blades slash into the current, throwing crisp spray across the lens; the kayak surges up a steep wave crest and is launched into the air. Whip-pan into the kayaker’s first-person POV as the boat rolls sideways beneath a fallen tree, bursts through a translucent curtain of water, and narrowly clears jagged wet rocks, with hands, paddle, and cobalt bow remaining physically consistent through splashes and momentary occlusion. The kayaker plants the paddle against the current, carves off the canyon wall in a tight rebound turn, and drops back into the torrent. At the climax, transition to a drone plunge from the cliff top, descending rapidly and orbiting the kayaker during the landing; the bow strikes a floating wooden crate with convincing impact, breaking it open—and, as a sudden playful reveal, a connected string of perfectly intact colorful paper lanterns springs out, unfurls, and bobs warmly across the cold rushing water while the kayak races onward. Continuous anatomically correct motion, realistic kayak inertia, paddle resistance, collisions, turbulent fluid dynamics, fabric movement, droplets, lens spray, and unwavering subject identity after every obstruction; no slow motion, no empty establishing shots, no cuts to screens, interfaces, dashboards, progress bars, charts, captions, logos, or explanatory text. Overcast cold cyan daylight, rain-darkened emerald rock walls, vivid orange-red life vest, warm multicolored lantern glow, cinematic widescreen diagonal compositions emphasizing speed, high contrast, natural motion blur, immersive photorealistic action-sports cinematography. Synchronized audio: roaring rapids, sharp paddle impacts, wood cracking, strained breath, water striking the hull, and a bright percussive musical accent when the lanterns burst free. 16:9 aspect ratio.
Generated with Grok Imagine Video v1.5 Text-to-Video on Atlas Cloud
Generated with Seedance 2.5 Text-to-Video on Atlas Cloud
Generated with Grok Imagine Video v1.5 Developer Text-to-Video on Atlas Cloud
An 8–10-second continuous-action cinematic sequence in an open-air night market moments before a torrential downpour: a blindfolded ramen master sprints confidently through a maze of crowded food stalls while continuously stretching a ribbon of silver-white noodles between his hands. Begin with a perfectly vertical bird’s-eye view, then plunge rapidly from above into the neon-lit market labyrinth as the first heavy raindrops strike awnings and wet pavement. Lock onto the flying noodle ribbon and race inches beside it as it whips over glowing charcoal braziers, startled diners, sizzling grills, and umbrellas snapping open in the wind; preserve seamless human motion, believable elastic noodle physics, complex foreground occlusion, splashing rain, sparks, and dense steam. Whip-pan to the chef vaulting a low crate without breaking his stride, then execute a fast 360-degree orbit around him as he twists and releases the noodles backward with precise momentum. The noodle ribbon arcs cleanly into a vigorously boiling copper pot behind him—at the exact moment of triumph, a mischievous orange tabby leaps from beneath the stall, snatches half the dangling noodles in its mouth, and darts away, creating a crisp visual punchline while the chef keeps running, unaware. Neon-noir photorealism, cyan-green and magenta reflections across rain-slick ground, warm furnace-orange fire as rhythmic visual accents, tactile steam and rain, stable choreography, sharp cinematic detail, energetic real-time pacing, no slow motion, no cuts that break spatial or character continuity. Audio: rising thunder, market chatter, pounding footsteps, umbrella snaps, sizzling charcoal, boiling water, noodle whooshes synchronized to camera moves, ending with a bright comedic cat meow and copper-pot splash. No screens, user interfaces, dashboards, progress bars, charts, captions, subtitles, logos, watermarks, or explanatory text. 16:9 aspect ratio.
Generated with Grok Imagine Video v1.5 Text-to-Video on Atlas Cloud
Generated with Seedance 2.5 Text-to-Video on Atlas Cloud
Generated with Grok Imagine Video v1.5 Developer Text-to-Video on Atlas Cloud
Grok Imagine Video turns prompts, stills, references, and existing MP4 files into practical paths for social clips, product visuals, brand stories, longer sequences, targeted edits, and creator tools.
Grok Imagine Video converts natural language prompts into short 480p or 720p clips. Use it to produce social teasers, campaign concepts, and rapid visual drafts without preparing source footage in advance.
Starting from a product image, the image to video endpoint adds prompt directed motion at 480p or 720p. Brands can turn catalog stills into product reveals, launch assets, and marketplace visuals.
With one to seven images, reference to video guides people, objects, or styles into a new clip. Creative teams can develop brand stories, character concepts, and product scenes grounded in supplied visuals.
Continue an existing MP4 with a two to ten second prompt driven extension that preserves input resolution up to 720p. This supports longer storyboard beats, episodic transitions, and follow on sequences from approved footage.
Edit an MP4 through natural language instructions while retaining its source duration, with output capped at 8.7 seconds. Teams can create alternate looks, localized campaign variants, or revised visual treatments.
Build prompt driven video tools around the five Grok Imagine Video modes for generation, animation, references, extension, and editing. Developers can serve creator workflows from one family while selecting 480p or 720p output.
Compare Grok Imagine Video workflows with selected Atlas Cloud alternatives across input methods, clip length, maximum resolution, and standard pricing.
| Model | Input Workflow | Clip or Segment Length | Maximum Resolution | Standard Price |
|---|---|---|---|---|
| Grok Imagine Video Text-to-Video | Text prompt | 1–15s | 720p | $0.05/sec |
| Grok Imagine Video Image-to-Video | Start image + text prompt | 1–15s | 720p | $0.05/sec |
| Grok Imagine Video Reference-to-Video | 1–7 reference images + text prompt | 1–10s | 720p | $0.05/sec |
| Grok Imagine Video Extend | 2–15s MP4 + text prompt | 2–10s extension | Matches input, up to 720p | $0.07/sec |
| Grok Imagine Video Edit | MP4 + text instruction | Matches input, up to 8.7s | - | $0.07/input sec |
| MiniMax H3 Text-to-Video | Text prompt | 4–15s | 2K | $0.038/sec |
| Seedance 2.0 Image-to-Video | Start image + text, optional last frame | 4–15s | Native 4K | $0.112/sec |
| Vidu Q3-Mix Reference to Video | 1–4 reference images + text prompt | 1–16s | 1440p SR, native up to 1080p | $0.125/run |
| Wan-2.7 Video-edit | Source video + text, optional images or audio | - | 1080p | $0.10/run |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced Grok Imagine Video models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run Grok Imagine Video, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
Grok Imagine Video is xAI’s video generation model family for creating, animating, extending, and editing short clips. Atlas Cloud provides separate API endpoints for text, image, reference, extension, and editing workflows.
Grok Imagine Video can generate a clip from text, animate a starting image, or use up to seven reference images to guide people, objects, and styles. Separate endpoints can continue an existing clip or edit an MP4 through natural language instructions.
Create an Atlas Cloud API key and send an authenticated POST request to /api/v1/model/generateVideo. Specify the required model ID, prompt, and any image or video input required by that route. The response includes a request ID, status, and output fields.
Choose text-to-video when the scene begins with a prompt, or image-to-video when a supplied image must become the starting frame. Reference-to-video provides broader guidance from multiple images, while extend-video and edit-video operate on existing MP4 files.
Text-to-video and image-to-video support clips from 1 to 15 seconds at 480p or 720p. Reference-to-video supports 1 to 10 seconds at the same resolutions, while common aspect ratios include 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, and 2:3. Extension and editing routes follow their own input limits.
Yes. xAI documents native audio generation as part of the Grok Imagine video workflow, allowing sound and visuals to be produced together. Atlas Cloud’s published schemas for these routes do not expose a separate audio toggle.
Provide between one and seven public HTTPS image URLs or base64 data URIs through the reference-to-video endpoint. Tags such as <IMAGE_0> can connect individual references to people, objects, or styles described in the prompt. This route returns clips up to 10 seconds at 480p or 720p.
Use extend-video with a 2 to 15 second MP4 and request an additional 2 to 10 seconds of prompt-directed action. For targeted changes, edit-video accepts an MP4 no longer than 8.7 seconds and preserves its duration. Both routes cap output resolution at 720p.
Standard Atlas Cloud pricing is $0.05 per second for text-to-video, image-to-video, and reference-to-video. Extend-video and edit-video use a standard rate of $0.07 per second. Editing is billed according to the input video duration because the output retains that duration.
First confirm that the selected endpoint matches your input type and that every required prompt, image URL, or video URL is present. Check the route-specific duration, resolution, aspect ratio, reference count, and MP4 limits before resubmitting. If the request succeeds but the scene is inaccurate, rewrite the prompt with explicit subject motion, camera movement, pacing, and visual details.
Guides, tutorials, and product updates to help you get the most out of Atlas Cloud.