Seedance 2.0 Mini & Fast API at Lowest Prices Worldwide — up to 68% off official pricing
Grok Imagine Video with Reference Guidance

Grok Imagine Video with Reference Guidance

Grok Imagine Video gives developers access to xAI’s video family for creating, animating, extending, and editing clips. Generate 1 to 15 second videos from prompts, guide people, objects, or styles with up to seven reference images, and work at 480p or 720p. Atlas Cloud brings these capabilities together through OpenAI-compatible endpoints with transparent pay-as-you-go pricing. Start building today.

Grok Imagine Video is developed by xAI. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.

Compare Grok Imagine Video Modes and Workflows

Choose the Grok Imagine Video endpoint that matches your source assets, desired transformation, and production workflow.

ModalityDescription
Grok Imagine Video T2V API (Text To Video)Turn natural language prompts into videos lasting 1 to 15 seconds at 480p or 720p. This endpoint suits concept visualization, social clips, and scenes created without source media.
Grok Imagine Video I2V API (Image To Video)Starting with a supplied image, this endpoint applies natural language motion prompts to produce video at 480p or 720p. Use it to animate artwork, product imagery, character frames, or campaign visuals.
Grok Imagine Video R2V API (Reference To Video)Guide video generation with 1 to 7 reference images representing people, objects, or visual styles. Outputs can run for up to 10 seconds at 480p or 720p, supporting reference led creative production.
Grok Imagine Video Extend API (Video Extension)Continue an existing 2 to 15 second MP4 with a prompt directed extension lasting 2 to 10 seconds. The output matches the input format and is capped at 720p, making this endpoint useful for lengthening established shots.
Grok Imagine Video Edit API (Video Editing)Provide an MP4 and natural language instructions to modify its visual content while retaining the source duration. Output is capped at 8.7 seconds, which fits targeted revisions and prompt based transformations of short footage.

Grok Imagine Video Across Five Creative Workflows

Grok Imagine Video turns text, starting frames, and up to seven references into short clips, then extends or edits existing MP4 footage through Atlas Cloud's unified, pay-as-you-go API.

Grok Imagine Video from Text

Write a natural-language prompt and generate a 1 to 15 second clip at 480p or 720p. Seven aspect ratios, including 16:9, 1:1, and 9:16, let each shot fit its intended canvas. Control subject, action, camera movement, and atmosphere directly in the prompt. It is a flexible starting point for story beats, campaign concepts, and social video.

Animate a Starting Frame

Turn a supplied starting-frame image into a 1 to 15 second video at 480p or 720p. The image establishes the opening composition while a natural-language motion prompt directs movement and scene development. Choose from seven aspect ratios when a different output shape is needed. This workflow suits product shots, illustrated scenes, and art-directed concepts that need motion without rebuilding the opening frame.

Grok Imagine Video Reference Control

Guide a video with 1 to 7 reference images representing people, objects, or visual styles. Refer to individual inputs in the prompt, then generate up to 10 seconds at 480p or 720p across seven supported aspect ratios. Because references shape the scene without serving only as a starting frame, creators gain more control over recurring visual elements. Use it for character-led scenes, product storytelling, or style-aware campaigns.

Extend the Storyline

Continue an existing 2 to 15 second MP4 with a prompt-driven extension lasting 2 to 10 seconds. Grok Imagine Video picks up from the final frame and returns the original clip plus the new segment, matching the input resolution up to 720p. Describe the next action, reveal, or camera move in plain language. This makes abrupt endings easier to turn into complete narrative beats.

Edit Clips with Grok Imagine Video

Give Grok Imagine Video an MP4 and natural-language instructions to transform the footage while retaining its source duration. Inputs may run up to 8.7 seconds, and billing follows the input video length. Ask for a visual change, atmosphere shift, or scene-level revision without rebuilding the clip from scratch. It is well suited to short creative variations and controlled post-generation refinements.

Five Workflows, One API

Move among text-to-video, image-to-video, reference-guided generation, extension, and editing through one unified API. Select the endpoint that matches the available source material and submit each job through an asynchronous prediction flow. Keep authentication and integration patterns consistent across the workflow. Pay-as-you-go access supports experiments and repeatable production pipelines. Developers can build broader video tools with less integration overhead.

One Prompt, Three Visions: Grok Imagine Video Compared

See how Grok Imagine Video and two leading alternatives interpret identical prompts through motion, continuity, camera control, lighting, and visual storytelling.

Prompt

An 8–10-second hyper-realistic outdoor adventure film set in an emerald gorge immediately after a torrential rainstorm. A kayaker in a cobalt-blue waterproof suit and orange-red life vest charges through violent whitewater from the first frame to the last. Open with an ultra-low waterline tracking shot racing beside the kayak as the paddle blades slash into the current, throwing crisp spray across the lens; the kayak surges up a steep wave crest and is launched into the air. Whip-pan into the kayaker’s first-person POV as the boat rolls sideways beneath a fallen tree, bursts through a translucent curtain of water, and narrowly clears jagged wet rocks, with hands, paddle, and cobalt bow remaining physically consistent through splashes and momentary occlusion. The kayaker plants the paddle against the current, carves off the canyon wall in a tight rebound turn, and drops back into the torrent. At the climax, transition to a drone plunge from the cliff top, descending rapidly and orbiting the kayaker during the landing; the bow strikes a floating wooden crate with convincing impact, breaking it open—and, as a sudden playful reveal, a connected string of perfectly intact colorful paper lanterns springs out, unfurls, and bobs warmly across the cold rushing water while the kayak races onward. Continuous anatomically correct motion, realistic kayak inertia, paddle resistance, collisions, turbulent fluid dynamics, fabric movement, droplets, lens spray, and unwavering subject identity after every obstruction; no slow motion, no empty establishing shots, no cuts to screens, interfaces, dashboards, progress bars, charts, captions, logos, or explanatory text. Overcast cold cyan daylight, rain-darkened emerald rock walls, vivid orange-red life vest, warm multicolored lantern glow, cinematic widescreen diagonal compositions emphasizing speed, high contrast, natural motion blur, immersive photorealistic action-sports cinematography. Synchronized audio: roaring rapids, sharp paddle impacts, wood cracking, strained breath, water striking the hull, and a bright percussive musical accent when the lanterns burst free. 16:9 aspect ratio.

Generated with Grok Imagine Video v1.5 Text-to-Video on Atlas Cloud

Generated with Seedance 2.5 Text-to-Video on Atlas Cloud

Generated with Grok Imagine Video v1.5 Developer Text-to-Video on Atlas Cloud

Prompt

An 8–10-second continuous-action cinematic sequence in an open-air night market moments before a torrential downpour: a blindfolded ramen master sprints confidently through a maze of crowded food stalls while continuously stretching a ribbon of silver-white noodles between his hands. Begin with a perfectly vertical bird’s-eye view, then plunge rapidly from above into the neon-lit market labyrinth as the first heavy raindrops strike awnings and wet pavement. Lock onto the flying noodle ribbon and race inches beside it as it whips over glowing charcoal braziers, startled diners, sizzling grills, and umbrellas snapping open in the wind; preserve seamless human motion, believable elastic noodle physics, complex foreground occlusion, splashing rain, sparks, and dense steam. Whip-pan to the chef vaulting a low crate without breaking his stride, then execute a fast 360-degree orbit around him as he twists and releases the noodles backward with precise momentum. The noodle ribbon arcs cleanly into a vigorously boiling copper pot behind him—at the exact moment of triumph, a mischievous orange tabby leaps from beneath the stall, snatches half the dangling noodles in its mouth, and darts away, creating a crisp visual punchline while the chef keeps running, unaware. Neon-noir photorealism, cyan-green and magenta reflections across rain-slick ground, warm furnace-orange fire as rhythmic visual accents, tactile steam and rain, stable choreography, sharp cinematic detail, energetic real-time pacing, no slow motion, no cuts that break spatial or character continuity. Audio: rising thunder, market chatter, pounding footsteps, umbrella snaps, sizzling charcoal, boiling water, noodle whooshes synchronized to camera moves, ending with a bright comedic cat meow and copper-pot splash. No screens, user interfaces, dashboards, progress bars, charts, captions, subtitles, logos, watermarks, or explanatory text. 16:9 aspect ratio.

Generated with Grok Imagine Video v1.5 Text-to-Video on Atlas Cloud

Generated with Seedance 2.5 Text-to-Video on Atlas Cloud

Generated with Grok Imagine Video v1.5 Developer Text-to-Video on Atlas Cloud

Grok Imagine Video Across Every Creative Stage

Grok Imagine Video turns prompts, stills, references, and existing MP4 files into practical paths for social clips, product visuals, brand stories, longer sequences, targeted edits, and creator tools.

Grok Imagine Video Social Clips

Grok Imagine Video converts natural language prompts into short 480p or 720p clips. Use it to produce social teasers, campaign concepts, and rapid visual drafts without preparing source footage in advance.

Animate Product Stills

Starting from a product image, the image to video endpoint adds prompt directed motion at 480p or 720p. Brands can turn catalog stills into product reveals, launch assets, and marketplace visuals.

Reference Guided Brand Stories

With one to seven images, reference to video guides people, objects, or styles into a new clip. Creative teams can develop brand stories, character concepts, and product scenes grounded in supplied visuals.

Grok Imagine Video Story Extensions

Continue an existing MP4 with a two to ten second prompt driven extension that preserves input resolution up to 720p. This supports longer storyboard beats, episodic transitions, and follow on sequences from approved footage.

Natural Language Video Edits

Edit an MP4 through natural language instructions while retaining its source duration, with output capped at 8.7 seconds. Teams can create alternate looks, localized campaign variants, or revised visual treatments.

Build with Grok Imagine Video

Build prompt driven video tools around the five Grok Imagine Video modes for generation, animation, references, extension, and editing. Developers can serve creator workflows from one family while selecting 480p or 720p output.

Where Grok Imagine Video Stands Among Video Models

Compare Grok Imagine Video workflows with selected Atlas Cloud alternatives across input methods, clip length, maximum resolution, and standard pricing.

ModelInput WorkflowClip or Segment LengthMaximum ResolutionStandard Price
Grok Imagine Video Text-to-VideoText prompt1–15s720p$0.05/sec
Grok Imagine Video Image-to-VideoStart image + text prompt1–15s720p$0.05/sec
Grok Imagine Video Reference-to-Video1–7 reference images + text prompt1–10s720p$0.05/sec
Grok Imagine Video Extend2–15s MP4 + text prompt2–10s extensionMatches input, up to 720p$0.07/sec
Grok Imagine Video EditMP4 + text instructionMatches input, up to 8.7s-$0.07/input sec
MiniMax H3 Text-to-VideoText prompt4–15s2K$0.038/sec
Seedance 2.0 Image-to-VideoStart image + text, optional last frame4–15sNative 4K$0.112/sec
Vidu Q3-Mix Reference to Video1–4 reference images + text prompt1–16s1440p SR, native up to 1080p$0.125/run
Wan-2.7 Video-editSource video + text, optional images or audio-1080p$0.10/run

How to Use Grok Imagine Video on Atlas Cloud

Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.

Create an Atlas Cloud Account

Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.

Why Use Grok Imagine Video on Atlas Cloud

Combining the advanced Grok Imagine Video models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.

Performance & flexibility

Low Latency:
GPU-optimized inference for real-time reasoning.

Unified API:
Run Grok Imagine Video, GPT, Gemini, and DeepSeek with one integration.

Transparent Pricing:
Predictable per-token billing with serverless options.

Enterprise & Scale

Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.

Reliability:
99.99% uptime, RBAC, and compliance-ready logging.

Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.

Grok Imagine Video API Questions, Answered

Grok Imagine Video is xAI’s video generation model family for creating, animating, extending, and editing short clips. Atlas Cloud provides separate API endpoints for text, image, reference, extension, and editing workflows.

Grok Imagine Video can generate a clip from text, animate a starting image, or use up to seven reference images to guide people, objects, and styles. Separate endpoints can continue an existing clip or edit an MP4 through natural language instructions.

Create an Atlas Cloud API key and send an authenticated POST request to /api/v1/model/generateVideo. Specify the required model ID, prompt, and any image or video input required by that route. The response includes a request ID, status, and output fields.

Choose text-to-video when the scene begins with a prompt, or image-to-video when a supplied image must become the starting frame. Reference-to-video provides broader guidance from multiple images, while extend-video and edit-video operate on existing MP4 files.

Text-to-video and image-to-video support clips from 1 to 15 seconds at 480p or 720p. Reference-to-video supports 1 to 10 seconds at the same resolutions, while common aspect ratios include 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, and 2:3. Extension and editing routes follow their own input limits.

Yes. xAI documents native audio generation as part of the Grok Imagine video workflow, allowing sound and visuals to be produced together. Atlas Cloud’s published schemas for these routes do not expose a separate audio toggle.

Provide between one and seven public HTTPS image URLs or base64 data URIs through the reference-to-video endpoint. Tags such as <IMAGE_0> can connect individual references to people, objects, or styles described in the prompt. This route returns clips up to 10 seconds at 480p or 720p.

Use extend-video with a 2 to 15 second MP4 and request an additional 2 to 10 seconds of prompt-directed action. For targeted changes, edit-video accepts an MP4 no longer than 8.7 seconds and preserves its duration. Both routes cap output resolution at 720p.

Standard Atlas Cloud pricing is $0.05 per second for text-to-video, image-to-video, and reference-to-video. Extend-video and edit-video use a standard rate of $0.07 per second. Editing is billed according to the input video duration because the output retains that duration.

First confirm that the selected endpoint matches your input type and that every required prompt, image URL, or video URL is present. Check the route-specific duration, resolution, aspect ratio, reference count, and MP4 limits before resubmitting. If the request succeeds but the scene is inaccurate, rewrite the prompt with explicit subject motion, camera movement, pacing, and visual details.

Explore More Families

Seedance 2.5

Seedance 2.5 API is now available on Atlas Cloud! It gives developers ByteDance's newest video model. It generates up to 30 seconds of native video in a single pass from text, a single image, or as many as 50 multimodal references, with synchronized audio and in-frame multilingual text. On Atlas Cloud you reach it through one key, with subject consistency and improved physics keeping long shots coherent. (Update: Seedance 2.5 1080P API Is Available NOW!)

View Family

Wan 3.0

Wan 3.0 API is the next generation of Alibaba's Wan video family, built to push long-form generation, multi-reference control, and audiovisual quality to new heights. Atlas Cloud already hosts Wan 2.7, 2.6, and 2.5, and Wan 3.0 runs on the same unified key with no separate setup. Start building today. Scroll down to the showcase to see what Wan 3.0 can create.

View Family

MiniMax H3

MiniMax H3 is MiniMax's multimodal video family for text, image, and reference guided creation. Across supported routes, it preserves subjects from reference media, offers flexible aspect ratios, and pairs generated sound with visuals through H3 Developer, with output profiles selected by endpoint. Atlas Cloud unifies the family behind one OpenAI-compatible key with transparent pay-as-you-go pricing from the standard rate of $0.038 per second. Start building today.

View Family

Seedream 5.0 Pro

Seedream 5.0 Pro API gives developers ByteDance's controllable image editing model on Atlas Cloud. It places edits precisely with anchors and coordinates, separates images into editable layers, fuses multiple references, and matches exact colors and materials, with multilingual text at 2K and 3K. On Atlas Cloud you reach it through one key!

View Family

Seedance 2.0

Seedance 2.0 is ByteDance’s production video model for precise shot creation. Turn prompts into video, animate a first-frame image with optional last-frame guidance, or shape results with reference media and optional web search. Atlas Cloud brings these workflows into one unified API with transparent pay-as-you-go pricing and one OpenAI-compatible key. Start building today.

View Family

GPT Image 2.5

The gpt-image-2.5 family from OpenAI gives developers a choice of Flare and Sunburst for production image workflows. Render at arbitrary resolutions up to 3840x2160 and select from five quality tiers, including xhigh and max, to match specific output requirements. Atlas Cloud provides ready-to-use REST inference with no cold starts and standard pricing from $0.004 per generation. Start building today.

View Family

GPT Image 2

The GPT Image 2 API gives developers access to OpenAI's latest image model, the successor to GPT Image 1.5. It generates and edits images with accurate text rendering across Latin and CJK scripts, plus strong composition for posters, mockups, and infographics. On Atlas Cloud you reach it through one unified API alongside 300+ models, with free credits, 99.99% uptime, and no OpenAI organization verification required.

View Family

Gemini Omni Flash

The gemini omni API brings Google DeepMind's natively multimodal Gemini Omni Flash family, including Gemini Omni 1.1 Flash, to developers. Create cinematic video with synchronized native audio, animate still images with precise start and end frame control, or revise existing footage through text guided edits that preserve untouched content. Atlas Cloud provides one OpenAI-compatible key, unified access, and transparent pay-as-you-go pricing. Start building today.

View Family

Grok Imagine

Grok Imagine Image is xAI's family for generating polished visuals and revising one or more reference images through natural language instructions. Its standard and quality endpoints cover text to image creation, single image changes, and indexed multi-image composition. Atlas Cloud brings these workflows into one API, with standard generation and editing priced at $0.02 per image. Start building today.

View Family

Google

Google's most powerful creative models are all available on Atlas Cloud. Veo 3.1 delivers cinematic video generation, Nano Banana 2 powers high-fidelity image creation, and Gemini brings multimodal intelligence to every workflow. Access the full Google model suite through one API key with Day-0 availability and pay-as-you-go pricing.

View Family

Seedance 2.0 Mini

The Seedance 2.0 Mini API is the lightest, lowest-cost tier of ByteDance's Seedance video line, built for teams where throughput and unit cost matter more than maximum polish. Use it for batch generation, rapid prototyping, and draft passes, all through one OpenAI-compatible key on Atlas Cloud.

View Family

ByteDance

From cinematic video generation to high-fidelity image creation, ByteDance's most powerful models are live on Atlas Cloud. Run Seedance and Seedream at scale with the lowest inference pricing and zero infrastructure overhead.

View Family

One API for All Media AI.

Explore all models