



The Grok Imagine API covers xAI's image, video, and speech models, from Image 2.0 to Video 1.5 and xAI TTS v1. Render 1K or 2K stills across 14 aspect ratios, push a scene to 15 seconds of 1080p motion, steer shots with up to 7 reference images, or narrate them in 20 languages. Atlas Cloud runs every mode on one endpoint, priced pay-as-you-go from $0.02 per image and $0.05 per second. Start building today.
Grok Imagine is developed by xAI. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.
Atlas Cloud provides you with the latest industry-leading creative models.
Match the right Grok Imagine API model to your input type, output length, resolution, and per-call price.
| Modality | Description |
|---|---|
| Grok Imagine Image 2.0 T2I API (Text to Image) | Write a prompt of up to 8,000 characters and this endpoint returns one to four images at 1K or 2K resolution. A low or medium quality tier lets you trade rendering effort against turnaround, with pricing at $0.04 per image. It fits concept exploration, social creatives, and batch visual production where volume matters. |
| Grok Imagine Image 2.0 Edit API (Image to Image) | Hand the endpoint as many as three reference images with a natural-language instruction, and edited results come back at 1K or 2K in the aspect ratio you set or in auto. The same low and medium quality tiers apply, and each image costs $0.04. Product restyling, background swaps, and fast design variations all run through one call. |
| Grok Imagine Image Quality T2I API (Text to Image) | Where finer detail matters, this endpoint turns text prompts into polished stills at 1K or 2K and returns up to four variations per request at $0.05 per image. Aspect ratios span square, portrait, landscape, and ultrawide formats. Reach for it on hero images, advertising creatives, and brand-grade product renders. |
| Grok Imagine Image Quality Edit API (Image to Image) | Feed one to eight source images with an editing instruction and receive revised versions at 1K or 2K for $0.05 per image. Multi-image references are addressed inline as <IMAGE_0> and <IMAGE_1>, so subjects and styles can be combined in a single instruction. Campaign refreshes, style transfers, and composite edits are the usual jobs. |
| Grok Imagine Image T2I API (Text to Image) | The original text-to-image endpoint keeps 1K and 2K output and four-image batching at the lowest image price in the family, $0.02 per image. That makes it a practical default for high-volume pipelines, thumbnail generation, and rapid prompt testing. |
| Grok Imagine Image Edit API (Image to Image) | Need lightweight edits at scale? This endpoint accepts one to eight images with a text instruction and returns as many as four edited results at 1K or 2K for $0.02 per image. Catalog cleanup, seasonal variants, and iterative design passes stay inexpensive. |
| Grok Imagine Video v1.5 T2V API (Text to Video) | A prompt alone produces one to fifteen seconds of video with native synchronized audio at 480p, 720p, or 1080p, across seven aspect ratios. Billing is $0.08 per second of output. Trailers, social spots, and narrative scenes that need sound baked in suit it best. |
| Grok Imagine Video v1.5 I2V API (Image to Video) | Supply a starting frame and a motion prompt, and v1.5 animates it for up to fifteen seconds at resolutions reaching 1080p with audio generated in the same pass. Cost runs $0.08 per second. Product spins, portrait animation, and still-to-motion campaign assets fit here. |
| Grok Imagine Video v1.5 R2V API (Reference to Video) | One to seven reference images guide the characters, objects, and style on screen, while up to three voices chosen from 26 presets drive the spoken audio. Output reaches fifteen seconds at 480p or 720p for $0.08 per second. Character-consistent storytelling, virtual try-on, and branded mascots benefit most. |
| Grok Imagine Video T2V API (Text to Video) | Text prompts become clips of one to fifteen seconds at 480p or 720p, with seven aspect ratios covering landscape, square, and vertical delivery. At $0.05 per second, it is the value option for social content and rapid concept boards. |
| Grok Imagine Video I2V API (Image to Video) | Anchor a still as the first frame, describe the motion you want, and the clip extends to fifteen seconds at 480p or 720p. Pricing holds at $0.05 per second, which keeps bulk animation of product photos and portraits affordable. |
| Grok Imagine Video R2V API (Reference to Video) | Between one and seven reference images shape who and what appears on screen without pinning down a fixed opening frame. Clips run up to ten seconds at 480p or 720p and cost $0.05 per second. Use it when identity and style must stay consistent from shot to shot. |
| Grok Imagine Video Extend API (Video Extension) | An existing mp4 of two to fifteen seconds can be continued by another two to ten seconds, guided by a prompt that sets what happens next. Output matches the source video and is capped at 720p, billed at $0.07 per second. It is handy for stretching a short cut to a required ad length. |
| Grok Imagine Video Edit API (Video to Video) | Natural-language instructions rewrite an existing mp4 while the original scene, motion, and framing survive. Output keeps the source duration, capped at 8.7 seconds, and billing follows the input video at $0.07 per second. Prop additions, wardrobe changes, and restyling happen without a reshoot. |
| xAI TTS v1 API (Text to Speech) | Up to 15,000 characters of text become natural speech with sub-second latency across 20 languages, with automatic language detection available. Delivery is tunable: playback speed from 0.7x to 1.5x, codecs including mp3, wav, and pcm, and sample rates up to 48 kHz. Voiceovers, IVR prompts, and in-app narration are the natural fits. |
Each part of the Grok Imagine API answers a different production need, covering 2K stills across 14 aspect ratios, reference-guided editing, video that carries its own synchronized audio, and speech in 20 languages.

Dialogue, ambience, and music are generated in the same pass as the motion, so clips up to 15 seconds arrive already in sync. No separate scoring or dubbing step is needed before publishing.

Grok Imagine Image 2.0 holds photorealistic texture at 1K or 2K across 14 aspect ratios, with low and medium quality tiers at $0.04 per image. One generation covers hero banners, print crops, and close-ups.

Feed in a photo and Grok Imagine Video v1.5 anchors it as the opening frame, then builds motion around it at up to 1080p. Catalog shots and portraits become showreels without any reshoot.

Up to three reference images can be edited together at 1K or 2K, with plain-language instructions changing only the elements you name. At $0.04 per edit, design iterations and on-brand variant sets stay inexpensive.

Between one and seven reference images carry characters, props, and style into clips up to 15 seconds long, with no fixed start frame. Virtual try-on, product placement, and character-driven series keep one recognizable look.

xAI TTS v1 ships inside the same suite, turning text into expressive speech across 20 languages with sub-second latency. Pair it with generated video for narrated explainers, localized ads, and in-app assistants.

Fine-grained control over pace, emphasis, and tone shapes each read, and sub-second synthesis makes auditioning a line almost instant. Across more than 80 voices, one engine covers audiobooks, dubbing, and character dialogue.
Each row runs the identical prompt through the Grok Imagine API and two rival models on Atlas Cloud, so motion, audio sync, detail, and typography can be judged on equal footing.
Live action cinematic night market scene, roughly 10 seconds. Open on a low angle tracking shot gliding past steaming stalls as a young wok chef in a soaked tank top tosses noodles, a column of flame bursting up from the wok and lighting his face orange. The camera whip pans right into a tight macro of shrimp searing and curling in the oil, droplets scattering, then pulls back and cranes up over the counter as he catches the spinning wok and plates the noodles in one continuous motion. Final beat: a scruffy market cat springs onto the counter, snatches a shrimp and bolts into the crowd while the chef spins around laughing, the camera breaking into a fast handheld chase behind the cat. Wet asphalt reflecting red lantern light, drifting steam, shallow depth of field, gritty documentary film look with fine grain. Audio: roaring gas burner, sizzling oil, market chatter and clattering ladles, a rising drum pattern that lands exactly on the cat's leap. 16:9 aspect ratio.
Generated with Grok Imagine Video v1.5 on Atlas Cloud
Generated with Veo3.1 on Atlas Cloud
Generated with Grok Imagine Video on Atlas Cloud
Extreme macro photograph shot at tabletop level, camera lens resting flat on the wooden desk surface at the eye height of a 1:64 scale figure, looking horizontally across a miniature modeling workbench in the exact instant the illusion of scale becomes real. In the mid-ground, a wave of blue-tinted epoxy resin has been frozen mid-curl, its surface just misted with a water sprayer so genuine droplets bead and cling to the glossy resin; three tiny hand-painted surfers in bright orange swim trunks ride down the face of the frozen wave, one crouched low with an arm dragging through the resin, and a spray of real water beads flings off the wave's lip, caught sharp in mid-air with tiny burst highlights on every droplet edge. Behind them, softly out of focus, the hand of a young woman modelmaker hovers in the air holding fine steel tweezers, suspended and enormous like a giant who has wandered into the world she built; only the lower half of her face is visible at the top of the frame, the corner of her mouth curved into a small delighted smile, all of it dissolved into creamy bokeh. Late-afternoon hard sunlight cuts through venetian blinds, painting a row of crisp diagonal light bars and clean shadows across the bare wood and over the resin sea, striping the surfers' shoulders. A single dropped, out-of-focus miniature part — a stray plastic sprue fragment — sits in the extreme foreground as a blurred occluding layer at the lower left, separating planes; strips of blue masking tape edge the resin sea, and negative space is left on the right side for the hovering tweezers. Complementary orange-and-blue palette: cool blue resin and blue tape against warm orange trunks, grounded by the warm honey tone of the oak desk. Shot on 100mm f/2.8 macro lens, ultra-shallow depth of field with buttery focus falloff, telephoto compression from an ultra-low camera position, visible fine film grain, true photographic realism, no illustration, no CGI look. Wide 16:9 aspect ratio, full-bleed horizontal composition.

Generated with Grok Imagine Image 2.0 on Atlas Cloud

Generated with Seedream v5.0 Pro on Atlas Cloud

Generated with Grok Imagine Image Quality on Atlas Cloud
Every workflow below runs on one OpenAI-compatible key, so a team can move a concept through Grok Imagine API image drafts, reference edits, video with synchronized audio, and localized narration without changing stacks.
Image 2.0 returns 1K drafts in roughly ten seconds on the low quality tier and up to four variants per call. Design teams scan many directions early, then re-render the winner at 2K.
Feed up to three reference images into Grok Imagine Image 2.0 Edit and describe the change in plain language. Ecommerce teams swap backgrounds, restyle packaging, and hold a single product consistent across a catalog.
xAI TTS v1 converts scripts into expressive speech across twenty languages and more than eighty voices with sub-second latency. Product teams pair it with generated video for localized ads, tutorials, and in-app narration.
Grok Imagine Video v1.5 writes synchronized music, effects, and dialogue alongside picture in clips up to fifteen seconds at 1080p. Marketers get a finished spot in one pass, with no separate audio session.
Need the same character in every shot? Reference-to-video accepts one to seven reference images plus an optional reference voice, so series creators keep identity and vocal tone steady across a fifteen-second clip.
Existing footage can be restyled through natural language editing or continued with a two to ten second extension that matches the source. Post teams fix props, wardrobe, and pacing without returning to set.
See how models from different providers stack up — compare performance, pricing, and unique strengths to make an informed decision.
| Model | Reference Image Limit | Output Num | Resolution | Aspect Ratio |
|---|---|---|---|---|
| Grok Imagine Image Quality | 8 | 1~4 | 2K, 1K | Auto, 1:1, 3:2, 2:3, 3:4, 4:3, 9:16, 16:9, 9:19.5, 19.5:9, 9:20, 20:9, 1:2, 2:1 |
| Nano Banana 2 | 14 | 1 | 4K, 2K, 1K | 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 |
| Nano Banana Pro | 10 | 1 | 4K, 2K, 1K | 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 |
| Seedream 5.0 Lite | 14 | 1~15 | 2K~4K+ | 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 |
| Qwen-Image | 3 | 1~6 | 512P~2K | Width[512, 2048]px, Height[512, 2048]px |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced Grok Imagine models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run Grok Imagine, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
The Grok Imagine API is Atlas Cloud's unified access point for xAI's Grok Imagine generation stack, covering image creation, image editing, video, and speech. One key reaches every mode, including the newest Grok Imagine Image 2.0 Text-to-Image and Grok Imagine Image 2.0 Edit models released on August 7, 2026. Billing is pay-as-you-go per successful generation, so you pay per call with no subscription.
Three image generations sit side by side: Grok Imagine Image at $0.02 per image, Grok Imagine Image Quality at $0.05, and Grok Imagine Image 2.0 from $0.04, each paired with its own edit endpoint. Video runs through Grok Imagine Video at $0.05 per second and Grok Imagine Video v1.5 at $0.08 per second, with extend and edit endpoints alongside them. xAI TTS v1 completes the suite with more than 80 voices across 20 languages.
Image 2.0 plans typography and layout the way a designer would, so dense compositions and small text stay legible instead of dissolving into texture. It also exposes a selectable quality tier, letting you trade render time against fidelity on the same prompt, which the earlier Image and Image Quality models do not offer. Both the Text-to-Image and Edit variants share that behavior.
Pricing is per successful generation with no minimum spend. Grok Imagine Image 2.0 is $0.04 per image at low quality and 1K, $0.06 at low quality with 2K or medium quality with 1K, and $0.08 at medium quality with 2K, while each input image on the Edit endpoint adds $0.01. For the other image models, Grok Imagine Image is $0.02 per image and Grok Imagine Image Quality is $0.05, and video starts at $0.05 per second.
Create an Atlas Cloud API key, then post your prompt and a model id such as xai/grok-imagine-image-2.0/text-to-image to the image generation endpoint. Rendering is asynchronous, so the call returns a request id that you poll on the prediction endpoint until its status moves from processing to completed. Because every Grok Imagine model follows the same pattern, adding video or speech later means changing one string. Start building today.
Outputs render at 1K (1024x1024) or 2K (2048x2048) across 14 aspect ratios, from square 1:1 and widescreen 16:9 to tall 9:20 and ultrawide 20:9 banner shapes. A single call can return up to four images, and prompts can run to 8,000 characters. On the Edit endpoint the aspect ratio defaults to auto and follows your first input image unless you set it explicitly.
Up to three source images per request, supplied as public URLs or base64 data URIs. When you pass more than one, cite them in the prompt as <IMAGE_0>, <IMAGE_1>, and <IMAGE_2> so the model knows which asset each instruction applies to. Each input image adds $0.01 to the call, and you can request one to four edited variants at a time.
Yes. Grok Imagine Video v1.5 produces native, synchronized audio in the same pass as the picture, so music, sound effects, and dialogue land in step with the motion instead of requiring a second pipeline. Its reference-to-video mode accepts an optional reference voice alongside one to seven reference images. Clips run up to 15 seconds, reaching 1080p on text-to-video and image-to-video.
Choose Image 2.0 when the frame carries typography, layout, or multi-part detail that has to hold together, and when you want direct control over the speed and fidelity trade-off. Grok Imagine Image Quality stays a simple flat rate of $0.05 per image with the same 1K and 2K output and 14 aspect ratios. If volume matters more than fine text, the original Grok Imagine Image at $0.02 per image remains the most economical route.
Medium is the default quality tier and takes roughly 84 seconds at 1K because it targets the model's best output. Setting quality to low cuts that to around 10 seconds, roughly 8x faster, at $0.04 rather than $0.06 for a 1K image. Draft and iterate at low quality, then rerun only the winning prompts at medium and 2K once the composition is settled.
Guides, tutorials, and product updates to help you get the most out of Atlas Cloud.