दुनिया भर में सबसे कम कीमतों पर Seedance 2.0 Mini & Fast API — आधिकारिक कीमत पर 68% तक की छूट
Grok Imagine Image 2.0 Text-to-Image
टेक्स्ट-से-इमेज

Grok Imagine Image 2.0 Text-to-Image API by xAI

xai/grok-imagine-image-2.0/text-to-image
Text-to-image

xAI Grok Imagine Image 2.0 generates polished visuals from natural-language prompts at 1K or 2K resolution, with 14 aspect ratios and selectable low/medium quality tiers.

मॉडल की तुलना करें

Grok Imagine Image 2.0 Text-to-Image को xAI द्वारा विकसित किया गया है। Atlas Cloud (Atlas Cloud AI LLC द्वारा संचालित) इस तक पहुँच प्रदान करता है, इसका स्वामी नहीं है। सभी ट्रेडमार्क उनके संबंधित स्वामियों की संपत्ति हैं।

1. Introduction

Grok Imagine Image 2.0 is xAI's image generation and editing model, released on August 7, 2026 and positioned around a single goal: producing images that hold up in real creative work rather than as one-off novelties. This README applies to the following API model identifiers:

  • xai/grok-imagine-image-2.0/text-to-image
  • xai/grok-imagine-image-2.0/edit

Developed by xAI as the successor to Grok Imagine Image Quality, Image 2.0 shipped first as the new Quality Mode on grok.com/imagine and the Grok iOS and Android apps, with general API availability following shortly after. Per the official announcement, the model was built to follow instructions closely down to fine detail, to plan typography and layout the way a designer would so that dense multi-part visuals hold together and small text stays sharp, and to preserve what the user supplies across successive generations and edits.

The model is exposed through two API variants that share the same underlying weights and differ only in input schema and conditioning path. The xai/grok-imagine-image-2.0/text-to-image variant produces images from a text prompt alone. The xai/grok-imagine-image-2.0/edit variant applies prompt-driven modifications to between one and three supplied source images. xAI describes editing as a first-class capability of 2.0 rather than a bolt-on, and the model's Arena standing in image editing is marginally stronger than its standing in pure text-to-image.


2. Key Features & Innovations

  • Designer-Style Typography and Layout Planning: xAI's central claim for 2.0 is that it plans typography and layout rather than treating text as texture. Dense, multi-part compositions — infographics, posters, itineraries, annotated diagrams, title screens — are intended to remain internally coherent with small type staying legible, historically the weakest area of generative image models.

  • Selectable Quality Tiers: A quality parameter selects between low and medium rendering. The distinction is substantial in practice rather than cosmetic: low returns in roughly a tenth of the time and at a lower per-image price, while medium — the default — produces the model's best output. Choosing the tier per request lets draft iteration and final rendering run against the same endpoint.

  • Editing as a First-Class Capability: Image 2.0 was trained with editing treated as a primary objective rather than a downstream adaptation. This shows in its benchmark position, where it places higher on the Image Edit Arena than on the Text-to-Image Arena.

  • Multi-Image Reference Editing: The edit variant accepts up to three source images in a single request, combining them without manual compositing. References are cited positionally in the prompt as <IMAGE_0>, <IMAGE_1>, <IMAGE_2>, and the output aspect ratio follows the first input unless overridden.

  • Input Preservation Across Iterations: The model is designed to hold onto what the user supplies across repeated generations and edits, supporting the iterative refine-and-re-edit loop that real production work depends on rather than drifting away from the source on each pass.

  • Style-Consistent World Building: Characters, locations, and props generated from separate prompts hold a single coherent visual style across images. xAI frames this as a stepping stone toward full video production workflows, where a consistent cast and set must survive across many shots.

  • Broad Aspect-Ratio Control: Both variants expose 13 fixed aspect ratios spanning 2:1 through 1:2, including ultra-tall and ultra-wide options such as 9:19.5, 19.5:9, 9:20 and 20:9. The edit variant additionally accepts auto, its default, which derives the output ratio from the first source image. On xAI's consumer surface the equivalent capability is presented as Smart Resize.

  • Dual-Resolution Output: Images are produced at 1K (1024×1024) or 2K (2048×2048). Outputs are returned as JPEG or PNG depending on resolution and content, and AtlasCloud rehosts every result to a durable URL rather than passing through xAI's short-lived image links.

  • Batch Generation: Both variants accept a num_images parameter (1–4 on AtlasCloud) to produce multiple candidates per request, supporting creative exploration and A/B selection inside production pipelines.

  • Long Prompt Budget: The model accepts prompts up to 8,000 characters, leaving room for the detailed layout, typography and style direction the model is built to act on.


3. Model Architecture & Technical Details

xAI has not published architecture, parameter count, or training-corpus details for Image 2.0. What the company has stated is the training objective: the model was trained for fidelity across photography, design, and illustration, with editing treated as a first-class capability alongside generation rather than as a secondary fine-tune.

Both API identifiers are served by the same underlying model. The distinction lies entirely in the request schema — text-to-image conditions on a prompt alone, while edit additionally conditions on one to three supplied source images.

API specifications:

ParameterText-to-ImageEdit
Required inputspromptprompt, image_urls
Reference images1–3
num_images1–41–4
aspect_ratio13 options (2:1 to 1:2), default 1:114 options (the same 13 plus auto), default auto
resolution1k / 2k1k / 2k
qualitylow / medium (default medium)low / medium (default medium)
Max prompt length8,000 characters8,000 characters

Quality tier and latency. The quality parameter is the single most consequential setting on this model, because it governs both cost and speed:

Quality1K latency2K latency
low~10 s~16 s
medium (default)~84 s~122 s

Text-to-image at the default medium tier is substantially slower than the Grok Imagine Image Quality generation path, and latency grows with batch size — a four-image 2K request takes roughly 140 seconds. Editing is considerably faster than generation at the same tier. Applications should treat medium text-to-image as an asynchronous operation, and should prefer quality: "low" for draft iteration, interactive previews, and any latency-sensitive path. Note that high is a valid value in the shared xAI image schema but is rejected for this model, which supports only low and medium.

Consumer product versus API surface. Several capabilities described in xAI's launch materials are interactive tools built into the Grok consumer applications rather than API parameters: the magic wand for region-local edits, segmentation-based area selection, background removal to transparency, and the preconfigured workflow templates. The consumer Multi-Ref Editing feature also accepts up to five inputs, whereas the API accepts three. The API surface provides prompt-driven editing of one to three source images.


4. Performance Highlights

xAI reports that Image 2.0 ranks second in the world on both the text-to-image and image-editing leaderboards. Scores below are overall Elo from the Arena Image Edit and Text-to-Image leaderboards as of August 7, 2026; xAI models are listed on Arena under the name SpaceXAI.

Image Edit Arena:

RankModelDeveloperElo
1GPT-Image-2OpenAI1463
2Grok Imagine Image 2.0xAI1439

Text-to-Image Arena:

RankModelDeveloperElo
1GPT-Image-2OpenAI1380
2Grok Imagine Image 2.0xAI1320

Models placing below Image 2.0 on these leaderboards at the time of measurement include Reve 2.1, Meta's Muse-Image, Alibaba's Qwen-Image-3.0-Pro, Google's Gemini image models, and ByteDance's Seedream family. The 24-point gap to first place in image editing is materially narrower than the 60-point gap in text-to-image, consistent with xAI's stated emphasis on editing quality.


5. Use Cases

  • Information-Dense Design: Infographics, posters, editorial layouts, itineraries, and annotated diagrams where multiple text blocks and visual elements must coexist legibly — the workload the model's typography and layout planning specifically targets.

  • Marketing and Product Imagery: Product photography, e-commerce imagery, editorial product posters, merchandise mockups, and UGC-style creative, generated from a brief or derived from an existing product shot via the edit variant.

  • Reference Composition: Merging up to three source images — a subject, a setting, and a prop, for instance — into a single coherent scene without manual masking or compositing.

  • Iterative Creative Refinement: Multi-turn workflows that draft at quality: "low" for fast, inexpensive exploration and re-render the selected candidate at medium, relying on the model's preservation of supplied inputs across passes.

  • Game and Application Assets: Character sprites, props, UI kits, icons, mascots, and streaming emoji, where a consistent visual style must hold across many separately generated pieces.

  • Pre-Production for Video: Establishing a character, their locations, and their props as style-consistent stills that serve as a coherent visual bible before committing to video generation.

  • Photo Editing and Retouching: Prompt-driven modification of an existing image — style transfer, object addition or removal, recomposition, and reimagining — through the edit variant without mask authoring.

समान मॉडल देखें

हर मीडिया AI के लिए एक ही API।

सभी मॉडल एक्सप्लोर करें