



Grok Imagine API 涵盖 xAI 的图像、视频和语音模型,包括 Image 2.0、Video 1.5 和 xAI TTS v1。支持生成跨 14 种宽高比的 1K 或 2K 静帧、生成最长 15 秒的 1080p 视频、通过最多 7 张参考图像引导镜头、或用 20 种语言进行配音。Atlas Cloud 在单一端点上运行所有功能,采用按量计费,起价为每张图像 $0.02、每秒 $0.05。立即开始使用。
Grok Imagine 由 xAI 开发。Atlas Cloud(由 Atlas Cloud AI LLC 运营)仅提供接入服务,并不拥有该模型。所有商标均归其各自所有者所有。
Atlas Cloud 为您提供最新的行业领先创意模型。
根据输入类型、输出时长、分辨率和每次调用价格,选择合适的 Grok Imagine API 模型。
| 模态 | 描述 |
|---|---|
| Grok Imagine Image 2.0 T2I API (Text to Image) | 输入最多 8000 个字符的提示词,该端点返回 1K 或 2K 分辨率的 1 到 4 张图像。低质量或中等质量级别让你在渲染耗时和出图速度之间权衡,每张图像定价为 $0.04。适用于概念探索、社交创意和批量视觉生成等大体量场景。 |
| Grok Imagine Image 2.0 Edit API (Image to Image) | 向端点提交最多 3 张参考图像和自然语言指令,编辑结果以你指定或自动判断的宽高比返回 1K 或 2K 版本。同样提供低质量和中等质量两个级别,每张图像 $0.04。产品重新设计、背景替换和快速设计变体都可以通过一次调用完成。 |
| Grok Imagine Image Quality T2I API (Text to Image) | 当细节质量至关重要时,该端点将文本提示转换为 1K 或 2K 的精品静图,每次请求返回最多 4 个变体,每张 $0.05。支持正方形、纵向、横向和超宽等多种宽高比。适用于主视觉、广告创意和品牌级产品渲染。 |
| Grok Imagine Image Quality Edit API (Image to Image) | 输入 1 到 8 张源图像和编辑指令,接收 1K 或 2K 的修改版本,每张 $0.05。多图引用可以通过 <IMAGE_0> 和 <IMAGE_1> 的方式内联指定,因此可以在单个指令中组合主体和风格。适用于活动更新、风格迁移和合成编辑。 |
| Grok Imagine Image T2I API (Text to Image) | 原始的文本转图像端点保持 1K 和 2K 输出及 4 张图像批处理,价格是产品线中最低的,每张 $0.02。这使其成为大体量管道、缩略图生成和快速提示词测试的实用默认选择。 |
| Grok Imagine Image Edit API (Image to Image) | 需要大规模轻量级编辑?该端点接受 1 到 8 张图像和文本指令,返回最多 4 个编辑结果,1K 或 2K,每张 $0.02。目录清理、季节性变体和迭代设计传递都能保持低成本。 |
| Grok Imagine Video v1.5 T2V API (Text to Video) | 仅通过提示词即可生成 1 到 15 秒的视频,包含原生同步音频,支持 480p、720p 或 1080p,覆盖 7 种宽高比。按输出秒数计费,$0.08/秒。适合预告片、社交短视频和需要音频的叙事场景。 |
| Grok Imagine Video v1.5 I2V API (Image to Video) | 提供起始帧和动作提示词,v1.5 可将其动画化至 15 秒,分辨率达 1080p,音频在同一过程中生成。成本为 $0.08/秒。适用于产品旋转、肖像动画和静止到动作的活动资产。 |
| Grok Imagine Video v1.5 R2V API (Reference to Video) | 1 到 7 张参考图像指导屏幕上出现的人物、对象和风格,同时从 26 个预设中选择最多 3 种声音作为语音音频。输出达 15 秒,480p 或 720p,$0.08/秒。最适合人物一致的故事讲述、虚拟试穿和品牌吉祥物。 |
| Grok Imagine Video T2V API (Text to Video) | 文本提示变成 1 到 15 秒的视频片段,480p 或 720p,支持 7 种宽高比,覆盖横向、正方形和纵向。$0.05/秒,是社交内容和快速概念板的经济选择。 |
| Grok Imagine Video I2V API (Image to Video) | 以静止图像作为首帧,描述你想要的动作,视频片段延伸至 15 秒,480p 或 720p。价格保持 $0.05/秒,使产品照片和肖像的批量动画化成本保持在可承受范围内。 |
| Grok Imagine Video R2V API (Reference to Video) | 1 到 7 张参考图像塑造屏幕上出现的人物和对象,无需固定开场帧。视频片段最长 10 秒,480p 或 720p,$0.05/秒。当身份和风格必须在各镜头间保持一致时使用。 |
| Grok Imagine Video Extend API (Video Extension) | 现有的 2 到 15 秒 mp4 视频可再延伸 2 到 10 秒,由指示后续内容的提示词指导。输出与源视频匹配,最高 720p,按 $0.07/秒计费。便于将短视频延伸到所需的广告长度。 |
| Grok Imagine Video Edit API (Video to Video) | 自然语言指令重写现有 mp4,原始场景、动作和构图保持不变。输出保持源视频时长,上限 8.7 秒,按输入视频计费,$0.07/秒。道具添加、服装更换和重新设计无需重新拍摄。 |
| xAI TTS v1 API (Text to Speech) | 最多 15000 个字符的文本以极低延迟转换为自然语音,支持 20 种语言,可自动检测语言。可调选项:播放速度 0.7x 到 1.5x、支持 mp3、wav 和 pcm 等编码格式、采样率最高 48 kHz。适合旁白、IVR 提示和应用内叙述。 |
探索 Grok Imagine API 提供的强大功能,涵盖从支持多语言文本的 2K 图像生成,到具备原生同步音频及多种创意模式的多模态视频生成。

Grok Imagine Image Quality API 提供高达 2K 分辨率的图像生成,确保每次输出都具有极其清晰的细节。通过在缩放时保留细腻的纹理和复杂的构图,用户可以制作出即使在超大画幅下展示也依然清晰的视觉内容。它是主视觉图、广告创意和品牌级产品渲染的终极解决方案。

Grok Imagine Image Quality API 在生成的图像中直接提供支持多语言的同类最佳文本渲染功能。通过准确还原任何语言的排版、文字符号和字符,用户可以将清晰可读的文案嵌入到视觉作品中,而无需进行手动后期编辑。这是广告创意、本地化营销活动和品牌级视觉效果的终极解决方案。

Grok Imagine API 能够生成具有自然光照、丰富纹理和逼真物理效果的写实图像输出。通过模拟真实世界的光学原理和材质表现,用户可以生成在视觉上与专业摄影无法区分的图像。它是产品渲染、主图和高端品牌视觉效果的终极解决方案。

Grok Imagine Image Quality API 支持更精准的提示词遵循,以及由参考输入驱动的高级图像编辑功能。通过解析详细指令并匹配上传参考图中的风格特征,用户可以以极高的精度完善和重塑视觉效果。它是广告创意、产品渲染和一致品牌级视觉效果的终极解决方案。

自动为每个片段生成同步的音乐、音效和对话,确保音频与画面动态在一次处理中保持对齐。片段无需单独的音频处理步骤,生成后即可直接使用。

它在单一套件中涵盖了文本生成视频、图像生成视频、参考生成视频以及视频编辑功能。您可以在生成和编辑任务之间无缝切换,而无需更换模型或集成。

Grok Imagine Video API 能够生成自然流畅的运动效果,并在不同帧之间保持稳定的物理特性和一致的主体。这减少了较长片段中的闪烁和伪影,使角色和场景从头到尾保持连贯。
Each row runs the identical prompt through the Grok Imagine API and two rival models on Atlas Cloud, so motion, audio sync, detail, and typography can be judged on equal footing.
Live action cinematic night market scene, roughly 10 seconds. Open on a low angle tracking shot gliding past steaming stalls as a young wok chef in a soaked tank top tosses noodles, a column of flame bursting up from the wok and lighting his face orange. The camera whip pans right into a tight macro of shrimp searing and curling in the oil, droplets scattering, then pulls back and cranes up over the counter as he catches the spinning wok and plates the noodles in one continuous motion. Final beat: a scruffy market cat springs onto the counter, snatches a shrimp and bolts into the crowd while the chef spins around laughing, the camera breaking into a fast handheld chase behind the cat. Wet asphalt reflecting red lantern light, drifting steam, shallow depth of field, gritty documentary film look with fine grain. Audio: roaring gas burner, sizzling oil, market chatter and clattering ladles, a rising drum pattern that lands exactly on the cat's leap. 16:9 aspect ratio.
Generated with Grok Imagine Video v1.5 on Atlas Cloud
Generated with Veo3.1 on Atlas Cloud
Generated with Grok Imagine Video on Atlas Cloud
Extreme macro photograph shot at tabletop level, camera lens resting flat on the wooden desk surface at the eye height of a 1:64 scale figure, looking horizontally across a miniature modeling workbench in the exact instant the illusion of scale becomes real. In the mid-ground, a wave of blue-tinted epoxy resin has been frozen mid-curl, its surface just misted with a water sprayer so genuine droplets bead and cling to the glossy resin; three tiny hand-painted surfers in bright orange swim trunks ride down the face of the frozen wave, one crouched low with an arm dragging through the resin, and a spray of real water beads flings off the wave's lip, caught sharp in mid-air with tiny burst highlights on every droplet edge. Behind them, softly out of focus, the hand of a young woman modelmaker hovers in the air holding fine steel tweezers, suspended and enormous like a giant who has wandered into the world she built; only the lower half of her face is visible at the top of the frame, the corner of her mouth curved into a small delighted smile, all of it dissolved into creamy bokeh. Late-afternoon hard sunlight cuts through venetian blinds, painting a row of crisp diagonal light bars and clean shadows across the bare wood and over the resin sea, striping the surfers' shoulders. A single dropped, out-of-focus miniature part — a stray plastic sprue fragment — sits in the extreme foreground as a blurred occluding layer at the lower left, separating planes; strips of blue masking tape edge the resin sea, and negative space is left on the right side for the hovering tweezers. Complementary orange-and-blue palette: cool blue resin and blue tape against warm orange trunks, grounded by the warm honey tone of the oak desk. Shot on 100mm f/2.8 macro lens, ultra-shallow depth of field with buttery focus falloff, telephoto compression from an ultra-low camera position, visible fine film grain, true photographic realism, no illustration, no CGI look. Wide 16:9 aspect ratio, full-bleed horizontal composition.

Generated with Grok Imagine Image 2.0 on Atlas Cloud

Generated with Seedream v5.0 Pro on Atlas Cloud

Generated with Grok Imagine Image Quality on Atlas Cloud
探索使用 Grok Imagine API 可以构建的内容,从照片级逼真的品牌视觉效果和多语言广告海报,到产品视频展示、人像动画以及基于参考的编辑。
Grok Imagine 图像质量 API 使创作者和开发者能够生成具有自然光照、丰富纹理和真实物理效果的逼真视觉效果。该 API 是追求工作室级别输出的营销团队和设计工作室的理想之选,可渲染清晰的 2K 分辨率和栩栩如生的材质细节——支持生成主图、广告创意和高端产品渲染图。
对于全球分发的创意内容,Grok Imagine Image Quality API 能够生成具备同类最佳文本渲染效果、准确的多语言排版以及直接在艺术作品中清晰集成字符的图像。此用例适用于广告代理商、本地化专家和品牌设计师,帮助他们制作需要将清晰易读、符合品牌形象的文案嵌入到最终图像中的视觉效果。
Grok Imagine Image Quality API 赋能设计师,通过更严格的提示词遵循、基于参考的输入以及精准的构图控制,对现有视觉内容进行优化和重塑。该 API 能够跨越多次编辑保持风格一致性,是迭代式创意生产和品牌一致性工作流的理想之选——支持概念细化、设计变体生成以及为商业活动打造精细的最终资产。
Grok Imagine Video Text-to-Video API 使创作者和开发者能够仅凭单一文本提示生成电影级视频片段,并配有原生音频和高达 720p 的分辨率。该 API 是追求生产级视频输出的营销团队和内容工作室的理想之选,它能渲染动态运动、自然的摄像机移动和同步音效——为品牌活动、社交媒体内容和沉浸式广告叙事提供支持。
对于希望为静态视觉作品注入生命的创作者而言,Grok Imagine Video 图生视频 API 可将静态图像转化为流畅、逼真的视频片段,并以源图像作为第一帧。该应用场景非常适合电子商务品牌、数字艺术家和广告团队,用于制作需要与原始资产保持视觉连续性的产品动画展示、人像动画和场景生动化内容。
对于需要对现有素材进行精确、定向修改的后期制作团队和创意机构,Grok Imagine Video Edit API 可将自然语言指令应用于现有视频,同时保留原始场景、运动和构图。该应用场景适合视频剪辑师、营销制作人和完善营销活动素材的品牌团队——能够在不破坏原有视频结构的情况下,实现道具添加、服装更换和视觉风格重塑。
查看不同厂商的模型表现 — 对比性能、价格和独特优势,做出明智决策。
| 模型 | 参考图片限制 | 输出数量 | 分辨率 | 宽高比 |
|---|---|---|---|---|
| Grok Imagine Image Quality | 8 | 1~4 | 2K, 1K | Auto, 1:1, 3:2, 2:3, 3:4, 4:3, 9:16, 16:9, 9:19.5, 19.5:9, 9:20, 20:9, 1:2, 2:1 |
| Nano Banana 2 | 14 | 1 | 4K, 2K, 1K | 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 |
| Nano Banana Pro | 10 | 1 | 4K, 2K, 1K | 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 |
| Seedream 5.0 Lite | 14 | 1~15 | 2K~4K+ | 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 |
| Qwen-Image | 3 | 1~6 | 512P~2K | Width[512, 2048]px, Height[512, 2048]px |
几分钟即可上手 — 按照以下简单步骤,通过 Atlas Cloud 平台集成和部署模型。
在 atlascloud.ai 注册并完成验证。新用户可获得免费额度,用于探索平台和测试模型。
将先进的 Grok Imagine 模型与 Atlas Cloud 的 GPU 加速平台相结合,提供无与伦比的性能、可扩展性和开发体验。
低延迟:
GPU 优化推理,实现实时响应。
统一 API:
一次集成,畅用 Grok Imagine、GPT、Gemini 和 DeepSeek。
透明定价:
按 Token 计费,支持 Serverless 模式。
开发者体验:
SDK、数据分析、微调工具和模板一应俱全。
可靠性:
99.99% 可用性、RBAC 权限控制、合规日志。
安全与合规:
SOC 2 Type II 认证、HIPAA 合规、美国数据主权。
Grok Imagine API 是 Atlas Cloud 为 xAI 的 Grok Imagine 生成栈提供的统一接入点,涵盖图像创建、图像编辑、视频和语音。一个密钥即可访问所有模式,包括 2026 年 8 月 7 日发布的最新 Grok Imagine Image 2.0 Text-to-Image 和 Grok Imagine Image 2.0 Edit 模型。按成功生成计费,每次调用付费,无需订阅。
三款图像生成模型并列提供:Grok Imagine Image 每张 0.02 美元,Grok Imagine Image Quality 每张 0.05 美元,Grok Imagine Image 2.0 起价 0.04 美元,各配独立编辑端点。视频通过 Grok Imagine Video 每秒 0.05 美元、Grok Imagine Video v1.5 每秒 0.08 美元提供,另有扩展和编辑端点。xAI TTS v1 补全整个套件,支持 20 种语言的 80 多种声音。
Image 2.0 像设计师一样规划排版和布局,使密集构图和小字保持可读,而不会溶解为纹理。它还暴露可选的质量层级,让用户能在同一提示词下用渲染时间换取保真度,这是早期 Image 和 Image Quality 模型所不具备的。Text-to-Image 和 Edit 两个变体均共享此特性。
按成功生成计费,无最低消费。Grok Imagine Image 2.0 低质量 1K 每张 0.04 美元,低质量 2K 或中质量 1K 每张 0.06 美元,中质量 2K 每张 0.08 美元,编辑端点每张输入图像额外加收 0.01 美元。其他图像模型中,Grok Imagine Image 每张 0.02 美元,Grok Imagine Image Quality 每张 0.05 美元,视频起价每秒 0.05 美元。
创建 Atlas Cloud API 密钥,然后将提示词和模型 ID(如 xai/grok-imagine-image-2.0/text-to-image)发送至图像生成端点。渲染为异步操作,调用返回请求 ID,需在预测端点轮询直到状态从 processing 变为 completed。由于所有 Grok Imagine 模型遵循同一模式,后续添加视频或语音只需改一个字符串。立即开始构建。
输出以 1K(1024x1024)或 2K(2048x2048)渲染,支持 14 种宽高比,从方形 1:1、宽屏 16:9 到竖屏 9:20 和超宽 20:9 横幅形状。单次调用最多返回四张图像,提示词最长可达 8,000 个字符。编辑端点的宽高比默认为 auto,跟随首张输入图像,除非显式设置。
每次请求最多三张源图像,以公开 URL 或 base64 数据 URI 提供。传入多张时,在提示词中以 <IMAGE_0>、<IMAGE_1>、<IMAGE_2> 引用,使模型明确每条指令对应哪个素材。每张输入图像使调用增加 0.01 美元,一次可请求一到四个编辑变体。
是。Grok Imagine Video v1.5 在与画面同一通道中生成原生同步音频,音乐、音效和对话与动作同步,无需额外管道。其参考视频模式接受可选参考语音,以及一到七张参考图像。片段最长 15 秒,文本转视频和图像转视频均可达 1080p。
当画面包含必须保持完整的排版、布局或多部分细节,且需要直接控制速度与保真度权衡时,选择 Image 2.0。Grok Imagine Image Quality 保持每张 0.05 美元的固定费率,提供相同的 1K 和 2K 输出及 14 种宽高比。若量比精细文字更重要,原始 Grok Imagine Image 每张 0.02 美元仍是最经济的选择。
中等质量是默认等级,1K 下约需 84 秒,因其追求模型最佳输出。将质量设为低质量可缩短至约 10 秒,提速约 8 倍,1K 图像费用从 0.06 美元降至 0.04 美元。建议先用低质量起草和迭代,构图确定后再以中等质量和 2K 重跑获胜提示词。
指南、教程与产品动态,助你充分发挥 Atlas Cloud 的价值。