


The MAI Image 2.5 API delivers Microsoft's photorealistic image generation and editing family through one endpoint, spanning text-to-image and editing in standard and Flash variants. Alongside reliable in-image text, it produces natural portraits and applies visual reasoning to keep scene layout coherent. Atlas Cloud offers pay-as-you-go pricing from $0.03 per image, Day-0 access, and one OpenAI-compatible key.
Atlas Cloud provides you with the latest industry-leading creative models.
Four endpoints span text-to-image generation and image editing, each offered in a standard and a Flash tier with transparent per-call pricing.
| Modality | Description |
|---|---|
| MAI Image 2.5 API (Text to Image) | Turn natural language prompts into high-quality, visually rich images with Microsoft's flagship text-to-image model. It targets marketing visuals, e-commerce photography, and commercial design work that needs production-ready output. Priced at $0.05 per image. |
| MAI Image 2.5 Flash (Text to Image) | When throughput and budget matter as much as fidelity, the Flash variant generates images on the same diffusion-based architecture at a lower cost of $0.03 per image. It fits high-volume generation and rapid prototyping pipelines that need consistent quality without inflating spend. |
| MAI Image 2.5 Edit API (Image Editing) | Feed an existing image together with a natural language instruction to apply precise, controllable edits instead of regenerating from scratch. The endpoint handles object removal, element replacement, and localized adjustments while keeping the surrounding composition intact, billed at $0.058 per edit. |
| MAI Image 2.5 Flash Edit (Image Editing) | Teams running high-throughput refinement pipelines get the same instruction-driven editing capability as the standard Edit model at a reduced cost of $0.038 per edit. This makes bulk catalog clean-up and large-batch asset updates practical while keeping per-operation spend low. |
Combining advanced models with Atlas Cloud's GPU-accelerated platform delivers unmatched speed, scalability, and creative control for image and video generation.

MAI-Image-2.5 generates expressive, natural-looking portraits with accurate facial structure, lighting, and skin texture from text prompts. The model renders film-quality aesthetics with consistent lighting that matches the described scene. It is designed for editorial, branding, and commercial campaigns where human-centric imagery needs to look finished without post-processing.

MAI-Image-2.5 offers enhanced reliability for text generation within images, handling product labels, signage, headlines, and branded copy with correct spacing and legibility. This addresses a consistent weak point in most image generation models and makes it practical for packaging mockups and advertising assets where readable text is required in the output. It is the right choice for design workflows where in-image text accuracy is non-negotiable.

The MAI-Image-2.5 Edit endpoint performs targeted modifications to specific image regions: removing unwanted elements, replacing or recoloring objects, updating text in existing signage, filling missing areas, and cleaning up visual defects like blur and noise. Edits maintain coherence and composition throughout, leaving untouched regions visually intact. It is the go-to tool for product refinement, catalog clean-up, and marketing asset updates.

MAI-Image-2.5 is built specifically for commercial and professional design applications, supporting branding, product mockups, and campaign-ready content from text prompts. The model maintains layout and composition integrity during both generation and editing, producing assets that are ready for use in advertising and product campaigns. It is the standard solution for design teams producing commercial visuals at scale.

MAI-Image-2.5 applies visual reasoning to understand spatial relationships, object placement, and lighting coherence across the full image. This makes it reliable for generating scenes where multiple elements need to coexist naturally, and for editing tasks where a modification needs to respect the surrounding context. It is suited for product-in-scene visualization and any workflow where contextual accuracy in the output matters.
The same prompt, generated by Mai Image 2.5 and other leading image models: product and cinematic lifestyle advertisement
Editorial macro still-life photograph: a row of old handmade stained-glass apothecary and perfume bottles of uneven heights — squat, tall, bulbous — standing on a weathered lime-plaster windowsill, their frosted glass streaked with tiny trapped air bubbles. The decisive moment: a bare human hand tilts one amber bottle and a thin clean thread of clear liquid hangs suspended in mid-air, caught and lit mid-fall, its surface tension and beading edge razor-sharp. Behind it the other bottles bend a low golden-hour sun into a scatter of overlapping jewel-colored caustics — amber, deep bottle-green, rose pink — crawling slowly across the coarse, calcified gray plaster. Low-angle golden-hour side-backlight rakes through the glass, throwing directional focused caustics and glowing rim-light along each bottle's silhouette. Composition built on the graphic repetition of the bottle array set against a wide expanse of neutral gray wall as negative space, shallow depth of field compressing the row, the pouring hand and glinting water thread as the sharp focal point. Limited gem-toned palette of amber, dark green and rose pink playing against the muted gray wall — restrained, never gaudy. Hyper-real macro detail: frosted glass grain, bubbles, chalky wall texture, the taut skin of the liquid, faint dust motes drifting in the beam. Quiet, with a trace of the magical. Shot on 100mm macro lens, natural window light, editorial product-photography finish. 16:9 aspect ratio.

Mai Image 2.5

Flux.2

Qwen Image 2.0
Late-night street noodle stall, the decisive split-second as a lamian master pulls a single ball of dough into a fan of dozens of perfectly parallel, hair-thin noodle strands suspended mid-air, the strands flaring outward like an unfolding hand-held fan while a burst of loose flour and rising steam explode at the same instant — this is action caught in motion, not a pose. His face is locked in total concentration, muscles taut with controlled force, brows furrowed, a bead of sweat at his temple; behind him, blurred out-of-focus onlookers hold their breath in anticipation. Realistic street documentary photography, telephoto lens language. Lit by the stall's single warm tungsten bulb as a hard side-backlight that rims every translucent noodle strand with glowing edges, catches the airborne flour dust as floating sparks and turns the billowing white steam luminous; the cold blue neon of the night street behind melts into soft bokeh orbs. Telephoto multi-layer depth compression flattens the master's hands, his focused face and the cascade of noodles into a single plane, the dense parallel strands acting as both a repeating graphic element and natural leading lines drawing the eye across the frame. Cool-warm color balance — foreground amber tungsten glow against the deep cold blue of the night — restrained and never garish. Rich tactile texture: gritty flour particles, faint sheen of gluten, the oil-soaked worn wooden cutting board, sweat on skin, and a subtle motion blur on the flying strands and drifting powder. Fine film grain, shallow depth of field, no over-rendering, no plastic skin, honest natural tonality. Shot on 135mm telephoto, wide aperture, candid reportage feel. 16:9 aspect ratio.

Mai Image 2.5

Flux.2

Qwen Image 2.0
From e-commerce catalogs and ad creative to packaging, spokesperson imagery, and product-in-scene visualization, the MAI Image 2.5 API supports the commercial design work teams ship every day.
Generate product shots on varied backgrounds from one reference, then recolor and clean defects across a full SKU range via the Edit endpoint. A complete variant set costs less than a single studio reshoot.
Performance marketers produce banner, social, and display variations with accurate text overlays and brand-consistent layouts. Because the Flash variant runs at $0.03 per image, testing dozens of concepts before scaling winners stays cheap.
Packaging mockups carry legible typography baked directly into box, bottle, and shelf-signage artwork rather than placed by hand. When wording changes, the Edit endpoint revises names, prices, or seasonal copy without rebuilding the layout.
A recognizable face, wardrobe, and overall identity stay stable across poses and layouts through successive edits. This keeps recurring campaign characters and ongoing social series looking coherent wherever they appear.
Reflections, stray objects, and motion blur come out cleanly while inpainting fills missing regions and the surrounding composition stays intact. At $0.058 per edit, cleaning thousands of legacy photos becomes a routine batch job.
The model reasons about lighting, scale, and spatial relationships to place products into lifestyle scenes with believable shadows and perspective. Concept and prototyping teams preview staging ideas well before booking a physical shoot.
Compare the MAI Image 2.5 API head to head with leading text-to-image models from Google, ByteDance, and Alibaba on provider, price, resolution, and text rendering.
| Model | Provider | Starting Price (per image) | Max Resolution | In-image Text Rendering | Cost-Optimized Fast Tier |
|---|---|---|---|---|---|
| MAI-Image-2.5 (Text to Image) | Microsoft | $0.05 | Up to ~1 MP (1360 px max per side) | √ | √ |
| MAI-Image-2.5 Flash (Text to Image) | Microsoft | $0.03 | Up to ~1 MP (1360 px max per side) | √ | √ |
| Nano Banana 2 (Text-to-Image) | $0.08 | Up to 4K | √ | √ | |
| Seedream v4.5 | ByteDance | $0.04 | Up to 2048×2048 | √ | - |
| Qwen Image 2.0 (Text-to-image) | Alibaba | $0.035 | Up to 2048×2048 | √ | - |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced MAI Image 2.5 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run MAI Image 2.5, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
The MAI Image 2.5 API gives developers programmatic access to Microsoft's MAI-Image-2.5, a photorealistic image generation and editing model built for commercial design work. It covers both text-to-image generation and instruction-based image editing, each offered in a standard and a Flash variant. On Atlas Cloud you reach all four endpoints through one OpenAI-compatible key with pay-as-you-go pricing from $0.03 per image.
Both variants share the same diffusion architecture, photorealism, and text rendering quality. Flash is the cost-optimized option at $0.03 per image for generation and $0.038 per edit, while the standard model runs $0.05 per image and $0.058 per edit. Reach for Flash on high-volume generation and rapid prototyping, and keep the standard model for work where maximum fidelity matters most.
Pricing on Atlas Cloud is pay-as-you-go with no subscription. Standard text-to-image is $0.05 per image and standard editing is $0.058 per edit, while the Flash variants drop to $0.03 per image and $0.038 per edit. You are billed per successful call, so a batch of variations costs only the flat per-image fee multiplied by the number of outputs.
Sign up on Atlas Cloud, create one API key, and point your existing OpenAI-compatible client at the MAI-Image-2.5 endpoint. Send a text prompt for generation, or an image plus an instruction for editing, and the response returns URLs to the finished images. Day-0 access means the model is available the moment it ships, with no waitlist. Start building today.
You set output dimensions with a size parameter in width by height format, defaulting to 1024 by 1024. Each side can range from 768 to 1360 pixels up to a ceiling of roughly one megapixel in total, which covers square, portrait, and landscape framing. The standard and Flash generation variants share the same resolution limits.
The Edit endpoint takes a single source image, supplied as a URL or base64 in JPEG or PNG, together with a natural language instruction. It performs targeted changes such as removing objects, replacing elements, recoloring, updating text in signage, and cleaning up defects, while leaving untouched regions intact. Because the model reasons about scene structure and lighting, edits stay coherent with the surrounding composition.
Text rendering is one of the model's biggest gains, and Microsoft reports its largest quality jump over the previous generation in exactly this category. It reliably produces product labels, signage, headlines, and branded copy with correct spacing and legibility. That reliability makes it practical for packaging mockups and advertising assets where readable in-image text is non-negotiable.
On the public Arena leaderboards, MAI-Image-2.5 launched at No. 2 for image editing and No. 3 for text-to-image generation. Its strengths lie in photorealistic portraits, commercial-grade text rendering, and identity-consistent editing rather than stylized or artistic output. For teams producing brand-ready visuals, that positioning makes it a strong alternative to general-purpose generators.
Microsoft built and tuned MAI-Image-2.5 specifically for commercial and professional design work, including packaging, product mockups, signage, and brand-forward visuals. The model is intended for production use across advertising, e-commerce, and marketing pipelines. For the exact usage rights that apply to your account, review the current Atlas Cloud terms before publishing generated assets.
Join the Discord community for the latest model updates, prompts, and support.