Seedance 2.5 Now Live — First on Atlas Cloud
MiniMax API

MiniMax API

As a premier suite of Large Language Models (LLMs) developed by MiniMax AI, MiniMax is engineered to redefine real-world productivity through cutting-edge artificial intelligence. The ecosystem features MiniMax M2.5, which is purpose-built for high-efficiency professional environments, and MiniMax M2.1, a model that offers significantly enhanced multi-language programming capabilities to master complex, large-scale technical tasks. By achieving SOTA performance in coding, agentic tool use, intelligent search, and office workflow automation, MiniMax empowers users to streamline a wide range of economically valuable operations with unparalleled precision and reliability.

Explore the Leading MiniMax

Atlas Cloud provides you with the latest industry-leading creative models.

Compare the MiniMax API Model Lineup

A clear look at what each MiniMax endpoint accepts, produces, and costs, spanning text, coding, agentic workflows, and music generation.

ModalityDescription
MiniMax M3 API (Text)Need a fast agentic model? MiniMax M3 converts text prompts into code, tool calls, and reasoning from a lightweight 10 billion activated parameter core built for low latency and scalability. Standard pricing is $0.6 per million input tokens and $2.4 per million output tokens, with a separate tier applied once a prompt passes 512K tokens. It fits long, tool-heavy agent sessions and modern application development.
MiniMax M2.7 (Text)MiniMax M2.7 turns natural language instructions into production code and multi-step agent actions, using the same 10 billion activated parameter design for high throughput. Standard pricing runs $0.3 per million input tokens and $1.2 per million output tokens, keeping large agentic runs economical. Reach for it when dependable coding and app-building capacity matter more than raw model size.
MiniMax M2.5 (Text)If your workload needs affordable, scalable reasoning, MiniMax M2.5 handles coding, agentic orchestration, and everyday application development from a compact 10 billion activated parameter engine. It lists at $0.3 per million input tokens and $1.2 per million output tokens, with cached input priced at $0.03 per million to lower the cost of repeated context. This makes it a practical default for high-volume, context-reuse pipelines.
MiniMax Music 2.6 API (Text to Music)Describe a style and supply structured lyrics with tags like [Verse] and [Chorus], and MiniMax Music 2.6 returns a complete vocal or instrumental track from a single generation request. Output formats include MP3, WAV, and PCM at sample rates up to 44.1 kHz and bitrates up to 256 kbps, accepting prompts up to 2,000 characters and lyrics up to 3,500. At $0.15 per generation, it suits soundtrack prototyping, jingles, and background music for apps and video.

New features of MiniMax + Showcase

Combining advanced models with Atlas Cloud's GPU-accelerated platform delivers unmatched speed, scalability, and creative control for image and video generation.

Advanced Coding and Agent Planning using MiniMax M2.5

Advanced Coding and Agent Planning using MiniMax M2.5

MiniMax M2.5 supports over 10 programming languages, including Rust, Go, and Python, to facilitate comprehensive full-stack development across Web, mobile, and desktop platforms. By integrating deep industry knowledge for professional document formatting and financial modeling, it enables seamless transitions from system architecture design to final deliverable testing. It is the definitive solution for complex software engineering and high-stakes office productivity workflows.

Rapid Response and Task Decision Efficiency using MiniMax M2.5

Rapid Response and Task Decision Efficiency using MiniMax M2.5

The M2.5 architecture achieves a 37% speed increase in end-to-end execution, significantly reducing complex task durations from 31.3 to 22.8 minutes on the SWE-bench. By optimizing task decomposition logic, the model requires 20% fewer tokens and search rounds to reach objectives in benchmarks like BrowseComp. It offers a streamlined solution for high-velocity decision-making while eliminating redundant computational overhead.

Interleaved Thinking That Holds Its Reasoning

Interleaved Thinking That Holds Its Reasoning

Reasoning tokens are preserved between turns, so the model resumes each step of a multi-turn agent conversation with its earlier thought process intact. This interleaved thinking lets it plan, act, observe, and refine without discarding the logic that led there. The payoff is steadier long-horizon decision-making, a decisive edge when a single task spans many sequential tool calls.

Evolutionary Architecture through Large-Scale RL using MiniMax M2.5

Evolutionary Architecture through Large-Scale RL using MiniMax M2.5

Built on a native Agent RL framework, MiniMax decouples its core engine from agent scaffolding to generalize across hundreds of thousands of diverse real-world environments. It incorporates a sophisticated process reward mechanism that utilizes real-time execution feedback to refine reasoning paths and ensure elite output quality. This creates a highly adaptive system capable of maintaining superior accuracy while maximizing overall operational response speed.

A 196K-Token Context Window in the MiniMax API

A 196K-Token Context Window in the MiniMax API

Every request to MiniMax M2.1 carries a 196,608-token context window, wide enough to hold sprawling monorepos, hundreds of pages of specifications, or long agent transcripts in a single pass. Because the full picture stays resident, the model tracks dependencies across files instead of losing the thread mid-task. Teams working on large codebases and dense research corpora gain coherence that short-context models cannot match.

Orchestrate Shell, Browser, and Code with the MiniMax API

Orchestrate Shell, Browser, and Code with the MiniMax API

Function calling turns MiniMax M2.1 into an active agent that orchestrates shell commands, browser sessions, Python execution, and MCP tools across long-horizon plans. It chains dozens of steps, reads execution feedback, and self-corrects rather than halting at the first result. Scoring 47.9 percent on Terminal-bench 2.0, it makes a dependable foundation for autonomous build, test, and deployment loops.

One MiniMax API Key for Every Model Tier

One MiniMax API Key for Every Model Tier

One OpenAI-compatible key unlocks the entire MiniMax lineup, from M2 and M2.1 through M2.5, M2.7, and M3, alongside Anthropic SDK and raw HTTP access. Switch tiers by editing a single model string, with no rewiring of your stack. Prompt caching bills repeated context at $0.03 per million tokens, so iterative agent loops and long system prompts stay affordable at scale.

One Prompt, Three Models: The MiniMax API Coding Face-Off

Feed the identical build prompt to the MiniMax API and two rival models, then watch how each one turns a single instruction into a working, single-file web page.

Prompt

A pristine, single-file, self-contained HTML page (inline CSS and JavaScript only — absolutely no external libraries, CDNs, fonts, or image URLs; every texture built purely from CSS gradients, box-shadows, and inset shadows) that renders an interactive, skeuomorphic LETTERPRESS TYPESETTING WORKBENCH: fill the viewport with a warm walnut-wood work surface (layered brown gradients + subtle grain via repeating linear-gradients and noise-like shadow stippling), lit by a single warm tungsten lamp from the upper-left so brushed-metal parts catch a soft specular highlight and every raised object casts a long soft shadow to the lower-right. Left side: a compartmentalized type case (California job case) with grid cells holding tiny embossed lead-type sorts (letters, numbers, punctuation), each sort a beveled metal slug with a raised mirror-reversed glyph on its face. Center: an open composing stick / galley tray where set type accumulates. Provide a full working typesetting flow: (1) the user types on the physical keyboard OR clicks glyphs in the type case, and for each character a lead sort animates flying from its case cell and drops into the composing stick with a crisp "clack" micro-bounce settle animation; (2) draggable thin lead spacer blocks between sorts let the user adjust letter-spacing/kerning by dragging, with live layout reflow; (3) a rotatable ink roller (brayer) the user drags across the locked-up forme, leaving a glossy deep-indigo ink sheen (animated moving specular reflection) on the raised type faces; (4) a pull-down impression lever the user presses (click-and-hold or button) that lowers a sheet of cream laid paper over the inked forme and stamps it — the climactic moment is the lever springing back while the freshly printed line appears on the paper with genuinely debossed letterpress typography: each letter pressed into the paper with an inset shadow giving a physical indentation, ink slightly bleeding/feathering into visible paper-fiber grain, a faint plate impression halo around the block. Palette: burnt umber, brushed brass gold, and warm cream/rice-paper off-white dominate, with deep indigo ink as the ONLY cool accent. Include tactile details: brass thumbscrews, a quoin lock, wood-toned buttons with realistic pressed/active states, and satisfying easing on every transition (spring-like cubic-bezier). Make it responsive (scales gracefully down to tablet width), keep all state in vanilla JS, and let the user compose a full short line of poetry, ink it, and pull a print they can visually admire — everything must feel handcrafted, weighty, and ceremonial, with the physics of the drop, the drag, and the impression conveyed through motion and shadow rather than any external asset.

Generated with MiniMax M3 on Atlas Cloud

Generated with Grok 4.5 on Atlas Cloud

Generated with MiniMax M2.7 on Atlas Cloud

Prompt

Create a single, complete, self-contained HTML file that opens directly in any modern browser with zero external dependencies — all CSS and JavaScript must be inlined, and no external libraries, CDNs, fonts, or image URLs are allowed (draw everything with Canvas 2D or SVG generated in code). Build an interactive "Brass Gear Train Workbench": a snap-together mechanical gear simulator rendered on a dusty engineering blueprint tabletop, where the user assembles a working geartrain and watches real physics-accurate rotation ripple through the chain. Core interaction and simulation requirements: - Render the main stage as a blueprint drafting surface: a fine cyan grid on deep engineering-blue paper with a faint film of dust and vignette, plus dimension tick marks and drafting annotations along the edges. - Provide a left-hand parts bin/sidebar holding gears of several distinct sizes (different tooth counts / modules, e.g. 12, 18, 24, 36, 48 teeth). The user drags a gear from the bin onto the stage; it snaps to the grid on release. - Each gear must be drawn as a real involute-style toothed wheel in brass tones — warm gold body, aged copper-patina rim, ivory highlight, visible hub, spokes, and a center bore — with subtle bevel shading so metal reads as metal. - Implement genuine mesh detection: when two placed gears sit at a center distance equal to the sum of their pitch radii (module × teeth / 2) within a tolerance, they automatically engage and lock into a driven chain; draw a meshing indicator where teeth interlock. Chains can branch through multiple gears. - Place a crank handwheel at the lower-left as the visual and mechanical origin. The user turns the crank (drag to rotate it, or hold a button); the entire meshed chain then rotates in sequence. Compute each downstream gear's angular velocity and direction strictly from the gear ratio (ω_out = ω_in × N_in / N_out, direction alternating each mesh), and animate teeth staying perfectly phase-locked at the contact point so no teeth visibly overlap or pass through each other. - The final gear in the chain drives a pointer needle across a numbered dial gauge; the needle must show a real reading derived from the cumulative gear ratio (display the computed output RPM and total ratio as live text). - If the user tries to place gears whose tooth geometry cannot mesh (incompatible spacing/module), that junction must jam: flash it red, halt propagation past the bad joint, and show a small "tooth mismatch" warning. - Add a light beam: a shaft of warm morning sunlight slanting from an upper window across the stage, with drifting dust motes animated in the beam and warm specular glints sliding across the gear teeth as they turn, so the machine looks like it is waking up. Visual and UX polish: - Art direction: retro engineering hand-drawn linework fused with photoreal metal — blueprint fine-grid layered under brass, aged bronze, patina green-blue, and ivory white; warm side-lighting with long soft drop shadows to sculpt volume and depth. - Composition should read the crank as the entry point with the drive chain cutting diagonally across the frame. - Include an uncluttered HUD showing crank speed, total gear ratio, output RPM, and gear count, plus small controls to reverse the crank, clear the stage, and adjust crank speed with a slider. - Everything must be smooth (requestAnimationFrame), responsive to window size, and keep running interactively — dragging, snapping, meshing, and reading updates should all respond in real time. Make it feel tactile, precise, and alive.

Generated with MiniMax M3 on Atlas Cloud

Generated with Grok 4.5 on Atlas Cloud

Generated with MiniMax M2.7 on Atlas Cloud

Build Production Software with the MiniMax API

From multi-language full-stack builds to autonomous office agents, the MiniMax API turns M2.1's coding depth and long-horizon tool use into workflows your team ships every day.

Full-Stack Development Across Languages

MiniMax M2.1 writes and reviews production code in Rust, Go, TypeScript, C++, and more, spanning low-level systems through application layers. Teams building polyglot services rely on it to keep every tier consistent.

Native Mobile Apps via the MiniMax API

When Android and iOS features need to ship together, the MiniMax API generates Kotlin, Swift, and Objective-C aligned to platform conventions. This lets small teams maintain both stores without splitting engineering effort.

Autonomous Coding Agents on the MiniMax API

Wire M2.1 into your agent loop through function calling to plan multi-step tasks, run tools, and self-correct across long horizons. Developers use this to build assistants that finish real tickets, not just snippets.

Digital Employees for Office Workflows

Acting as a digital employee, M2.1 executes end-to-end tasks across administration, finance, data science, and HR from plain text commands. Operations teams deploy it to clear repetitive back-office work at scale.

Whole-Codebase Reasoning at 196K Context

Need to refactor across an unfamiliar repository? With a 196.61K token context window, M2.1 reads hundreds of files at once, tracing dependencies and surfacing the changes a migration actually requires.

Interactive Web Interfaces with the MiniMax API

Strong design comprehension lets the MiniMax API build polished web frontends, complex interactions, and even 3D scientific scene simulations. Product teams ship data-rich dashboards and visualizations straight from a prompt.

How the MiniMax API Compares to Other Coding and Agent LLMs

Weigh the MiniMax API against rival agentic and coding models on Atlas Cloud, side by side on context window, output ceiling, and pay-as-you-go token pricing.

ModelContext WindowMax OutputInput Price ($/1M)Output Price ($/1M)
MiniMax M3524K tokens512K tokens$0.60 / $1.20 (>512K)$2.40 / $4.80 (>512K)
MiniMax M2.7196K tokens196K tokens$0.30$1.20
MiniMax M2.5196K tokens196K tokens$0.30$1.20
Kimi K2.7 Code262K tokens262K tokens$0.95$4.00
GLM 4.7203K tokens203K tokens$0.60$2.20
DeepSeek V4 Flash1M tokens393K tokens$0.14$0.28

How to Use MiniMax on Atlas Cloud

Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.

Create an Atlas Cloud Account

Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.

Why Use MiniMax on Atlas Cloud

Combining the advanced MiniMax models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.

Performance & flexibility

Low Latency:
GPU-optimized inference for real-time reasoning.

Unified API:
Run MiniMax, GPT, Gemini, and DeepSeek with one integration.

Transparent Pricing:
Predictable per-token billing with serverless options.

Enterprise & Scale

Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.

Reliability:
99.99% uptime, RBAC, and compliance-ready logging.

Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.

MiniMax API Questions, Answered

The MiniMax API gives developers programmatic access to the MiniMax family of large language models built by MiniMax AI, a series engineered around coding, agentic tool use, and long-horizon reasoning. On Atlas Cloud you reach these models through one OpenAI-compatible endpoint with transparent pay-as-you-go pricing and Day-0 access to new releases.

Atlas Cloud hosts the M-series text models, including MiniMax M2.5, MiniMax M2.7, and the flagship MiniMax M3, alongside MiniMax Music 2.6 for text-to-music generation. Every model shares the same authentication and endpoint, so switching between them means changing a single model identifier in your request.

These models are optimized for end-to-end software engineering, handling multi-file editing, compile-run-fix loops, and test-validated repair across many programming languages. They also hold up well in agentic settings that need function calling, multi-step planning, and recovery from execution errors, which makes them a natural fit for autonomous coding agents and tool-driven workflows.

Because the MiniMax API is OpenAI-compatible, you can point the standard OpenAI SDK at the Atlas Cloud base URL, pass your key, and set the model field to the version you want. Streaming, function calling, and JSON-structured output all run through the same interface, so most existing integrations need only a base URL and key change to get started.

Context windows vary by version. MiniMax M2 handles roughly 200K tokens, enough to load a large codebase or lengthy documentation in a single request, while the newer M3 extends this to about 1 million tokens. Reach for the higher-context model when your task spans many files or long reference material.

Billing is pay-as-you-go and metered per token, with no subscription required. The M2-series text models start at $0.30 per million input tokens and $1.20 per million output tokens, while the higher-capacity M3 is priced at $0.60 per million input and $2.40 per million output, and cached input is billed at a lower rate.

Yes. MiniMax has open-sourced the M-series weights under a modified MIT license that adds an attribution requirement for very large commercial deployments. You can self-host the weights for evaluation, then move to the Atlas Cloud API when you want managed infrastructure, consistent uptime, and no GPU maintenance.

MiniMax aims for frontier-level coding and agentic quality at a fraction of the cost of larger proprietary models, using a mixture-of-experts design that activates only about 10 billion of its 230 billion parameters per token. If your workload is code generation, refactoring, or tool-using agents on a tight budget, it delivers a strong balance of capability, speed, and price. Benchmark standing shifts with each release, so validate the current version against your own tasks before committing.

Explore More Families

Seedance 2.5

The Seedance 2.5 API delivers ByteDance's newest video generation model, the successor to Seedance 2.0 built on a unified multimodal architecture. It renders up to 30 seconds of footage in one pass, keeps subjects consistent under believable physics, draws text and multilingual subtitles directly in frame. Atlas Cloud brings Day-0 access on the same unified endpoints that already serve Seedance 2.0 and 1.5. Start building today.

View Family

MiniMax H3

The MiniMax H3 API opens MiniMax's general purpose multimodal video model, which reads text, images, video and audio as one context instead of one task at a time. Clips run 5 to 15 seconds at 24 FPS across aspect ratios from 21:9 to 9:16, and one prompt can swap characters, replace backgrounds, rewrite dialogue or clone a voice from a reference clip. Atlas Cloud serves it all through one OpenAI-compatible endpoint. Start building today.

View Family

Seedream 5.0 Pro

Seedream 5.0 Pro API gives developers ByteDance's controllable image editing model on Atlas Cloud. It places edits precisely with anchors and coordinates, separates images into editable layers, fuses multiple references, and matches exact colors and materials, with multilingual text at 2K and 3K. On Atlas Cloud you reach it through one key!

View Family

Seedance 2.0

The Seedance 2.0 API gives you production access to ByteDance's multimodal video model — quad-modal inputs (text, image, video, audio) and an industry-leading "Universal Reference" system that locks composition, camera movement, and character actions across shots. Integrate director-level control with one API call, a flat $0.09/s, instant key, and no waitlist — backed by enterprise-grade uptime and compliance. Seedance 2.0 Native 4K is now live!

View Family

GPT Image 2

The GPT Image 2 API gives developers access to OpenAI's latest image model, the successor to GPT Image 1.5. It generates and edits images with accurate text rendering across Latin and CJK scripts, plus strong composition for posters, mockups, and infographics. On Atlas Cloud you reach it through one unified API alongside 300+ models, with free credits, 99.99% uptime, and no OpenAI organization verification required.

View Family

Gemini Omni Flash

The Gemini Omni API brings Google DeepMind's multimodal video generation and editing model, introduced at Google I/O 2026, to your stack. Gemini Omni fuses Gemini's reasoning engine with generative media, accepting any mix of text, images, video, and audio to produce consistent, knowledge-grounded output. Refine results through natural conversation, swapping objects, rewriting scenes, and shifting styles while physics, characters, and continuity stay intact. Atlas Cloud serves the full Gemini Omni Flash lineup, text-to-video, image-to-video with up to 7 reference images, and reference-to-video, through one unified API with transparent per-second pricing from $0.112 and no subscription. Start building today.

View Family

Grok Imagine

The Grok Imagine API gives developers xAI's image, video, and audio generation in one suite. It produces up to 2K images with multilingual text rendering, plus video up to 15 seconds with native, synchronized audio and reference-based editing. On Atlas Cloud one key runs every Grok Imagine mode, so you move between image, video, and audio without separate setups, from $0.02 per image and $0.05 per second.

View Family

Google

Google's most powerful creative models are all available on Atlas Cloud. Veo 3.1 delivers cinematic video generation, Nano Banana 2 powers high-fidelity image creation, and Gemini brings multimodal intelligence to every workflow. Access the full Google model suite through one API key with Day-0 availability and pay-as-you-go pricing.

View Family

Seedance 2.0 Mini

The Seedance 2.0 Mini API is the lightest, lowest-cost tier of ByteDance's Seedance video line, built for teams where throughput and unit cost matter more than maximum polish. Use it for batch generation, rapid prototyping, and draft passes, all through one OpenAI-compatible key on Atlas Cloud.

View Family

ByteDance

From cinematic video generation to high-fidelity image creation, ByteDance's most powerful models are live on Atlas Cloud. Run Seedance and Seedream at scale with the lowest inference pricing and zero infrastructure overhead.

View Family

Alibaba

Atlas Cloud brings together Alibaba's full model lineup under one API: Qwen for language and image tasks, Wan for video generation up to 1080p. Access every model pay-as-you-go with no subscriptions. The Alibaba API is available via a single base URL using your existing OpenAI-compatible client.

View Family

OpenAI

Atlas Cloud gives you access to the full OpenAI API lineup, from GPT Image 2 for image generation to Sora 2 for video. Every model is available pay-as-you-go with no monthly commitment. Plug in with a single base URL swap using the OpenAI-compatible API.

View Family

One API for All Media AI.

Explore all models