TWO WEEKS ONLY | 20% OFF Seedream 5.0 Pro!
DeepSeek V4 Pro
LLM
PRO

DeepSeek V4 Pro

DeepSeek V4 Pro is a state-of-the-art large language model combining efficient sparse attention, strong reasoning, and integrated agent capabilities for robust long-context understanding and versatile AI applications.

DeepSeek V4 Pro
MiniMax M2.7
Kimi K2.5
Qwen3.5 122B A10B
CATEGORY
Series
48 of 61 models
New
NEW
HOT
Kimi K3
LLM

Kimi K3

Top-performing open-weight model optimized for frontier reasoning, coding, and enterprise AI applications.

1048.6K CONTEXT:
Input type:
Output type:
Context:1048.58K
Input:$3/M tokens
Output:$15/M tokens
Max Output:1048.58K
$3/15M in/out
Cache-Based
NEW
HOT
Grok 4.5
LLM

Grok 4.5

Flagship conversational model built for real-time knowledge exploration, sharp reasoning, and highly engaging AI interactions.

500.0K CONTEXT:
Input type:
Output type:
Context:500.00K
Input:$2/M tokens
Output:$6/M tokens
Max Output:500.00K
$2/6M in/out
Cache-Based
NEW
HOT
KwaiKAT
LLM
PRO

KAT Coder Pro V2.5

High-end coding agent built for complex software engineering, repository-scale tasks, and autonomous development workflows.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.74/M tokens
Output:$2.96/M tokens
Max Output:262.14K
$0.74/2.96M in/out
Cache-Based
NEW
HOT
KwaiKAT
LLM

KAT Coder Air V2.5

Fast, lightweight coding model designed for interactive development, rapid code iteration, and efficient coding agents.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.15/M tokens
Output:$0.6/M tokens
Max Output:262.14K
$0.15/0.6M in/out
Cache-Based
NEW
HOT
Doubao Seed Character
LLM

Doubao Seed Character

Flagship model delivering premium reasoning, coding, multimodal understanding, and enterprise-grade performance.

131.1K CONTEXT:
Input type:
Output type:
Context:131.07K
Input:$0.2/M tokens
Output:$0.8/M tokens
Max Output:32.77K
$0.2/0.8M in/out
Cache-Based
Gradient-Based
NEW
HOT
Doubao Seed 2.1 Turbo
LLM
TURBO

Doubao Seed 2.1 Turbo

Turbo model optimized for ultra-low latency, high throughput, and responsive AI experiences.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.45/M tokens
Output:$2.25/M tokens
Max Output:262.14K
$0.45/2.25M in/out
Cache-Based
NEW
HOT
Doubao Seed 2.1 Pro
LLM
PRO

Doubao Seed 2.1 Pro

Flagship model delivering premium reasoning, coding, multimodal understanding, and enterprise-grade performance.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.9/M tokens
Output:$4.5/M tokens
Max Output:262.14K
$0.9/4.5M in/out
Cache-Based
NEW
HOT
GLM 5.2
LLM

GLM 5.2

Agent-oriented model built for complex reasoning, tool use, and autonomous task execution.

1048.6K CONTEXT:
Input type:
Output type:
Context:1048.58K
Input:$1.26/M tokens
Output:$3.96/M tokens
Max Output:131.07K
$1.26/3.96M in/out
Cache-Based
NEW
HOT
Kimi K2.7 Code
LLM

Kimi K2.7 Code

Powerful coding model for programming, debugging, and AI developer workflows.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.95/M tokens
Output:$4/M tokens
Max Output:262.14K
$0.95/4M in/out
Cache-Based
NEW
HOT
MiniMax M3
LLM

MiniMax M3

MiniMax M3 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.

524.3K CONTEXT:
Input type:
Output type:
Context:524.30K
Input:$0.3/M tokens
Output:$1.2/M tokens
Max Output:524.29K
$0.3/1.2M in/out
Cache-Based
Gradient-Based
NEW
HOT
Grok Build 0.1
LLM

Grok Build 0.1

Specialized coding model optimized for software development, code generation, debugging, refactoring, and developer workflows.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$1/M tokens
Output:$2/M tokens
Max Output:262.14K
$1/2M in/out
Cache-Based
Gradient-Based
NEW
HOT
Grok 4.3
LLM

Grok 4.3

Advanced conversational AI model optimized for natural dialogue, knowledge exploration, reasoning, and interactive chat experiences.

1000.0K CONTEXT:
Input type:
Output type:
Context:1000.00K
Input:$1.25/M tokens
Output:$2.5/M tokens
Max Output:1000.00K
$1.25/2.5M in/out
Cache-Based
Gradient-Based
NEW
HOT
Gemini 3.5 Flash
LLM

Gemini 3.5 Flash

Fast and cost-efficient multimodal model designed for high-throughput applications, real-time interactions, and everyday AI tasks.

1048.6K CONTEXT:
Input type:
Output type:
Context:1048.58K
Input:$1.5/M tokens
Output:$9/M tokens
Max Output:65.54K
$1.5/9M in/out
Cache-Based
NEW
HOT
DeepSeek V4 Pro
LLM
PRO

DeepSeek V4 Pro

DeepSeek V4 Pro is a state-of-the-art large language model combining efficient sparse attention, strong reasoning, and integrated agent capabilities for robust long-context understanding and versatile AI applications.

1048.6K CONTEXT:
Input type:
Output type:
Context:1048.58K
Input:$1.68/M tokens
Output:$3.38/M tokens
Max Output:393.22K
$1.68/3.38M in/out
Cache-Based
NEW
HOT
DeepSeek V4 Flash 0731
LLM

DeepSeek V4 Flash 0731

DeepSeek V4 Flash is a state-of-the-art large language model combining efficient sparse attention, strong reasoning, and integrated agent capabilities for robust long-context understanding and versatile AI applications.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.14/M tokens
Output:$0.28/M tokens
Max Output:131.07K
$0.14/0.28M in/out
Cache-Based
NEW
HOT
DeepSeek V4 Flash
LLM

DeepSeek V4 Flash

DeepSeek V4 Flash is a state-of-the-art large language model combining efficient sparse attention, strong reasoning, and integrated agent capabilities for robust long-context understanding and versatile AI applications.

1048.6K CONTEXT:
Input type:
Output type:
Context:1048.58K
Input:$0.14/M tokens
Output:$0.28/M tokens
Max Output:393.22K
$0.14/0.28M in/out
Cache-Based
NEW
HOT
MiMo V2.5
LLM

MiMo V2.5

Processes text, image, audio, and video to text. Delivers strong agentic performance with high token efficiency and low inference cost.

1024.0K CONTEXT:
Input type:
Output type:
Context:1024.00K
Input:$0.14/M tokens
Output:$0.28/M tokens
Max Output:131.07K
$0.14/0.28M in/out
Cache-Based
NEW
HOT
MiMo V2.5 Pro
LLM
PRO

MiMo V2.5 Pro

Optimized for complex reasoning, software engineering, and long-horizon agentic tasks. Supports autonomous execution of extensive tool-calling workflows.

1024.0K CONTEXT:
Input type:
Output type:
Context:1024.00K
Input:$0.435/M tokens
Output:$0.87/M tokens
Max Output:131.07K
$0.435/0.87M in/out
Cache-Based
NEW
HOT
Hunyuan 3
LLM

Hunyuan 3

Flagship foundation model built for deep reasoning, complex problem-solving, and sophisticated agentic workflows.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.2/M tokens
Output:$0.8/M tokens
Max Output:131.07K
$0.2/0.8M in/out
Cache-Based
HOT
Kimi K2.6
LLM

Kimi K2.6

Enhanced model for reasoning, coding, and productivity.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.95/M tokens
Output:$4/M tokens
Max Output:262.14K
$0.95/4M in/out
Cache-Based
NEW
Qwen3.6 35B A3B
LLM

Qwen3.6 35B A3B

The latest Qwen reasoning model.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.186/M tokens
Output:$1.114/M tokens
Max Output:65.54K
$0.186/1.114M in/out
NEW
Qwen3.6 Plus
LLM

Qwen3.6 Plus

Versatile model for chat, and productivity workflows.

1000.0K CONTEXT:
Input type:
Output type:
Context:1000.00K
Input:$0.325/M tokens
Output:$1.95/M tokens
Max Output:65.54K
$0.325/1.95M in/out
Cache-Based
Gradient-Based
NEW
HOT
Doubao Seed Evolving
LLM

Doubao Seed Evolving

Self-improving research model designed for adaptive reasoning, exploration, and continuous learning workflows.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.9/M tokens
Output:$4.5/M tokens
Max Output:262.14K
$0.9/4.5M in/out
Cache-Based
NEW
HOT
Doubao Seed 2.0 Pro
LLM
PRO

Doubao Seed 2.0 Pro

Professional-grade model built for advanced workloads, complex analysis, and enterprise AI applications.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.5/M tokens
Output:$3/M tokens
Max Output:131.07K
$0.5/3M in/out
Cache-Based
Gradient-Based
NEW
HOT
Doubao Seed 2.0 Code Preview
LLM
PREVIEW

Doubao Seed 2.0 Code Preview

Developer-focused model specialized in coding agents, repository understanding, and software engineering.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.5/M tokens
Output:$3/M tokens
Max Output:131.07K
$0.5/3M in/out
Cache-Based
Gradient-Based
NEW
HOT
Doubao Seed 2.0 Lite
LLM

Doubao Seed 2.0 Lite

Ultra-efficient model focused on lightweight AI tasks, rapid inference, and large-scale deployment.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.25/M tokens
Output:$2/M tokens
Max Output:131.07K
$0.25/2M in/out
Cache-Based
Gradient-Based
NEW
HOT
Doubao Seed 2.0 Mini
LLM

Doubao Seed 2.0 Mini

Small yet capable model designed for edge scenarios, automation, and cost-sensitive services.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.1/M tokens
Output:$0.4/M tokens
Max Output:131.07K
$0.1/0.4M in/out
Cache-Based
Gradient-Based
NEW
HOT
Doubao Seed 1.8
LLM

Doubao Seed 1.8

Next-generation assistant model with improved instruction following and deeper contextual understanding.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.25/M tokens
Output:$2/M tokens
Max Output:65.54K
$0.25/2M in/out
Cache-Based
Gradient-Based
NEW
HOT
Doubao Seed 1.6 Flash
LLM

Doubao Seed 1.6 Flash

High-speed model engineered for instant responses, real-time interaction, and massive request workloads.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.075/M tokens
Output:$0.3/M tokens
Max Output:32.77K
$0.075/0.3M in/out
Cache-Based
Gradient-Based
NEW
HOT
Doubao Seed 1.6
LLM

Doubao Seed 1.6

Versatile foundation model providing reliable conversation, knowledge understanding, and content creation.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.25/M tokens
Output:$2/M tokens
Max Output:65.54K
$0.25/2M in/out
Cache-Based
Gradient-Based
NEW
HOT
GLM 5.1
LLM

GLM 5.1

GLM-5.1 is Z.AI’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while delivering more natural conversational experiences and superior front-end aesthetics.

202.8K CONTEXT:
Input type:
Output type:
Context:202.75K
Input:$1.26/M tokens
Output:$3.96/M tokens
Max Output:202.75K
$1.26/3.96M in/out
Cache-Based
NEW
HOT
MiniMax M2.7
LLM

MiniMax M2.7

MiniMax-M2.7 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.

196.6K CONTEXT:
Input type:
Output type:
Context:196.61K
Input:$0.3/M tokens
Output:$1.2/M tokens
Max Output:196.61K
$0.3/1.2M in/out
Cache-Based
NEW
Qwen3.5 122B A10B
LLM

Qwen3.5 122B A10B

Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.3/M tokens
Output:$2.4/M tokens
Max Output:65.54K
$0.3/2.4M in/out
NEW
Qwen3.5 35B A3B
LLM

Qwen3.5 35B A3B

Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.225/M tokens
Output:$1.8/M tokens
Max Output:65.54K
$0.225/1.8M in/out
NEW
Qwen3.5 27B
LLM

Qwen3.5 27B

Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.27/M tokens
Output:$2.16/M tokens
Max Output:65.54K
$0.27/2.16M in/out
NEW
Qwen3.5 397BA17B
LLM

Qwen3.5 397BA17B

Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.55/M tokens
Output:$3.5/M tokens
Max Output:65.54K
$0.55/3.5M in/out
Cache-Based
HOT
MiniMax M2.5
LLM

MiniMax M2.5

MiniMax-M2.5 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.

196.6K CONTEXT:
Input type:
Output type:
Context:196.61K
Input:$0.295/M tokens
Output:$1.2/M tokens
Max Output:196.61K
$0.295/1.2M in/out
Cache-Based
NEW
HOT
GLM 5v Turbo
LLM
TURBO

GLM 5v Turbo

GLM-5v Turbo is Z.AI’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while delivering more natural conversational experiences and superior front-end aesthetics.

202.8K CONTEXT:
Input type:
Output type:
Context:202.75K
Input:$1.2/M tokens
Output:$4/M tokens
Max Output:131.07K
$1.2/4M in/out
Cache-Based
NEW
HOT
GLM 5
LLM

GLM 5

GLM-5 is Z.AI’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while delivering more natural conversational experiences and superior front-end aesthetics.

202.8K CONTEXT:
Input type:
Output type:
Context:202.75K
Input:$0.95/M tokens
Output:$3.15/M tokens
Max Output:202.75K
$0.95/3.15M in/out
Cache-Based
HOT
Kimi K2.5
LLM

Kimi K2.5

Powerful model for long-context and intelligent workflows.

262.1K CONTEXT:
Input type:
Output type:
Context:262.14K
Input:$0.49/M tokens
Output:$2.5/M tokens
Max Output:262.14K
$0.49/2.5M in/out
Cache-Based
NEW
Qwen3.7 Max
LLM

Qwen3.7 Max

Flagship model for advanced reasoning, coding, and complex tasks.

1000.0K CONTEXT:
Input type:
Output type:
Context:1000.00K
Input:$2.5/M tokens
Output:$7.5/M tokens
Max Output:67.07K
$2.5/7.5M in/out
Cache-Based
NEW
Qwen3.7 Plus
LLM

Qwen3.7 Plus

Balanced model combining strong capability, speed, and efficiency.

1000.0K CONTEXT:
Input type:
Output type:
Context:1000.00K
Input:$0.4/M tokens
Output:$1.6/M tokens
Max Output:67.07K
$0.4/1.6M in/out
Cache-Based
Gradient-Based
NEW
Qwen3.5 Plus
LLM

Qwen3.5 Plus

Efficient model for everyday tasks and AI assistants.

1000.0K CONTEXT:
Input type:
Output type:
Context:1000.00K
Input:$0.4/M tokens
Output:$2.4/M tokens
Max Output:67.07K
$0.4/2.4M in/out
Cache-Based
Gradient-Based
NEW
Qwen3.5 Flash
LLM

Qwen3.5 Flash

Fast model optimized for instant responses and large-scale usage.

1000.0K CONTEXT:
Input type:
Output type:
Context:1000.00K
Input:$0.1/M tokens
Output:$0.4/M tokens
Max Output:67.07K
$0.1/0.4M in/out
NEW
HOT
GLM 4.7
LLM

GLM 4.7

GLM-4.7 is Z.AI’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while delivering more natural conversational experiences and superior front-end aesthetics.

202.8K CONTEXT:
Input type:
Output type:
Context:202.75K
Input:$0.52/M tokens
Output:$1.85/M tokens
Max Output:202.75K
$0.52/1.85M in/out
Cache-Based
NEW
HOT
DeepSeek V3.2
LLM

DeepSeek V3.2

DeepSeek V3.2 is a state-of-the-art large language model combining efficient sparse attention, strong reasoning, and integrated agent capabilities for robust long-context understanding and versatile AI applications.

163.8K CONTEXT:
Input type:
Output type:
Context:163.84K
Input:$0.26/M tokens
Output:$0.38/M tokens
Max Output:163.84K
$0.26/0.38M in/out
Cache-Based
NEW
HOT
GPT 5.6 Sol
LLM

GPT 5.6 Sol

High-intelligence model built for deep reasoning, ambitious problem-solving, and frontier AI workloads.

1050.0K CONTEXT:
Input type:
Output type:
Context:1050.00K
Input:$5/M tokens
Output:$30/M tokens
Max Output:131.07K
$5/30M in/out
Cache-Based
Gradient-Based
NEW
HOT
GPT 5.6 Terra
LLM

GPT 5.6 Terra

Dependable general-purpose model designed for practical workflows, grounded analysis, and production applications.

1050.0K CONTEXT:
Input type:
Output type:
Context:1050.00K
Input:$2.5/M tokens
Output:$15/M tokens
Max Output:131.07K
$2.5/15M in/out
Cache-Based
Gradient-Based