Model by Inception AI. A diffusion large language model that runs incredibly quickly (500+ tokens/second) while matching Claude 3.5 Haiku and GPT-4o-mini. 1st in speed on Copilot arena, and matching 2nd in quality.
Context Window
32.8K
Max Output
16.4K
Input Price (Auto)
$0.25/1M
Output Price (Auto)
$1.00/1M
Cache Read (Auto)
$0.13/1M
Benchmarks
Benchmarks
Performance metrics and benchmarks
No benchmark data is available yet for this model.
Providers
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Mercury Coder Small with similar models from the same provider or model family.
Mercury 2
mercury-2Inception Labs' fastest reasoning model with tool calling and structured outputs support.
Inkling Small
thinkingmachines/Inkling-SmallThe direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.
Inkling Small Thinking
thinkingmachines/Inkling-Small:thinkingThe reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.
KAT Coder Air V2.5
kwaipilot/kat-coder-air-v2.5Fast, cost-efficient KAT Coder model for code generation, editing, debugging, and agentic software-development workflows.
KAT Coder Pro V2.5
kwaipilot/kat-coder-pro-v2.5Higher-capability KAT Coder model for complex code generation, repository-scale editing, debugging, and agentic software-development workflows.
KAT Coder Pro V2
kwaipilot/kat-coder-pro-v2Latest high-performance coding model in Kwaipilot's KAT Coder series, built for complex software engineering, SaaS integration, and large-scale production workflows.