Inception Labs' fastest reasoning model with tool calling and structured outputs support.
Context Window
128.0K
Max Output
50.0K
Input Price (Auto)
$0.25/1M
Output Price (Auto)
$0.75/1M
Cache Read (Auto)
$0.025/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
21.9
Coding Index
31.1
Agentic Index
9.5
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
77.0%
Better than 64% of models compared
HLE
Humanity's Last Exam
17.1%
Better than 72% of models compared
IFBench
Instruction-following benchmark
69.8%
Better than 84% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
70.8%
Better than 63% of models compared
AA-LCR
Long context reasoning evaluation
40.7%
Better than 46% of models compared
Coding
SciCode
Python programming for scientific computing
38.7%
Better than 65% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
26.5%
Better than 68% of models compared
Last updated Aug 16, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Mercury 2 with similar models from the same provider or model family.
Mercury Coder Small
mercury-coder-smallModel by Inception AI. A diffusion large language model that runs incredibly quickly (500+ tokens/second) while matching Claude 3.5 Haiku and GPT-4o-mini. 1st in speed on Copilot arena, and matching 2nd in quality.
Gemma 4 31B MeroMero v2
Gemma-4-31B-MeroMero-v2Gemma 4 31B MeroMero v2 is a LoRA finetune for emotive dialogue, relationship scenes, creative writing, and multimodal roleplay.
Ornith 1.5 9B
ornith-ai/ornith-1.5-9bOrnith 1.5 9B is an FP8 dense open-weight reasoning model built for agentic coding, tool use, visual understanding, and efficient long-context work.
Gemma 4 26B A4B Uncensored
google/gemma-4-26b-a4b-uncensoredGemma 4 26B A4B Uncensored is an FP8 open-weight multimodal mixture-of-experts model LoRA-tuned for fewer refusals across chat, coding, tool use, and long-context work.
DeepSeek V4 Flash Vision Exp
deepseek/deepseek-v4-flash-vision-expAn experimental vision-enabled DeepSeek V4 Flash model that adds image understanding while retaining the text, reasoning, coding, tool-calling, and agent capabilities of the base model. This route is served directly by DeepSeek, so privacy and logging guarantees are limited.
Qwen 3.6 35B A3B Uncensored
qwen/qwen3.6-35b-a3b-uncensoredQwen 3.6 35B A3B Uncensored is an NVFP4 open-weight mixture-of-experts model LoRA-tuned for fewer refusals across chat, coding, tool use, and multimodal tasks.