Inception Labs' fastest reasoning model with tool calling and structured outputs support.
Context Window
128.0K
Max Output
50.0K
Avg output tokens (7d)
441 tokens
Input Price (Auto)
$0.25/1M
Output Price (Auto)
$0.75/1M
Cache Read (Auto)
$0.025/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
11.5
Coding Index
31.1
Agentic Index
4.0
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
1.7%
Better than 22% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
280 Elo
Better than 18% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
646 Elo
Better than 29% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
3.8%
Better than 26% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
43.7%
Better than 42% of models compared
Reasoning
HLE
Humanity's Last Exam
17.1%
Better than 68% of models compared
IFBench
Instruction-following benchmark
69.8%
Better than 84% of models compared
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
0.0%
Better than 17% of models compared
SciCode
Python programming for scientific computing
37.7%
Better than 20% of models compared
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
77.0%
Better than 61% of models compared
Terminal-Bench Hard (legacy)
Agentic coding and terminal use
26.5%
Better than 68% of models compared
T²-Bench Telecom (legacy)
Conversational AI agents in dual-control scenarios
70.8%
Better than 63% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
43.7%
Better than 42% of models compared
Last updated Sep 13, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Mercury 2 with similar models from the same provider or model family.
Mercury 2.5 Preview
inception/mercury-2.5-previewMercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.
Mercury Coder Small
mercury-coder-smallModel by Inception AI. A diffusion large language model that runs incredibly quickly (500+ tokens/second) while matching Claude 3.5 Haiku and GPT-4o-mini. 1st in speed on Copilot arena, and matching 2nd in quality.
Schematron V2 Small
inference-net/schematron-v2-smallInference.net's 3B-parameter HTML-to-JSON extraction model, focused on accuracy for complex schemas and long web pages. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.
Schematron V2 Turbo
inference-net/schematron-v2-turboInference.net's 3B-parameter HTML-to-JSON extraction model, optimized for throughput and low cost on high-volume workloads. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.
DeepSeek V4.1 Flash TEE
TEE/deepseek-v4.1-flashDeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.
GPT Astra Latest
openai/gpt-astra-latestCompatibility alias that routes to GPT 6 Astra, the latest supported GPT Astra model.