Provider logo

Mercury 2

mercury-2
Provider logo

Mercury 2

mercury-2

Inception Labs' fastest reasoning model with tool calling and structured outputs support.

Context Window

128.0K

Max Output

50.0K

Avg output tokens (7d)

441 tokens

37%

Input Price (Auto)

$0.25/1M

Output Price (Auto)

$0.75/1M

Cache Read (Auto)

$0.025/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

Sourced from Artificial Analysis.

Intelligence Index

11.5

Better than 48% of models compared

Coding Index

31.1

Better than 38% of models compared

Agentic Index

4.0

Better than 30% of models compared

Agentic work

AutomationBench-AA

Workflow automation with guardrail penalties

1.7%

Better than 22% of models compared

AA-Briefcase

Agentic knowledge work (Elo)

280 Elo

Better than 18% of models compared

GDPval-AA v2

Economically valuable tasks (Elo)

646 Elo

Better than 29% of models compared

Document reasoning

GDP.pdf

Professional PDF reasoning: all-pass rate

3.8%

Better than 26% of models compared

AA-LCR v1.1

Long context reasoning with updated grading

43.7%

Better than 42% of models compared

Reasoning

HLE

Humanity's Last Exam

17.1%

Better than 68% of models compared

IFBench

Instruction-following benchmark

69.8%

Better than 84% of models compared

Coding

Terminal-Bench v4.0

Practical coding and terminal tasks

0.0%

Better than 17% of models compared

SciCode

Python programming for scientific computing

37.7%

Better than 20% of models compared

Legacy benchmarks

GPQA Diamond (legacy)

Graduate-level scientific reasoning

77.0%

Better than 61% of models compared

Terminal-Bench Hard (legacy)

Agentic coding and terminal use

26.5%

Better than 68% of models compared

T²-Bench Telecom (legacy)

Conversational AI agents in dual-control scenarios

70.8%

Better than 63% of models compared

AA-LCR (unversioned / legacy)

Long context reasoning evaluation

43.7%

Better than 42% of models compared

Last updated Sep 13, 2026

Artificial Analysis

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare Mercury 2 with similar models from the same provider or model family.

Mercury 2.5 Preview

inception/mercury-2.5-preview

Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.

Mercury Coder Small

mercury-coder-small

Model by Inception AI. A diffusion large language model that runs incredibly quickly (500+ tokens/second) while matching Claude 3.5 Haiku and GPT-4o-mini. 1st in speed on Copilot arena, and matching 2nd in quality.

Schematron V2 Small

inference-net/schematron-v2-small

Inference.net's 3B-parameter HTML-to-JSON extraction model, focused on accuracy for complex schemas and long web pages. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.

Schematron V2 Turbo

inference-net/schematron-v2-turbo

Inference.net's 3B-parameter HTML-to-JSON extraction model, optimized for throughput and low cost on high-volume workloads. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.

DeepSeek V4.1 Flash TEE

TEE/deepseek-v4.1-flash

DeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.

GPT Astra Latest

openai/gpt-astra-latest

Compatibility alias that routes to GPT 6 Astra, the latest supported GPT Astra model.