Gemini 2.0 Pro Reasoner
Note: This model is now being routed to Gemini 2.5 Pro because Google no longer has the Gemini 2.0 Pro model available and Gemini 2.5 Pro is an across the board improvement. 'DeepGemini', fusion of Gemini 2.5 Pro and Deepseek R1.
- Audio Input
Pricing
Auto routing · per 1M tokens- Input
- $1.29
- Output
- $5.00
- Cache read
- $0.32
Specifications
- Context window
- 128K
- Max output
- 65.5K
Benchmarks
Benchmarks
Sourced from Artificial Analysis.
Intelligence Index
16.1
Coding Index
33.3
Agentic Index
1.6
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
2.2%
Better than 21% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
298 Elo
Better than 16% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
459 Elo
Better than 24% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
10.2%
Better than 41% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
69.0%
Better than 60% of models compared
Reasoning
HLE
Humanity's Last Exam
22.5%
Better than 71% of models compared
IFBench
Instruction-following benchmark
48.7%
Better than 58% of models compared
CritPt
Research-level physics reasoning
2.6%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
0.0%
Better than 16% of models compared
SciCode
Python programming for scientific computing
46.3%
Better than 43% of models compared
LiveCodeBench
Contamination-free coding benchmark
80.1%
Better than 92% of models compared
Math
AIME 2025
American Invitational Mathematics Examination 2025
87.7%
Better than 85% of models compared
AIME
American Invitational Mathematics Examination
88.7%
Better than 95% of models compared
Math-500
Diverse mathematical problem solving benchmark
96.7%
Better than 86% of models compared
Knowledge
MMLU-Pro
Professional and academic subject knowledge
86.2%
Better than 95% of models compared
AA-Omniscience Accuracy
Proportion of correctly answered questions
39.1%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
90.9%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
84.4%
Better than 76% of models compared
Terminal-Bench Hard (legacy)
Agentic coding and terminal use
26.5%
Better than 68% of models compared
T²-Bench Telecom (legacy)
Conversational AI agents in dual-control scenarios
54.1%
Better than 55% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
69.0%
Better than 60% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
0.0%
Last updated Oct 3, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…