Google's Gemma 4 31B instruction-tuned model with thinking explicitly enabled, exposing reasoning traces for complex multimodal and coding workflows.
Added Apr 2, 2026
Model weightsContext Window
262.1K
Max Output
131.1K
Avg output tokens (7d)
1.6K tokens
Input Price (Auto)
$0.100/1M
Output Price (Auto)
$0.34/1M
Cache Read (Auto)
$0.100/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
29.7
Coding Index
43.4
Agentic Index
14.4
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
85.7%
Better than 83% of models compared
HLE
Humanity's Last Exam
23.6%
Better than 79% of models compared
IFBench
Instruction-following benchmark
75.6%
Better than 93% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
59.9%
Better than 57% of models compared
AA-LCR
Long context reasoning evaluation
68.3%
Better than 76% of models compared
GDPval-AA
Economically valuable tasks
15.5%
CritPt
Research-level physics reasoning
1.4%
Coding
SciCode
Python programming for scientific computing
43.4%
Better than 80% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
36.4%
Better than 83% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
20.0%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
85.0%
Last updated Aug 16, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Gemma 4 31B Thinking with similar models from the same provider or model family.
Gemma 4 31B
google/gemma-4-31b-itGoogle's Gemma 4 31B instruction-tuned model for heavier reasoning, coding, agentic workflows, and long-context multimodal understanding. This route keeps tokenizer thinking disabled for faster direct answers.
Gemma 4 26B A4B Uncensored
google/gemma-4-26b-a4b-uncensoredGemma 4 26B A4B Uncensored is an FP8 open-weight multimodal mixture-of-experts model LoRA-tuned for fewer refusals across chat, coding, tool use, and long-context work.
Gemma 4 26B A4B
google/gemma-4-26b-a4b-itGoogle's Gemma 4 26B A4B instruction-tuned model built for scalable reasoning, coding, long-context, and multimodal workflows. This route is tuned for faster direct answers while preserving multimodal and structured output support.
Gemma 4 26B A4B Thinking
google/gemma-4-26b-a4b-it:thinkingGoogle's Gemma 4 26B A4B instruction-tuned model with structured reasoning for more deliberate coding, multimodal analysis, and long-context problem solving.
Gemma 4 31B MeroMero v2
Gemma-4-31B-MeroMero-v2Gemma 4 31B MeroMero v2 is a LoRA finetune for emotive dialogue, relationship scenes, creative writing, and multimodal roleplay.
Gemma 4 31B IT TEE
TEE/gemma-4-31b-itGemma 4 31B Instruct is a 30.7B dense model with a long context window, multilingual performance, function calling, and configurable reasoning. Running inside a TEE (Trusted Execution Environment), with provider attestation support.