Google's Gemma 4 12B Instruct is an open-weight multimodal model for text, image, audio, and video understanding, with tool calling and structured output support.
Added Aug 1, 2026
Context Window
262.1K
Max Output
32.8K
Avg output tokens (7d)
164 tokens
Input Price (Auto)
$0.050/1M
Output Price (Auto)
$0.25/1M
Cache Read (Auto)
$0.025/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
9.4
Agentic work
T²-Bench Telecom (legacy)
Legacy fallback · Conversational AI agents in dual-control scenarios
31.9%
Better than 40% of models compared
Document reasoning
AA-LCR v1.1
Long context reasoning with updated grading
35.0%
Better than 36% of models compared
Reasoning
HLE
Humanity's Last Exam
6.3%
Better than 40% of models compared
IFBench
Instruction-following benchmark
45.2%
Better than 52% of models compared
Coding
Terminal-Bench Hard (legacy)
Legacy fallback · Agentic coding and terminal use
11.4%
Better than 47% of models compared
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
66.1%
Better than 42% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
35.0%
Better than 36% of models compared
Last updated Sep 11, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Gemma 4 12B Instruct with similar models from the same provider or model family.
Gemma 4 12B Semancer
gemma-4-12b-it-semancerGemma 4 12B Semancer is an open-weight finetune by UnstableLlama trained on a custom occult-philosophy dataset. It is designed for in-depth philosophical discussion of topics such as truth, free will, consciousness, and meaning, with answers developed from first principles. It supports image understanding, tool calling, optional reasoning, and a 131,072-token context window.
Gemma 4 12B StationKeeper
gemma-4-12b-it-station-keeperGemma 4 12B StationKeeper is an open-weight roleplay finetune with image understanding, tool calling, optional reasoning, and a 131,072-token context window.
Gemma 4 E2B Instruct
gemma-4-e2b-itGoogle's Gemma 4 E2B Instruct is a compact open-weight multimodal model with 2B active parameters, supporting text, image, audio, video, tools, and structured output.
Gemma 4 E4B Instruct
gemma-4-e4b-itGoogle's Gemma 4 E4B Instruct is an efficient open-weight multimodal model with 4B active parameters, supporting text, image, audio, video, tools, and structured output.
Gemma 4 26B A4B Chimera X
gemma-4-26b-a4b-it-chimeraxChimera X is a Gemma 4 26B A4B multimodal mixture-of-experts fine-tune for creative writing, expressive dialogue, and roleplay.
Gemma 4 26B A4B Dark Soul
gemma-4-26b-a4b-it-darksoulDark Soul is a Gemma 4 26B A4B multimodal mixture-of-experts fine-tune for creative writing, expressive dialogue, and roleplay.