DiffusionGemma is a high-speed diffusion-based version of Gemma 4 26B A4B. It supports optional reasoning and a 262,144-token context window.
Added Sep 19, 2026
Context Window
262.1K
Max Output
32.8K
Input Price (Auto)
$0.050/1M
Output Price (Auto)
$0.15/1M
Cache Read (Auto)
$0.025/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
9.5
Coding Index
19.7
Agentic work
GDPval-AA v2
Economically valuable tasks (Elo)
501 Elo
Better than 21% of models compared
Document reasoning
AA-LCR v1.1
Long context reasoning with updated grading
19.7%
Better than 23% of models compared
Reasoning
HLE
Humanity's Last Exam
10.8%
Better than 55% of models compared
IFBench
Instruction-following benchmark
59.5%
Better than 71% of models compared
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
66.9%
Better than 43% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
19.7%
Better than 23% of models compared
Last updated Sep 19, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare DiffusionGemma with similar models from the same provider or model family.
Gemma 4 26B A4B Cybersecurity
google/gemma-4-26b-a4b-it-cybersecurityGemma 4 26B A4B Cybersecurity is a cybersecurity-focused variant based on the uncensored model, with provider moderation for illegal activities. It supports optional reasoning, image understanding, tool calling, and a 262,144-token context window.
Gemma 4 26B A4B
google/gemma-4-26b-a4b-itGoogle's Gemma 4 26B A4B instruction-tuned model built for scalable reasoning, coding, long-context, and multimodal workflows. This route is tuned for faster direct answers while preserving multimodal and structured output support. Requests containing video cost $0.10 per million input tokens and $0.40 per million output tokens.
Gemma 4 26B A4B Thinking
google/gemma-4-26b-a4b-it:thinkingGoogle's Gemma 4 26B A4B instruction-tuned model with structured reasoning for more deliberate coding, multimodal analysis, and long-context problem solving. Requests containing video cost $0.10 per million input tokens and $0.40 per million output tokens.
Gemma 4 31B
google/gemma-4-31b-itGoogle's Gemma 4 31B instruction-tuned model for heavier reasoning, coding, agentic workflows, and long-context multimodal understanding. This route keeps tokenizer thinking disabled for faster direct answers. Requests containing video cost $0.14 per million input tokens and $0.40 per million output tokens.
Gemma 4 31B Thinking
google/gemma-4-31b-it:thinkingGoogle's Gemma 4 31B instruction-tuned model with thinking explicitly enabled, exposing reasoning traces for complex multimodal and coding workflows. Requests containing video cost $0.14 per million input tokens and $0.40 per million output tokens.
Gemma 4 31B Split-Untied
google/gemma4-31b-splituntiedBlazed-Forge's Split-Untied is a text-only Gemma 4 31B community finetune with an untied BF16 output head, built for creative writing, roleplay, expressive dialogue, and tool use.