Nvidia's latest Nemotron 3 Nano model with 30B total parameters (3B active) using hybrid Mamba-Transformer MoE architecture. Features excellent throughput and strong reasoning capabilities.
Added Mar 1, 2026
Model weightsContext Window
256.0K
Max Output
262.1K
Input Price (Auto)
$0.17/1M
Output Price (Auto)
$0.68/1M
Cache Read (Auto)
$0.085/1M
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
7.2
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
39.9%
Better than 15% of models compared
HLE
Humanity's Last Exam
4.6%
Better than 26% of models compared
IFBench
Instruction-following benchmark
37.5%
Better than 31% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
25.4%
Better than 29% of models compared
AA-LCR
Long context reasoning evaluation
8.7%
Better than 19% of models compared
GDPval-AA
Economically valuable tasks
0.0%
CritPt
Research-level physics reasoning
0.0%
Coding
SciCode
Python programming for scientific computing
23.0%
Better than 25% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
12.1%
Better than 48% of models compared
LiveCodeBench
Contamination-free coding benchmark
36.0%
Better than 44% of models compared
Math
AIME 2025
American Invitational Mathematics Examination 2025
13.3%
Better than 16% of models compared
Knowledge
MMLU-Pro
Professional and academic subject knowledge
57.9%
Better than 21% of models compared
AA-Omniscience Accuracy
Proportion of correctly answered questions
11.4%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
91.0%
Last updated Aug 16, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Nvidia Nemotron 3 Nano 30B with similar models from the same provider or model family.
Nvidia Nemotron 3 Nano Omni
nvidia/nemotron-3-nano-omni-30b-a3b-reasoningNvidia's Nemotron 3 Nano Omni 30B-A3B reasoning model. It accepts multimodal context on supported providers and returns text responses for perception and agentic workflows.
Nvidia Nemotron 3.5 Lightning
nvidia/nemotron-3.5-lightningNvidia Nemotron 3.5 Lightning is an open-weight 30B-A3B hybrid Mamba-Transformer mixture-of-experts model for high-throughput agentic, coding, tool-use, instruction-following, and long-context workloads. Thinking is disabled on this variant. Included in the NanoGPT subscription.
Nvidia Nemotron 3.5 Lightning Thinking
nvidia/nemotron-3.5-lightning:thinkingNvidia Nemotron 3.5 Lightning is an open-weight 30B-A3B hybrid Mamba-Transformer mixture-of-experts model for high-throughput agentic, coding, tool-use, instruction-following, and long-context workloads. This variant enables its reasoning trace. Included in the NanoGPT subscription.
Nvidia Nemotron 3 Ultra 550B
nvidia/nemotron-3-ultra-550b-a55bNvidia's Nemotron 3 Ultra 550B A55B model from the Nemotron 3 family. It uses a hybrid Mamba-Transformer MoE architecture. Provider-specific context limits vary, with the longest current route supporting up to 1M context.
Nvidia Nemotron 3 Ultra 550B Thinking
nvidia/nemotron-3-ultra-550b-a55b:thinkingNvidia's Nemotron 3 Ultra 550B A55B model from the Nemotron 3 family. It uses a hybrid Mamba-Transformer MoE architecture. Provider-specific context limits vary, with the longest current route supporting up to 1M context. Thinking enabled.
Nvidia Nemotron 3 Super 120B
nvidia/nemotron-3-super-120b-a12bNvidia's Nemotron 3 Super 120B A12B model from the March 2026 Nemotron 3 release. It uses a hybrid Mamba-Transformer MoE architecture and targets agentic and coding workloads with a 262K context window here.