Private AI
Nvidia's Nemotron 3 Ultra 550B A55B model from the Nemotron 3 family. It uses a hybrid Mamba-Transformer MoE architecture. Provider-specific context limits vary, with the longest current route supporting up to 1M context.
Added Jun 4, 2026
Model weightsContext Window
1.0M
Max Output
65.5K
Avg output tokens (7d)
1.4K tokens
Input Price (Auto)
$0.53/1M
Output Price (Auto)
$2.63/1M
Cache Read (Auto)
$0.26/1M
Capabilities
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
37.8
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Coding Index
49.3
Agentic Index
27.4
GPQA Diamond
Graduate-level scientific reasoning
86.7%
Better than 86% of models compared
HLE
Humanity's Last Exam
26.6%
Better than 85% of models compared
IFBench
Instruction-following benchmark
81.4%
Better than 99% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
83.3%
Better than 73% of models compared
AA-LCR
Long context reasoning evaluation
67.0%
Better than 86% of models compared
GDPval-AA
Economically valuable tasks
33.1%
CritPt
Research-level physics reasoning
3.1%
SciCode
Python programming for scientific computing
39.9%
Better than 72% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
36.4%
AA-Omniscience Accuracy
Proportion of correctly answered questions
21.5%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
28.5%
Last updated Aug 4, 2026
Artificial AnalysisBetter than 83% of models compared