Hermes 4 Large with thinking enabled. Streams visible reasoning before the final answer and supports structured outputs.
Context Window
128.0K
Max Output
8.2K
Avg output tokens (7d)
1.1K tokens
Input Price (Auto)
$0.30/1M
Output Price (Auto)
$1.20/1M
Cache Read (Auto)
$0.15/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
8.8
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
72.7%
Better than 55% of models compared
HLE
Humanity's Last Exam
10.9%
Better than 60% of models compared
IFBench
Instruction-following benchmark
32.7%
Better than 21% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
22.2%
Better than 24% of models compared
AA-LCR
Long context reasoning evaluation
22.7%
Better than 31% of models compared
CritPt
Research-level physics reasoning
0.3%
Coding
SciCode
Python programming for scientific computing
25.2%
Better than 29% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
11.4%
Better than 47% of models compared
LiveCodeBench
Contamination-free coding benchmark
68.6%
Better than 77% of models compared
Math
AIME 2025
American Invitational Mathematics Examination 2025
69.7%
Better than 64% of models compared
Knowledge
MMLU-Pro
Professional and academic subject knowledge
82.9%
Better than 82% of models compared
AA-Omniscience Accuracy
Proportion of correctly answered questions
30.1%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
94.5%
Last updated Aug 16, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Hermes 4 Large (Thinking) with similar models from the same provider or model family.
Hermes 4 Large
nousresearch/hermes-4-405bAdvanced reasoning model built on Llama-3.1-405B with hybrid thinking modes. Features internal deliberation capabilities, excels at math, code, STEM, and logical reasoning while supporting structured outputs with improved steerability and neutral alignment.
Hermes 3 70B
nousresearch/hermes-3-llama-3.1-70bHermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, better roleplaying, reasoning, multi-turn conversation, and long context coherence. This 70B model is a competitive finetune of Llama-3.1-70B focused on aligning LLMs to the user with powerful steering capabilities.
Hermes 4 (Thinking)
NousResearch/Hermes-4-70B:thinkingHermes 4 70B with thinking enabled. Emits explicit reasoning content before final answer when streamed.
Hermes 4 Medium
nousresearch/hermes-4-70bEfficient reasoning model based on Llama-3.1-70B. Offers hybrid thinking capabilities with strong performance in math, code, and logical reasoning tasks. Supports structured outputs and JSON mode with enhanced steerability.
Hermes High
hermes-highCurrently points to Claude 4.7 Opus. Highest tier for Hermes-style agent work: most expensive, strongest intelligence, and best suited for hard multi-step tasks.
Hermes Low
hermes-lowCurrently points to Gemini 3.1 Flash Lite. Lowest tier for Hermes-style agent work: lowest cost for routine, fast, high-volume tasks.