Hermes 4 70B with thinking enabled. Emits explicit reasoning content before final answer when streamed.
Added Sep 17, 2025
Model weightsContext Window
128.0K
Max Output
8.2K
Input Price (Auto)
$0.20/1M
Output Price (Auto)
$0.40/1M
Cache Read (Auto)
$0.10/1M
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
9.9
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
69.9%
Better than 50% of models compared
HLE
Humanity's Last Exam
8.8%
Better than 54% of models compared
IFBench
Instruction-following benchmark
31.3%
Better than 16% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
22.5%
Better than 24% of models compared
AA-LCR
Long context reasoning evaluation
8.0%
Better than 18% of models compared
CritPt
Research-level physics reasoning
0.0%
Coding
SciCode
Python programming for scientific computing
34.1%
Better than 50% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
4.5%
Better than 29% of models compared
LiveCodeBench
Contamination-free coding benchmark
65.3%
Better than 73% of models compared
Math
AIME 2025
American Invitational Mathematics Examination 2025
68.7%
Better than 63% of models compared
Knowledge
MMLU-Pro
Professional and academic subject knowledge
81.1%
Better than 73% of models compared
AA-Omniscience Accuracy
Proportion of correctly answered questions
23.5%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
95.4%
Last updated Aug 16, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Hermes 4 (Thinking) with similar models from the same provider or model family.
Hermes 3 70B
nousresearch/hermes-3-llama-3.1-70bHermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, better roleplaying, reasoning, multi-turn conversation, and long context coherence. This 70B model is a competitive finetune of Llama-3.1-70B focused on aligning LLMs to the user with powerful steering capabilities.
Hermes 4 Medium
nousresearch/hermes-4-70bEfficient reasoning model based on Llama-3.1-70B. Offers hybrid thinking capabilities with strong performance in math, code, and logical reasoning tasks. Supports structured outputs and JSON mode with enhanced steerability.
Hermes 4 Large
nousresearch/hermes-4-405bAdvanced reasoning model built on Llama-3.1-405B with hybrid thinking modes. Features internal deliberation capabilities, excels at math, code, STEM, and logical reasoning while supporting structured outputs with improved steerability and neutral alignment.
Hermes 4 Large (Thinking)
nousresearch/hermes-4-405b:thinkingHermes 4 Large with thinking enabled. Streams visible reasoning before the final answer and supports structured outputs.
Hermes High
hermes-highCurrently points to Claude 4.7 Opus. Highest tier for Hermes-style agent work: most expensive, strongest intelligence, and best suited for hard multi-step tasks.
Hermes Low
hermes-lowCurrently points to Gemini 3.1 Flash Lite. Lowest tier for Hermes-style agent work: lowest cost for routine, fast, high-volume tasks.