Step 3.5 Flash
StepFun's most capable open-source reasoning model with visible reasoning traces. Built on a sparse Mixture-of-Experts architecture with 196B total parameters and only 11B active per token, it achieves frontier-level performance in math, logic, and agentic coding while reaching up to 350 tokens/sec. Supports 256K context. NOTE: This model runs via StepFun, which may log and train on your prompts.
- Reasoning
Added Feb 2, 2026
Model weightsPricing
Auto routing · per 1M tokens- Input
- $0.10
- Output
- $0.30
Specifications
- Context window
- 256K
- Max output
- 256K
- Parameters
- 196B / 11B
- Total / active
- Avg output (7d)
- 1K tokens
- Longer than 71% of models
Benchmarks
Benchmarks
Sourced from Artificial Analysis.
Intelligence Index
16.6
Agentic work
T²-Bench Telecom (legacy)
Legacy fallback · Conversational AI agents in dual-control scenarios
94.4%
Better than 92% of models compared
Document reasoning
AA-LCR v1.1
Long context reasoning with updated grading
50.0%
Better than 44% of models compared
Reasoning
HLE
Humanity's Last Exam
21.1%
Better than 69% of models compared
IFBench
Instruction-following benchmark
64.6%
Better than 76% of models compared
CritPt
Research-level physics reasoning
2.5%
Coding
Terminal-Bench Hard (legacy)
Legacy fallback · Agentic coding and terminal use
27.3%
Better than 69% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
23.6%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
85.7%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
83.1%
Better than 73% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
50.0%
Better than 44% of models compared
Last updated Oct 3, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…