StepFun's most capable open-source reasoning model with visible reasoning traces. Built on a sparse Mixture-of-Experts architecture with 196B total parameters and only 11B active per token, it achieves frontier-level performance in math, logic, and agentic coding while reaching up to 350 tokens/sec. Supports 256K context. NOTE: This model runs via StepFun, which may log and train on your prompts.
Added Feb 2, 2026
Model weightsContext Window
256.0K
Max Output
256.0K
Avg output tokens (7d)
1.2K tokens
Input Price (Auto)
$0.10/1M
Output Price (Auto)
$0.30/1M
Cache Read (Auto)
$0.050/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
17.0
Agentic work
T²-Bench Telecom (legacy)
Legacy fallback · Conversational AI agents in dual-control scenarios
94.4%
Better than 92% of models compared
Document reasoning
AA-LCR v1.1
Long context reasoning with updated grading
63.0%
Better than 57% of models compared
Reasoning
HLE
Humanity's Last Exam
21.1%
Better than 73% of models compared
IFBench
Instruction-following benchmark
64.6%
Better than 76% of models compared
CritPt
Research-level physics reasoning
2.5%
Coding
Terminal-Bench Hard (legacy)
Legacy fallback · Agentic coding and terminal use
27.3%
Better than 70% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
23.6%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
85.7%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
83.1%
Better than 73% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
50.0%
Better than 47% of models compared
Last updated Sep 5, 2026, 12:00 AM
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Step 3.5 Flash with similar models from the same provider or model family.
Step 3.5 Flash 2603
stepfun-ai/step-3.5-flash-2603Step 3.5 Flash 2603 is optimized for high-frequency agentic and coding workflows with improved token efficiency and faster reasoning. NOTE: This model runs via StepFun, which may log and train on your prompts.
Step 3.7 Flash Thinking
stepfun/step-3.7-flash:thinkingStep 3.7 Flash Thinking is StepFun's high-efficiency multimodal MoE model with visible reasoning enabled for deeper agentic coding, long-context reasoning, tool use, and native image/video understanding. ⚠️ Note: This model routes through StepFun, so privacy and logging guarantees may be limited.
DeepSeek V4.1 Flash TEE
TEE/deepseek-v4.1-flashDeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.
Agnes 3.0 Flash
agnes-3.0-flashAgnes 3.0 Flash is a low-cost model for coding, tool use, and multi-turn agent tasks. It supports text and image input, optional thinking, and a 512K-token context window.
Ling 3.0 Flash VL
inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL is inclusionAI's native multimodal Mixture-of-Experts model with 124B total parameters and 5.5B active parameters per token. It combines image and video understanding with reasoning and tool use for document analysis, charts, visual verification, and interface-based agent tasks. Thinking is enabled by default and can be turned off in settings.
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This is a rate-limited beta with limited capacity, intended for testing rather than production use.