MiniMax M3 is the non-thinking route for MiniMax's open-weights frontier model, built for coding, agent workflows, tool use, and multimodal understanding from step zero. It keeps native thinking disabled for faster direct answers. MiniMax reports 59.0% on SWE-Bench Pro and 66.0% on Terminal Bench 2.1, with Sparse Attention designed to scale context to 1M. It starts with a 512K context cap on NanoGPT for now.
Added Jun 1, 2026
Model weightsContext Window
512.0K
Max Output
80.0K
Avg output tokens (7d)
401 tokens
Input Price (Auto)
$0.23/1M
Output Price (Auto)
$0.96/1M
Cache Read (Auto)
$0.050/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
29.6
Coding Index
58.6
Agentic Index
30.8
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
21.3%
Better than 48% of models compared
AutomationBench-AA Tasks Completed
Fully completed workflows without guardrail violations
4.4%
Better than 21% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
1096 Elo
Better than 62% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
1304 Elo
Better than 71% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
9.8%
Better than 43% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
83.0%
Better than 97% of models compared
Reasoning
HLE
Humanity's Last Exam
39.0%
Better than 90% of models compared
IFBench
Instruction-following benchmark
82.9%
Better than 99% of models compared
CritPt
Research-level physics reasoning
3.7%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
2.0%
Better than 52% of models compared
SciCode
Python programming for scientific computing
47.1%
Better than 49% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
16.7%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
18.4%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
92.9%
Better than 96% of models compared
Terminal-Bench Hard (legacy)
Agentic coding and terminal use
42.4%
Better than 89% of models compared
T²-Bench Telecom (legacy)
Conversational AI agents in dual-control scenarios
88.9%
Better than 83% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
83.0%
Better than 97% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
40.2%
Last updated Sep 10, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare MiniMax M3 with similar models from the same provider or model family.
MiniMax M3 Thinking
minimax/minimax-m3:thinkingMiniMax M3 Thinking is the adaptive-thinking version of MiniMax's open-weights frontier model for coding, agent workflows, tool use, long-context tasks, and native multimodal understanding. MiniMax reports 59.0% on SWE-Bench Pro and 66.0% on Terminal Bench 2.1, with Sparse Attention designed to scale context to 1M. It starts with a 512K context cap on NanoGPT for now.
MiniMax Latest
minimax/minimax-latestCompatibility alias that routes to the newest MiniMax text model. Currently routes to MiniMax M3 (adaptive thinking).
MiniMax M2.7
minimax/minimax-m2.7MiniMax M2.7 is the first model deeply involved in iterating on its own training. It excels in real-world software engineering (SWE-Pro 56.22%), end-to-end project delivery (VIBE-Pro 55.6%), and complex office workflows with strong tool-use compliance and agentic capabilities.
MiniMax M2.7 Turbo
minimax/minimax-m2.7-turboMiniMax M2.7 Turbo is the highspeed and higher priced route for M2.7.
MiniMax M2.5
minimax/minimax-m2.5MiniMax M2.5 is a productivity-focused flagship model that builds on M2.1 with stronger coding and real-world office workflow performance (Word, Excel, PowerPoint), plus better tool-use planning and token efficiency.
MiniMax M2.1
minimax/minimax-m2.1MiniMax M2.1 builds on M2 with enhanced context understanding and improved complex tool use. 230B parameter MoE model (10B active) optimized for agentic workflows and long-horizon tasks.