Private AI
Browse and discover the best AI language models for conversations, coding, and creative writing.
Claude Opus 5.5 is Anthropic's flagship model for agentic coding, computer use, complex knowledge work, and long-running tasks, with improved efficiency and more natural communication.
Features
Context
1M
Max Output
128K
Date Added
Sep 22, 2026
Avg output
3.2K
Not included in subscription
Input
$4.00/1M
Output
$20.00/1M
Est. per message
$0.0684
Cache:
Read $0.20/1M · Write $5.00/1M (5m) / $8.00/1M (1h)
Abliteration.ai's default unrestricted large text reasoning model is derived from GLM-5.3 and weight-modified to reduce refusals compared with the base model. It supports native tool calling, structured output, automatic prompt caching, and a one-million-token context window.
Features
Context
1M
Max Output
1M
Date Added
Aug 31, 2026
Avg output
878
Not included in subscription
Input
$3.00/1M
Output
$5.00/1M
Est. per message
$0.0074
Cache:
Read $0.30/1M
View Providers
An experimental uncensored version of the full GLM 5.3 reasoning model for chat, creative writing, coding, and tool use. It may show unusual behavior or inconsistent response quality.
Features
Context
1M
Max Output
131.1K
Parameters
744B / 40B
Date Added
Sep 26, 2026
Avg output
1.2K
Not included in subscription
Input
$1.25/1M
Output
$2.25/1M
Est. per message
$0.0041
Cache:
Read $0.30/1M
View Providers
GLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model with provider-dependent vision support, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.
Features
Context
1M
Max Output
32.8K
Parameters
320B / 18B
Date Added
Jul 29, 2026
Avg output
745
Included in subscription
Input
$0.20/1M
Output
$0.80/1M
Est. per message
$0.0008
Cache:
Read $0.07/1M
View Providers
Claude Fable 5.1 improves on Fable 5 across agentic coding, long-running workflows, front-end and visual code generation, finance, analysis, and knowledge work, with more concise plans and summaries. Anthropic retains prompts and outputs for 30 days; Zero Data Retention is not available.
Features
Context
1M
Max Output
128K
Date Added
Sep 1, 2026
Avg output
1.4K
Not included in subscription
Input
$10.00/1M
Output
$50.00/1M
Est. per message
$0.0815
Cache:
Read $0.25/1M · Write $12.50/1M (5m) / $20.00/1M (1h)
GLM-5.3 with higher reasoning enabled for harder long-horizon coding, autonomous agent workflows, and complex engineering tasks.
Features
Context
1M
Max Output
131.1K
Date Added
Aug 14, 2026
Avg output
2K
Included in subscription
Input
$1.00/1M
Output
$3.20/1M
Est. per message
$0.0074
Cache:
Read $0.20/1M
View Providers
Kimi K3 is Moonshot AI's open-weight, always-thinking multimodal model for long-context reasoning, coding, tool use, and native image/video understanding.
Features
Context
1M
Max Output
1M
Date Added
Jul 16, 2026
Avg output
1K
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0122
Cache:
Read $0.20/1M
View Providers
GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.
Features
Context
1M
Max Output
131.1K
Date Added
Aug 14, 2026
Avg output
935
Included in subscription
Input
$1.00/1M
Output
$3.20/1M
Est. per message
$0.0040
Cache:
Read $0.20/1M
View Providers
OpenAI's most capable model for the hardest end-to-end work, with state-of-the-art performance in computer use, browsing, software engineering, science, and professional tasks. Astra is designed to carry long, multistep workflows across code, browsers, and professional software through to a finished result.
Features
Context
1.1M
Max Output
128K
Date Added
Sep 3, 2026
Avg output
480
Not included in subscription
Input
$10.00/1M
Output
$50.00/1M
Est. per message
$0.0340
Cache:
Read $1.00/1M · Write $12.50/1M
>272k input tier: Input $20.00/1M · Output $75.00/1M · Cache read $2.00/1M · Cache write $25.00/1M
Google's fast multimodal model for agentic workloads, including coding, tool use, image understanding, PDF and document extraction, audio, and video. Its capabilities, limits, reasoning behavior, and pricing currently mirror Gemini 3.7 Flash.
Features
Context
1M
Max Output
65.5K
Date Added
Sep 2, 2026
Avg output
1.5K
Not included in subscription
Input
$0.75/1M
Output
$3.75/1M
Est. per message
$0.0063
Cache:
Read $0.08/1M · Write $0.04/1M
MiMo V2.6 Flash Uncensored with maximum thinking enabled. Supports a 1M-token context window and tool calling.
Features
Context
1M
Max Output
65.5K
Parameters
309B / 15B
Date Added
Sep 26, 2026
Avg output
1.2K
Not included in subscription
Input
$0.50/1M
Output
$1.50/1M
Est. per message
$0.0023
Cache:
Read $0.20/1M
GPT 6 Astra in Pro reasoning mode. Uses additional model work for difficult tasks, with higher latency and token usage at the same per-token rates. Reasoning effort remains independently configurable.
Features
Context
1.1M
Max Output
128K
Date Added
Sep 4, 2026
Avg output
2.1K
Not included in subscription
Input
$10.00/1M
Output
$50.00/1M
Est. per message
$0.1158
Cache:
Read $1.00/1M · Write $12.50/1M
>272k input tier: Input $20.00/1M · Output $75.00/1M · Cache read $2.00/1M · Cache write $25.00/1M
Claude Sonnet 5.5 handles everyday coding, tool-based workflows, visual analysis, and polished writing with a 1M-token context window. This version skips up-front thinking and can reason between tool calls for faster responses; use Thinking for deeper planning and complex problems.
Features
Context
1M
Max Output
128K
Date Added
Sep 28, 2026
Avg output
1.4K
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0157
Cache:
Read $0.10/1M · Write $2.50/1M (5m) / $4.00/1M (1h)
Claude Sonnet 5.5 with adaptive thinking for complex coding, multi-step tool use, visual analysis, and detailed planning. Supports a 1M-token context window and five reasoning effort levels, with stronger thinking available when a task needs it.
Features
Context
1M
Max Output
128K
Date Added
Sep 28, 2026
Avg output
3.3K
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0349
Cache:
Read $0.10/1M · Write $2.50/1M (5m) / $4.00/1M (1h)
Compatibility alias that routes to the newest version of Claude Opus. Currently routes to Claude Opus 5.5.
Features
Context
1M
Max Output
128K
Date Added
Mar 29, 2026
Not included in subscription
Input
$4.00/1M
Output
$20.00/1M
Est. per message
$0.0140
Cache:
Read $0.20/1M · Write $5.00/1M (5m) / $8.00/1M (1h)
GPT-6.1 Sol delivers near-Astra performance at a lower cost for complex coding, computer use, and professional work. It supports text and image input, tool calling, structured outputs, and reasoning from low through max.
Features
Context
1.1M
Max Output
128K
Date Added
Sep 29, 2026
Avg output
543
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0074
Cache:
Read $0.10/1M · Write $2.50/1M
>272k input tier: Input $4.00/1M · Output $15.00/1M · Cache read $0.20/1M · Cache write $5.00/1M
Claude Opus 4.6 supports advanced coding, agentic workflows, and long-context analysis with up to a 1M-token context window.
Features
Context
1M
Max Output
128K
Date Added
Feb 5, 2026
Avg output
258
Not included in subscription
Input
$5.00/1M
Output
$25.00/1M
Est. per message
$0.0115
Cache:
Read $0.50/1M · Write $6.25/1M (5m) / $10.00/1M (1h)
ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
Features
Context
1M
Max Output
131.1K
Parameters
320B / 18B
Date Added
Aug 26, 2026
Avg output
1.3K
Included in subscription
Input
$0.10/1M
Output
$0.30/1M
Est. per message
$0.0005
Cache:
Read $0.03/1M
View Providers