Private AI
Browse and discover the best AI language models for conversations, coding, and creative writing.
Claude Opus 5.5 is Anthropic's flagship model for agentic coding, computer use, complex knowledge work, and long-running tasks, with improved efficiency and more natural communication.
Features
Context
1M
Max Output
128K
Date Added
Sep 22, 2026
Avg output
2.6K
Not included in subscription
Input
$4.00/1M
Output
$20.00/1M
Est. per message
$0.0552
Cache:
Read $0.20/1M · Write $5.00/1M (5m) / $8.00/1M (1h)
Abliteration.ai's default unrestricted large text reasoning model is derived from GLM-5.3 and weight-modified to reduce refusals compared with the base model. It supports native tool calling, structured output, automatic prompt caching, and a one-million-token context window.
Features
Context
1M
Max Output
1M
Date Added
Aug 31, 2026
Avg output
985
Not included in subscription
Input
$3.00/1M
Output
$5.00/1M
Est. per message
$0.0079
Cache:
Read $0.30/1M
View Providers
GLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model with provider-dependent vision support, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.
Features
Context
1M
Max Output
32.8K
Parameters
320B / 18B
Date Added
Jul 29, 2026
Avg output
782
Included in subscription
Input
$0.20/1M
Output
$0.80/1M
Est. per message
$0.0008
Cache:
Read $0.07/1M
View Providers
GLM-5.3 with higher reasoning enabled for harder long-horizon coding, autonomous agent workflows, and complex engineering tasks.
Features
Context
1M
Max Output
131.1K
Date Added
Aug 14, 2026
Avg output
1.9K
Included in subscription
Input
$1.00/1M
Output
$3.20/1M
Est. per message
$0.0069
Cache:
Read $0.20/1M
View Providers
Claude Fable 5.1 improves on Fable 5 across agentic coding, long-running workflows, front-end and visual code generation, finance, analysis, and knowledge work, with more concise plans and summaries. Anthropic retains prompts and outputs for 30 days; Zero Data Retention is not available.
Features
Context
1M
Max Output
128K
Date Added
Sep 1, 2026
Avg output
1.4K
Not included in subscription
Input
$10.00/1M
Output
$50.00/1M
Est. per message
$0.0806
Cache:
Read $0.25/1M · Write $12.50/1M (5m) / $20.00/1M (1h)
An experimental uncensored version of the full GLM 5.3 reasoning model for chat, creative writing, coding, and tool use. It may show unusual behavior or inconsistent response quality.
Features
Context
1M
Max Output
131.1K
Parameters
744B / 40B
Date Added
Sep 26, 2026
Avg output
1.2K
Not included in subscription
Input
$1.25/1M
Output
$2.25/1M
Est. per message
$0.0039
Cache:
Read $0.30/1M
View Providers
Kimi K3 is Moonshot AI's open-weight, always-thinking multimodal model for long-context reasoning, coding, tool use, and native image/video understanding.
Features
Context
1M
Max Output
1M
Date Added
Jul 16, 2026
Avg output
1.1K
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0128
Cache:
Read $0.20/1M
View Providers
Google's fast multimodal model for agentic workloads, including coding, tool use, image understanding, PDF and document extraction, audio, and video. Its capabilities, limits, reasoning behavior, and pricing currently mirror Gemini 3.7 Flash.
Features
Context
1M
Max Output
65.5K
Date Added
Sep 2, 2026
Avg output
1.5K
Not included in subscription
Input
$0.75/1M
Output
$3.75/1M
Est. per message
$0.0062
Cache:
Read $0.08/1M · Write $0.04/1M
GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.
Features
Context
1M
Max Output
131.1K
Date Added
Aug 14, 2026
Avg output
789
Included in subscription
Input
$1.00/1M
Output
$3.20/1M
Est. per message
$0.0035
Cache:
Read $0.20/1M
View Providers
Claude Opus 4.6 supports advanced coding, agentic workflows, and long-context analysis with up to a 1M-token context window.
Features
Context
1M
Max Output
128K
Date Added
Feb 5, 2026
Avg output
295
Not included in subscription
Input
$5.00/1M
Output
$25.00/1M
Est. per message
$0.0124
Cache:
Read $0.50/1M · Write $6.25/1M (5m) / $10.00/1M (1h)
Claude Sonnet 5.5 handles everyday coding, tool-based workflows, visual analysis, and polished writing with a 1M-token context window. This version skips up-front thinking and can reason between tool calls for faster responses; use Thinking for deeper planning and complex problems.
Features
Context
1M
Max Output
128K
Date Added
Sep 28, 2026
Avg output
976
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0118
Cache:
Read $0.20/1M · Write $2.50/1M (5m) / $4.00/1M (1h)
OpenAI's most capable model for the hardest end-to-end work, with state-of-the-art performance in computer use, browsing, software engineering, science, and professional tasks. Astra is designed to carry long, multistep workflows across code, browsers, and professional software through to a finished result.
Features
Context
1.1M
Max Output
128K
Date Added
Sep 3, 2026
Avg output
408
Not included in subscription
Input
$10.00/1M
Output
$50.00/1M
Est. per message
$0.0304
Cache:
Read $1.00/1M · Write $12.50/1M
>272k input tier: Input $20.00/1M · Output $75.00/1M · Cache read $2.00/1M · Cache write $25.00/1M
ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
Features
Context
1M
Max Output
131.1K
Parameters
320B / 18B
Date Added
Aug 26, 2026
Avg output
1.3K
Included in subscription
Input
$0.10/1M
Output
$0.30/1M
Est. per message
$0.0005
Cache:
Read $0.03/1M
View Providers
MiMo V2.6 Flash Uncensored with maximum thinking enabled. Supports a 1M-token context window and tool calling.
Features
Context
1M
Max Output
65.5K
Parameters
309B / 15B
Date Added
Sep 26, 2026
Avg output
1.9K
Not included in subscription
Input
$0.50/1M
Output
$1.50/1M
Est. per message
$0.0033
Cache:
Read $0.20/1M
Qwen 3.8 27B Uncensored with thinking enabled for more deliberate creative work, coding, multimodal analysis, tool use, and long-context problem solving.
Features
Context
524.3K
Max Output
32.8K
Parameters
27B
Date Added
Jul 29, 2026
Avg output
1.3K
Included in subscription
Input
$0.15/1M
Output
$1.20/1M
Est. per message
$0.0017
Cache:
Read $0.13/1M
View Providers
GLM-5.2 with thinking enabled for harder long-horizon coding, autonomous agent workflows, complex engineering optimization, and real-world development tasks.
Features
Context
1M
Max Output
131.1K
Parameters
744B / 40B
Date Added
Jun 15, 2026
Avg output
1.6K
Included in subscription
Input
$0.42/1M
Output
$1.32/1M
Est. per message
$0.0026
Cache:
Read $0.08/1M
View Providers
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, agentic workflows, command-line work, and multi-step professional tasks.
Features
Context
1.1M
Max Output
128K
Date Added
Jul 9, 2026
Avg output
689
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0089
Cache:
Read $0.20/1M · Write $2.50/1M
>272k input tier: Input $4.00/1M · Output $15.00/1M · Cache read $0.40/1M · Cache write $5.00/1M
Gemini 3.1 Pro preview is built for tasks where simple answers are not enough. Stronger core reasoning for complex coding, math, and long-context workflows, with multimodal support and a reported 77.1% verified score on ARC-AGI-2. NOTE: Inputs > 200k tokens are charged at 2x input and 1.5x output rates.
Features
Context
1M
Max Output
65.5K
Date Added
Feb 19, 2026
Avg output
727
Not included in subscription
Input
$2.00/1M
Output
$12.00/1M
Est. per message
$0.0107
Cache:
Read $0.20/1M · Write $0.38/1M