Private AI
Browse and discover the best AI language models for conversations, coding, and creative writing.
Claude Opus 5.5 is Anthropic's flagship model for agentic coding, computer use, complex knowledge work, and long-running tasks, with improved efficiency and more natural communication.
Features
Context
1M
Max Output
128K
Date Added
Sep 22, 2026
Avg output
2.5K
Not included in subscription
Input
$4.00/1M
Output
$20.00/1M
Est. per message
$0.0533
Cache:
Read $0.20/1M · Write $5.00/1M (5m) / $8.00/1M (1h)
Abliteration.ai's default unrestricted large text reasoning model is derived from GLM-5.3 and weight-modified to reduce refusals compared with the base model. It supports native tool calling, structured output, automatic prompt caching, and a one-million-token context window.
Features
Context
1M
Max Output
1M
Date Added
Aug 31, 2026
Avg output
1K
Not included in subscription
Input
$3.00/1M
Output
$5.00/1M
Est. per message
$0.0082
Cache:
Read $0.30/1M
View Providers
GLM-5.3 with higher reasoning enabled for harder long-horizon coding, autonomous agent workflows, and complex engineering tasks.
Features
Context
1M
Max Output
131.1K
Date Added
Aug 14, 2026
Avg output
1.9K
Included in subscription
Input
$1.00/1M
Output
$3.20/1M
Est. per message
$0.0069
Cache:
Read $0.20/1M
View Providers
GLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model with provider-dependent vision support, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.
Features
Context
1M
Max Output
32.8K
Parameters
320B / 18B
Date Added
Jul 29, 2026
Avg output
818
Included in subscription
Input
$0.20/1M
Output
$0.80/1M
Est. per message
$0.0009
Cache:
Read $0.07/1M
View Providers
Claude Fable 5.1 improves on Fable 5 across agentic coding, long-running workflows, front-end and visual code generation, finance, analysis, and knowledge work, with more concise plans and summaries. Anthropic retains prompts and outputs for 30 days; Zero Data Retention is not available.
Features
Context
1M
Max Output
128K
Date Added
Sep 1, 2026
Avg output
1.4K
Not included in subscription
Input
$10.00/1M
Output
$50.00/1M
Est. per message
$0.0783
Cache:
Read $0.25/1M · Write $12.50/1M (5m) / $20.00/1M (1h)
Kimi K3 is Moonshot AI's open-weight, always-thinking multimodal model for long-context reasoning, coding, tool use, and native image/video understanding.
Features
Context
1M
Max Output
1M
Date Added
Jul 16, 2026
Avg output
1.1K
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0128
Cache:
Read $0.20/1M
View Providers
An experimental uncensored version of the full GLM 5.3 reasoning model for chat, creative writing, coding, and tool use. It may show unusual behavior or inconsistent response quality.
Features
Context
1M
Max Output
131.1K
Parameters
744B / 40B
Date Added
Sep 26, 2026
Avg output
1.2K
Not included in subscription
Input
$1.25/1M
Output
$2.25/1M
Est. per message
$0.0040
Cache:
Read $0.30/1M
View Providers
GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.
Features
Context
1M
Max Output
131.1K
Date Added
Aug 14, 2026
Avg output
851
Included in subscription
Input
$1.00/1M
Output
$3.20/1M
Est. per message
$0.0037
Cache:
Read $0.20/1M
View Providers
Google's fast multimodal model for agentic workloads, including coding, tool use, image understanding, PDF and document extraction, audio, and video. Its capabilities, limits, reasoning behavior, and pricing currently mirror Gemini 3.7 Flash.
Features
Context
1M
Max Output
65.5K
Date Added
Sep 2, 2026
Avg output
1.4K
Not included in subscription
Input
$0.75/1M
Output
$3.75/1M
Est. per message
$0.0062
Cache:
Read $0.08/1M · Write $0.04/1M
Claude Opus 4.6 supports advanced coding, agentic workflows, and long-context analysis with up to a 1M-token context window.
Features
Context
1M
Max Output
128K
Date Added
Feb 5, 2026
Avg output
318
Not included in subscription
Input
$5.00/1M
Output
$25.00/1M
Est. per message
$0.0129
Cache:
Read $0.50/1M · Write $6.25/1M (5m) / $10.00/1M (1h)
Claude Sonnet 5.5 handles everyday coding, tool-based workflows, visual analysis, and polished writing with a 1M-token context window. This version skips up-front thinking and can reason between tool calls for faster responses; use Thinking for deeper planning and complex problems.
Features
Context
1M
Max Output
128K
Date Added
Sep 28, 2026
Avg output
1.1K
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0127
Cache:
Read $0.20/1M · Write $2.50/1M (5m) / $4.00/1M (1h)
Compatibility chat alias that routes to GPT 5.6 Sol, the flagship GPT-5.6 tier for complex reasoning, coding, and agentic workflows.
Features
Context
1.1M
Max Output
128K
Date Added
May 3, 2026
Avg output
682
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0088
Cache:
Read $0.20/1M · Write $2.50/1M
>272k input tier: Input $4.00/1M · Output $15.00/1M · Cache read $0.40/1M · Cache write $5.00/1M
Claude Sonnet 5.5 with adaptive thinking for complex coding, multi-step tool use, visual analysis, and detailed planning. Supports a 1M-token context window and five reasoning effort levels, with stronger thinking available when a task needs it.
Features
Context
1M
Max Output
128K
Date Added
Sep 28, 2026
Avg output
2.6K
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0277
Cache:
Read $0.20/1M · Write $2.50/1M (5m) / $4.00/1M (1h)
OpenAI's most capable model for the hardest end-to-end work, with state-of-the-art performance in computer use, browsing, software engineering, science, and professional tasks. Astra is designed to carry long, multistep workflows across code, browsers, and professional software through to a finished result.
Features
Context
1.1M
Max Output
128K
Date Added
Sep 3, 2026
Avg output
465
Not included in subscription
Input
$10.00/1M
Output
$50.00/1M
Est. per message
$0.0333
Cache:
Read $1.00/1M · Write $12.50/1M
>272k input tier: Input $20.00/1M · Output $75.00/1M · Cache read $2.00/1M · Cache write $25.00/1M
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, agentic workflows, command-line work, and multi-step professional tasks.
Features
Context
1.1M
Max Output
128K
Date Added
Jul 9, 2026
Avg output
733
Not included in subscription
Input
$2.00/1M
Output
$10.00/1M
Est. per message
$0.0093
Cache:
Read $0.20/1M · Write $2.50/1M
>272k input tier: Input $4.00/1M · Output $15.00/1M · Cache read $0.40/1M · Cache write $5.00/1M
ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
Features
Context
1M
Max Output
131.1K
Parameters
320B / 18B
Date Added
Aug 26, 2026
Avg output
1.3K
Included in subscription
Input
$0.10/1M
Output
$0.30/1M
Est. per message
$0.0005
Cache:
Read $0.03/1M
View Providers
DeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This is a rate-limited beta with limited capacity, intended for testing rather than production use.
Features
Context
1M
Max Output
384K
Date Added
Sep 8, 2026
Avg output
1.4K
Included in subscription
Input
$0.13/1M
Output
$0.52/1M
Est. per message
$0.0009
Cache:
Read $0.0060/1M
View Providers
Qwen 3.8 27B Uncensored is an FP8 open-weight multimodal model LoRA-tuned for fewer refusals across chat, coding, reasoning, tool use, and long-context work.
Features
Context
524.3K
Max Output
32.8K
Parameters
27B
Date Added
Jul 29, 2026
Avg output
364
Included in subscription
Input
$0.15/1M
Output
$1.20/1M
Est. per message
$0.0006
Cache:
Read $0.13/1M
View Providers