Private AI
Browse and discover the best AI language models for conversations, coding, and creative writing.
Qwen3.8 Max
Qwen3.8 Max is Qwen's 2.4T-parameter flagship model for advanced reasoning, coding, knowledge work, full-stack development, data analysis, and long-running agent workflows. It supports text, image, and video input, selectable thinking effort, tool calling, structured output, and a near-million-token context window.
Features
Context
991.0K
Max Output
131.1K
Date Added
Aug 3, 2026
Pricing
Input:
$2.00/1M
Output:
$6.00/1M
Cache:
Read $0.25/1M · Write $2.50/1M
Est./msg:
$0.0050
Subscription
Not included in subscription
DeepSeek V4 Flash Latest
Compatibility alias that routes to the newest dated DeepSeek V4 Flash release. Currently routes to DeepSeek V4 Flash 0731. ⚠️ This route goes directly to DeepSeek, so privacy and logging guarantees are limited.
Features
Context
1.0M
Max Output
384.0K
Parameters
284B / 13B
Date Added
Aug 2, 2026
Pricing
Input:
$0.14/1M
Output:
$0.28/1M
Cache:
Read $0.0028/1M
Est./msg:
$0.0003
Subscription
Included in subscription
DeepSeek V4 Flash 0731 (Thinking)
DeepSeek V4 Flash 0731 Thinking enables reasoning by default on the re-post-trained Mixture-of-Experts model with a 1M-token context window, built for coding, reasoning, and agent workflows.
DeepSeek V4 Flash 0731 Cheaper
DeepSeek V4 Flash 0731 Cheaper is the same re-post-trained Mixture-of-Experts model with a 1M-token context window. This route goes directly to DeepSeek to use its lower cached-input pricing. ⚠️ Privacy and logging guarantees are limited.
DeepSeek V4 Flash 0731 Cheaper (Thinking)
DeepSeek V4 Flash 0731 Cheaper Thinking enables reasoning by default on the same re-post-trained Mixture-of-Experts model with a 1M-token context window. This route goes directly to DeepSeek to use its lower cached-input pricing. ⚠️ Privacy and logging guarantees are limited.
Gemma 4 12B Instruct
Google's Gemma 4 12B Instruct is an open-weight multimodal model for text, image, audio, and video understanding, with tool calling and structured output support.
Gemma 4 E2B Instruct
Google's Gemma 4 E2B Instruct is a compact open-weight multimodal model with 2B active parameters, supporting text, image, audio, video, tools, and structured output.
Gemma 4 E4B Instruct
Google's Gemma 4 E4B Instruct is an efficient open-weight multimodal model with 4B active parameters, supporting text, image, audio, video, tools, and structured output.
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a re-post-trained Mixture-of-Experts model with a 1M-token context window, built for coding, reasoning, and agent workflows.
Inkling Small
The direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.
Inkling Small Thinking
The reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.
Kimi K3 TEE
Kimi K3 is Moonshot AI's open-weight, always-thinking multimodal model for long-context reasoning, coding, tool use, and native image/video understanding. Running inside a TEE (Trusted Execution Environment), with provider attestation support.
Celeris 1
Celeris 1 is a diffusion language model built for ultra-low-latency classification, extraction, judging, query rewriting, and other short structured responses.
Qwen3.7 Flash
Qwen3.7 Flash is Qwen's fast multimodal model for coding, search and computer-use agents, visual understanding, object recognition, spatial reasoning, and stable end-to-end task execution.
Qwen3.7 Flash Thinking
Qwen3.7 Flash with thinking enabled for deeper multimodal reasoning, coding, search and computer-use agents, spatial reasoning, and multi-step task execution.
Claude Opus 5
Claude Opus 5 is Anthropic's flagship model for advanced coding, long-running agentic tasks, research, and complex knowledge work.
Fugu Ultra v1.1
Sakana AI's upgraded Fugu Ultra release with stronger coding, agentic task execution, and advanced reasoning through dynamic orchestration of frontier models.
Ling 3.0 Flash
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.