Private AI
Browse and discover the best AI language models for conversations, coding, and creative writing.
DeepSeek V4 Flash Vision Exp
An experimental vision-enabled DeepSeek V4 Flash model that adds image understanding while retaining the text, reasoning, coding, tool-calling, and agent capabilities of the base model. This route is served directly by DeepSeek, so privacy and logging guarantees are limited.
Features
Context
1.0M
Max Output
384.0K
Parameters
284B / 13B
Date Added
Aug 21, 2026
Pricing
Input:
$0.22/1M
Output:
$0.66/1M
Cache:
Read $0.0070/1M
Est./msg:
$0.0006
Subscription
Included in subscription
Ox Alpha
Ox Alpha is an experimental stealth model with long-context reasoning, vision, tool calling, and structured output support. Prompts and responses are logged and retained by the provider.
Features
Context
1.0M
Max Output
32.8K
Date Added
Aug 21, 2026
Pricing
Input:
$0.05/1M
Output:
$0.05/1M
Cache:
Read $0.03/1M
Est./msg:
$0.0001
Subscription
Included in subscription
Qwen 3.6 35B A3B Uncensored
Qwen 3.6 35B A3B Uncensored is an FP8 open-weight mixture-of-experts model tuned for fewer refusals across chat, coding, tool use, and multimodal tasks.
Qwen 3.8 27B Uncensored
Qwen 3.8 27B Uncensored is an FP8 open-weight multimodal model tuned for fewer refusals across chat, coding, tool use, and long-context work.
Ornith 1.5 35B
Ornith 1.5 35B A3B is an open-weight mixture-of-experts model for agentic coding, tool use, image understanding, and long-context work. This variant disables thinking for faster direct responses.
Ornith 1.5 35B Thinking
Ornith 1.5 35B A3B is an open-weight mixture-of-experts model for agentic coding, reasoning, tool use, image understanding, and long-context work. This variant enables thinking by default.
Qwen3.5 0.8B
Qwen3.5 0.8B is a lightweight open-weight multimodal model from Alibaba for fast reasoning, visual understanding, tool use, and JSON output.
Qwen3.5 2B
Qwen3.5 2B is a small open-weight multimodal model from Alibaba for efficient reasoning, coding, visual understanding, tool use, and JSON output.
Qwen3.5 4B
Qwen3.5 4B is a compact open-weight multimodal model from Alibaba for reasoning, coding, visual understanding, tool use, and structured output.
Qwen3.8 27B
Qwen3.8 27B is an open-weight multimodal model from Alibaba for coding, visual understanding, tool use, and structured output. This variant keeps thinking disabled for faster direct responses.
Qwen3.8 27B Thinking
Qwen3.8 27B is an open-weight multimodal model from Alibaba for reasoning, coding, visual understanding, tool use, and structured output. This variant enables thinking by default.
Dots3-Note Preview
Dots Studio's open-weight multimodal Mixture-of-Experts model activates 16B of 280B parameters for long-context reasoning, coding, visual and document understanding, tool use, and long-horizon agent workflows. Prompts and completions of this model may be logged.
Kimi K2.7 Code TEE
Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built for long-horizon software engineering workflows. It supports native image input, tool calling, and forced thinking mode while running inside a Trusted Execution Environment (TEE) with attestation support.
GLM 5.3 Preview
A preview of GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses. While in preview, prompts and responses may be logged and used for training.
GLM 5.3 Preview Thinking
The GLM-5.3 preview with higher reasoning enabled for harder long-horizon coding, autonomous agent workflows, and complex engineering tasks. While in preview, prompts and responses may be logged and used for training.
Gemini 3.7 Flash
Google's frontier-performance Flash model for multimodal and agentic workloads, including coding, tool use, image understanding, PDF and document extraction, audio, and video. Google reports 65.3% on DeepSWE v1.1, up from 49.0% for Gemini 3.6 Flash, and 34% on GDP.pdf, up from 14%.
Muse Glimmer 30B TEE
Meta's Muse Glimmer 30B is a dense, open-weight multimodal model for long-horizon agentic and coding workflows. Running inside a TEE (Trusted Execution Environment), with provider attestation support.
ByteDance Seed 2.1 Turbo
ByteDance Seed 2.1 Turbo is a multimodal model for coding and long-horizon agent workflows, including end-to-end software delivery and multi-step task execution. It supports text, image, and video input with a 262k context window.