Private AI
Browse and discover the best AI language models for conversations, coding, and creative writing.
Gemma 4 26B A4B Uncensored Thinking
Gemma 4 26B A4B Uncensored with thinking enabled for more deliberate coding, multimodal analysis, tool use, and long-context problem solving.
Features
Context
131.1K
Max Output
32.8K
Parameters
26B / 4B
Date Added
Aug 24, 2026
Pricing
Input:
$0.08/1M
Output:
$0.33/1M
Cache:
Read $0.04/1M
Est./msg:
$0.0002
Subscription
Included in subscription
View Providers
Gemma 4 31B MeroMero v2 Thinking
Gemma 4 31B MeroMero v2 with thinking enabled for more deliberate emotive dialogue, relationship scenes, creative writing, and multimodal roleplay.
Features
Context
65.5K
Max Output
32.8K
Parameters
31B
Date Added
Aug 24, 2026
Pricing
Input:
$0.08/1M
Output:
$0.33/1M
Cache:
Read $0.04/1M
Est./msg:
$0.0002
Subscription
Included in subscription
View Providers
Ornith 1.5 397B
Ornith 1.5 397B is an open-weight mixture-of-experts model built for agentic coding, tool use, image understanding, and long-context workflows. This variant disables thinking for faster direct responses.
Ornith 1.5 397B Thinking
Ornith 1.5 397B with thinking enabled for deeper agentic coding, reasoning, tool use, image understanding, and long-context workflows.
Ornith 1.5 9B Thinking
Ornith 1.5 9B with thinking enabled for more deliberate agentic coding, tool use, visual understanding, and long-context work.
Qwen 3.6 35B A3B Uncensored Thinking
Qwen 3.6 35B A3B Uncensored with thinking enabled for more deliberate coding, multimodal analysis, tool use, and complex chat tasks.
Qwen 3.8 27B Obliterated
Qwen 3.8 27B Obliterated is an open-weight multimodal model LoRA-tuned for fewer refusals across chat, coding, reasoning, tool use, and long-context work.
Qwen 3.8 27B Obliterated Thinking
Qwen 3.8 27B Obliterated with thinking enabled for more deliberate creative work, coding, multimodal analysis, tool use, and long-context problem solving.
Qwen3.8 27B TEE
Qwen3.8 27B is an open-weight dense vision-language model from Alibaba for reasoning, coding, professional workflows, multimodal interaction, tool use, and structured output. Running inside a TEE (Trusted Execution Environment), with provider attestation support.
Gemma 4 31B MeroMero v2
Gemma 4 31B MeroMero v2 is a LoRA finetune for emotive dialogue, relationship scenes, creative writing, and multimodal roleplay.
Ornith 1.5 9B
Ornith 1.5 9B is an FP8 dense open-weight model built for agentic coding, tool use, visual understanding, and efficient long-context work. This variant disables thinking for faster direct responses.
Gemma 4 26B A4B Uncensored
Gemma 4 26B A4B Uncensored is an FP8 open-weight multimodal mixture-of-experts model LoRA-tuned for fewer refusals across chat, coding, tool use, and long-context work.
DeepSeek V4 Flash Vision Exp
An experimental vision-enabled DeepSeek V4 Flash model that adds image understanding while retaining the text, reasoning, coding, tool-calling, and agent capabilities of the base model. This route is served directly by DeepSeek, so privacy and logging guarantees are limited.
Ox Alpha
Ox Alpha is an experimental stealth model with long-context reasoning, vision, tool calling, and structured output support. Prompts and responses are logged and retained by the provider.
Qwen 3.6 35B A3B Uncensored
Qwen 3.6 35B A3B Uncensored is an NVFP4 open-weight mixture-of-experts model LoRA-tuned for fewer refusals across chat, coding, tool use, and multimodal tasks.
Ornith 1.5 35B
Ornith 1.5 35B A3B is an open-weight mixture-of-experts model for agentic coding, tool use, image understanding, and long-context work. This variant disables thinking for faster direct responses.
Ornith 1.5 35B Thinking
Ornith 1.5 35B A3B is an open-weight mixture-of-experts model for agentic coding, reasoning, tool use, image understanding, and long-context work. This variant enables thinking by default.
Qwen3.5 0.8B
Qwen3.5 0.8B is a lightweight open-weight multimodal model from Alibaba for fast reasoning, visual understanding, tool use, and JSON output.