Private AI
Browse and discover the best AI language models for conversations, coding, and creative writing.
Muse Spark 1.2
Meta's Muse Spark 1.2 is a multimodal reasoning model for complex agentic and coding tasks, with tool calling, structured output, and a one-million-token context window.
Features
Context
1.0M
Max Output
65.5K
Date Added
Aug 5, 2026
Pricing
Input:
$1.25/1M
Output:
$4.25/1M
Cache:
Read $0.15/1M
Est./msg:
$0.0034
Subscription
Not included in subscription
Muse Spark 1.2 Contributor (Data Used for Training)
A much cheaper opt-in version of Muse Spark 1.2 with the same multimodal coding and agentic capabilities. Prompts and outputs sent to this Contributor model may be used by Meta for training and to improve its products; use the standard Muse Spark 1.2 model if you do not want your data used for training.
Features
Context
1.0M
Max Output
65.5K
Date Added
Aug 5, 2026
Pricing
Input:
$0.10/1M
Output:
$0.20/1M
Cache:
Read $0.0020/1M
Est./msg:
$0.0002
Subscription
Included in subscription
Pokee-Isaac 28B
Pokee-Isaac is a 28B agentic model with a roughly 10-million-token context window, function calling, and OpenAI-compatible structured output. Pokee bills in $0.01 increments, rounding each non-zero request up to the next cent.
Qwen3.8 Max
Qwen3.8 Max is Qwen's 2.4T-parameter flagship model for coding, knowledge work, full-stack development, data analysis, and long-running agent workflows in non-thinking mode. It supports text, image, video, PDF input, tool calling, structured output, and a near-million-token context window.
Qwen3.8 Max Thinking
Qwen3.8 Max Thinking enables generation-time reasoning for deeper coding, knowledge work, data analysis, and long-running agent workflows. It supports text, image, video, PDF input, tool calling, structured output, and a near-million-token context window.
DeepSeek V4 Flash Latest
Compatibility alias that routes to the newest dated DeepSeek V4 Flash release. Currently uses Fireworks first for DeepSeek V4 Flash 0731, with automatic failover when needed. ⚠️ Privacy and logging guarantees are limited.
DeepSeek V4 Flash 0731 (Thinking)
DeepSeek V4 Flash 0731 Thinking enables reasoning by default on the re-post-trained Mixture-of-Experts model with a 1M-token context window, built for coding, reasoning, and agent workflows.
DeepSeek V4 Flash 0731 Cheaper
DeepSeek V4 Flash 0731 Cheaper is the same re-post-trained Mixture-of-Experts model with a 1M-token context window. This route uses Fireworks first, with automatic failover when needed. ⚠️ Privacy and logging guarantees are limited.
DeepSeek V4 Flash 0731 Cheaper (Thinking)
DeepSeek V4 Flash 0731 Cheaper Thinking enables reasoning by default on the same re-post-trained Mixture-of-Experts model with a 1M-token context window. This route uses Fireworks first, with automatic failover when needed. ⚠️ Privacy and logging guarantees are limited.
Gemma 4 12B Instruct
Google's Gemma 4 12B Instruct is an open-weight multimodal model for text, image, audio, and video understanding, with tool calling and structured output support.
Gemma 4 E2B Instruct
Google's Gemma 4 E2B Instruct is a compact open-weight multimodal model with 2B active parameters, supporting text, image, audio, video, tools, and structured output.
Gemma 4 E4B Instruct
Google's Gemma 4 E4B Instruct is an efficient open-weight multimodal model with 4B active parameters, supporting text, image, audio, video, tools, and structured output.
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a re-post-trained Mixture-of-Experts model with a 1M-token context window, built for coding, reasoning, and agent workflows.
Inkling Small
The direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.
Inkling Small Thinking
The reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.
Kimi K3 TEE
Kimi K3 is Moonshot AI's open-weight, always-thinking multimodal model for long-context reasoning, coding, tool use, and native image/video understanding. Running inside a TEE (Trusted Execution Environment), with provider attestation support.
Celeris 1
Celeris 1 is a diffusion language model built for ultra-low-latency classification, extraction, judging, query rewriting, and other short structured responses.
Qwen3.7 Flash
Qwen3.7 Flash is Qwen's fast multimodal model for coding, search and computer-use agents, visual understanding, object recognition, spatial reasoning, and stable end-to-end task execution.