Pareto is an experimental multimodal composite model from Unbiased AI for research, coding, agentic workflows, and general-purpose tasks. It runs several frontier and open models for each request and returns a single selected or synthesized answer. This experimental release supports text and image input, tool calling, a 262K-token context window, and up to 131K output tokens. Its underlying model mix and behavior may change as it is evaluated. Request content may be retained for abuse prevention and security monitoring for up to 30 days, and is not used for model training without explicit consent.
Added Sep 18, 2026
Context Window
262.1K
Max Output
131.1K
Input Price (Auto)
$2.50/1M
Output Price (Auto)
$7.50/1M
Cache Read (Auto)
$0.25/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
No benchmark data is available yet for this model.
Providers
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Pareto with similar models from the same provider or model family.
Qwen3.8 Omni Flash
qwen/qwen3.8-omni-flashQwen3.8 Omni Flash is a fast omni-modal reasoning model for understanding text, images, audio, and video. It is especially suited to meeting summaries, transcripts and subtitles, speaker-aware audiovisual analysis, and long-form content review. It returns text and supports a nearly one-million-token context window, tool calling, and structured output.
Gemma 4 31B Split-Untied
slowburn/gemma4-31b-splituntiedSlowburn's Split-Untied is a text-only Gemma 4 31B community finetune with an untied BF16 output head, built for creative writing, roleplay, expressive dialogue, and tool use.
Schematron V2 Small
inference-net/schematron-v2-smallInference.net's 3B-parameter HTML-to-JSON extraction model, focused on accuracy for complex schemas and long web pages. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.
Schematron V2 Turbo
inference-net/schematron-v2-turboInference.net's 3B-parameter HTML-to-JSON extraction model, optimized for throughput and low cost on high-volume workloads. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.
DeepSeek V4.1 Flash TEE
TEE/deepseek-v4.1-flashDeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.
GPT Astra Latest
openai/gpt-astra-latestCompatibility alias that routes to GPT 6 Astra, the latest supported GPT Astra model.