Browse all Microsoft text models
Provider logo

WizardLM-2 8x22B

microsoft/wizardlm-2-8x22b
Provider logo

WizardLM-2 8x22B

microsoft/wizardlm-2-8x22b

Microsoft's advanced Wizard model. The most popular role-playing model.

Added Apr 15, 2025

Model weights

Context Window

65.5K

Max Output

8.2K

Avg output tokens (7d)

525 tokens

47%

Input Price (Auto)

$0.49/1M

Output Price (Auto)

$0.49/1M

Cache Read (Auto)

$0.25/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

No benchmark data is available yet for this model.

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare WizardLM-2 8x22B with similar models from the same provider or model family.

Phi 4 Mini

phi-4-mini-instruct

Phi 4 Mini by Microsoft. A small multilingual model.

Phi 4 Multimodal

phi-4-multimodal-instruct

Phi 4 by Microsoft. A small multimodal model that can handle images and text.

Mixtral 8x22B

mistralai/mixtral-8x22b-instruct-v0.1

Mixtral 8x22B is a powerful sparse Mixture of Experts (MoE) model with 141B total parameters and 39B active per token. Features a 64K context window, exceptional math performance, and cost-efficient inference. Supports English, French, German, Spanish, and Italian. Apache 2.0 licensed.

Schematron V2 Small

inference-net/schematron-v2-small

Inference.net's 3B-parameter HTML-to-JSON extraction model, focused on accuracy for complex schemas and long web pages. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.

Schematron V2 Turbo

inference-net/schematron-v2-turbo

Inference.net's 3B-parameter HTML-to-JSON extraction model, optimized for throughput and low cost on high-volume workloads. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.

DeepSeek V4.1 Flash TEE

TEE/deepseek-v4.1-flash

DeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.