Qwen25 VL 72b model with 32k context window
Added May 10, 2025
Context Window
32.0K
Max Output
32.8K
Input Price (Auto)
$0.70/1M
Output Price (Auto)
$0.70/1M
Cache Read (Auto)
$0.35/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
No benchmark data is available yet for this model.
Providers
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Qwen25 VL 72b with similar models from the same provider or model family.
Qwen2.5 72B
qwen/qwen-2.5-72b-instructGreat multilingual support, strong at mathematics and coding, supports roleplay and chatbots.
Qwen2.5 VL 72B TEE
TEE/qwen2.5-vl-72b-instructQwen2.5 Vision-Language 72B model with multimodal capabilities. Running inside a TEE (Trusted Execution Environment), with provider attestation support.
Qwen3 VL 235B A22B Instruct Original
qwen3-vl-235b-a22b-instruct-originalNote: direct via Alibaba, a Chinese entity - privacy and logging guarantees may be limited. Qwen3 Vision‑Language model (235B MoE, ≈22B active) tuned for instruction following and grounded visual QA. Excels at image understanding, dense OCR, charts and diagrams, and multi‑image context. Use this variant when you want concise, direct answers grounded in the visuals.
Qwen3 Next 80B A3B (Instruct)
Qwen/Qwen3-Next-80B-A3B-InstructBased on the new Qwen3‑Next architecture (hybrid attention, highly sparse MoE, training‑stability optimizations, and multi‑token prediction), the Qwen3‑Next‑80B‑A3B‑Instruct model delivers extreme efficiency with only 3B active parameters per pass. It performs comparably to Qwen3‑235B‑A22B‑Instruct‑2507 and shows clear advantages on ultra‑long context tasks (up to 256K tokens).
Qwen3 Coder 30B A3B Instruct
qwen3-coder-30b-a3b-instructQwen3 Coder 30B with 3B active parameters, optimized for code generation and technical tasks
Qwen 3 235b A22B 2507
Qwen/Qwen3-235B-A22B-Instruct-2507Qwen 3 235b A22B Instruct 2507 the updated version of Qwen3 235B A22B, with significant improvements in performance. This model is non-thinking.