Sarvam 105B is a 105B parameter chat completion model from Sarvam AI with multilingual support, streaming, tool calling, reasoning controls, and a 128k context window.
Added May 12, 2026
Context Window
131.1K
Max Output
4.1K
Input Price (Auto)
$0.045/1M
Output Price (Auto)
$0.18/1M
Cache Read (Auto)
$0.028/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
11.9
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
73.8%
Better than 57% of models compared
HLE
Humanity's Last Exam
11.0%
Better than 60% of models compared
IFBench
Instruction-following benchmark
34.4%
Better than 25% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
46.8%
Better than 51% of models compared
AA-LCR
Long context reasoning evaluation
0.0%
Better than 6% of models compared
Coding
SciCode
Python programming for scientific computing
26.4%
Better than 31% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
1.5%
Better than 16% of models compared
Last updated Aug 16, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Sarvam 105B with similar models from the same provider or model family.
Sarvam 30B
sarvam-30bSarvam 30B is a 30B parameter chat completion model from Sarvam AI with multilingual support, streaming, tool calling, reasoning controls, and a 64k context window.
Gemma 4 31B MeroMero v2
Gemma-4-31B-MeroMero-v2Gemma 4 31B MeroMero v2 is a LoRA finetune for emotive dialogue, relationship scenes, creative writing, and multimodal roleplay.
Ornith 1.5 9B
ornith-ai/ornith-1.5-9bOrnith 1.5 9B is an FP8 dense open-weight reasoning model built for agentic coding, tool use, visual understanding, and efficient long-context work.
Gemma 4 26B A4B Uncensored
google/gemma-4-26b-a4b-uncensoredGemma 4 26B A4B Uncensored is an FP8 open-weight multimodal mixture-of-experts model LoRA-tuned for fewer refusals across chat, coding, tool use, and long-context work.
DeepSeek V4 Flash Vision Exp
deepseek/deepseek-v4-flash-vision-expAn experimental vision-enabled DeepSeek V4 Flash model that adds image understanding while retaining the text, reasoning, coding, tool-calling, and agent capabilities of the base model. This route is served directly by DeepSeek, so privacy and logging guarantees are limited.
Qwen 3.6 35B A3B Uncensored
qwen/qwen3.6-35b-a3b-uncensoredQwen 3.6 35B A3B Uncensored is an NVFP4 open-weight mixture-of-experts model LoRA-tuned for fewer refusals across chat, coding, tool use, and multimodal tasks.