Provider logo

Sarvam 105B

sarvam-105b
Provider logo

Sarvam 105B

sarvam-105b

Sarvam 105B is a 105B parameter chat completion model from Sarvam AI with multilingual support, streaming, tool calling, reasoning controls, and a 128k context window.

Added May 12, 2026

Context Window

131.1K

Max Output

4.1K

Avg output tokens (7d)

868 tokens

64%

Input Price (Auto)

$0.054/1M

Output Price (Auto)

$0.21/1M

Cache Read (Auto)

$0.034/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

Sourced from Artificial Analysis.

Intelligence Index

8.8

Better than 36% of models compared

Agentic work

T²-Bench Telecom (legacy)

Legacy fallback · Conversational AI agents in dual-control scenarios

46.8%

Better than 51% of models compared

Document reasoning

AA-LCR v1.1

Long context reasoning with updated grading

0.0%

Better than 6% of models compared

Reasoning

HLE

Humanity's Last Exam

11.0%

Better than 57% of models compared

IFBench

Instruction-following benchmark

34.4%

Better than 25% of models compared

Coding

Terminal-Bench Hard (legacy)

Legacy fallback · Agentic coding and terminal use

1.5%

Better than 16% of models compared

Legacy benchmarks

GPQA Diamond (legacy)

Graduate-level scientific reasoning

73.8%

Better than 55% of models compared

AA-LCR (unversioned / legacy)

Long context reasoning evaluation

0.0%

Better than 6% of models compared

Last updated Sep 13, 2026

Artificial Analysis

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare Sarvam 105B with similar models from the same provider or model family.

Schematron V2 Small

inference-net/schematron-v2-small

Inference.net's 3B-parameter HTML-to-JSON extraction model, focused on accuracy for complex schemas and long web pages. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.

Schematron V2 Turbo

inference-net/schematron-v2-turbo

Inference.net's 3B-parameter HTML-to-JSON extraction model, optimized for throughput and low cost on high-volume workloads. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.

DeepSeek V4.1 Flash TEE

TEE/deepseek-v4.1-flash

DeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.

GPT Astra Latest

openai/gpt-astra-latest

Compatibility alias that routes to GPT 6 Astra, the latest supported GPT Astra model.

GPT Luna Latest

openai/gpt-luna-latest

Compatibility alias that routes to GPT 5.6 Luna, the latest supported GPT Luna model.

GPT Sol Latest

openai/gpt-sol-latest

Compatibility alias that routes to GPT 5.6 Sol, the latest supported GPT Sol model.