Sarvam 105B
Sarvam 105B is Sarvam AI's flagship 105B+ parameter Mixture-of-Experts reasoning model for Indian-language and English chat, complex reasoning, coding, long-context document analysis, and agentic tool-use workflows.
Provider
Sarvam AI
Model family
Sarvam
Chat LLM
Cost tier
105b
Status
Current
Release Mar 6, 2026
Why teams choose it
Best evaluated as an India-focused sovereign AI model
not as a universal replacement for every global frontier model.
Sarvam reports especially strong Indian-language and agentic benchmark results for its class.
Sarvam reports especially strong Indian-language and agentic benchmark results for its class.
The model exposes an OpenAI-compatible chat completions API
which lowers integration friction for teams already using OpenAI-style clients.
Tradeoffs to know
- Published benchmark results are primarily vendor-reported and should be validated with an independent task-specific eval.
- Thinking mode is on by default in Sarvam docs, so small max_tokens settings can be consumed by reasoning tokens.
- Global model comparison should account for language mix, latency region, context length, and deployment constraints.
When not to use this
- Self-hosting outcomes depend on hardware, quantization, and ops maturity—budget time beyond swapping an API hostname.
- May demand more instrumentation than SaaS-managed APIs to duplicate latency, failover, and support guarantees.
- Benchmark prompts and regressions continuously before rewriting entire routing tables around weights.
Technical specs
- Inputs
- text
- Outputs
- text
- Capabilities
- Indian-language chat, Reasoning, Coding, Long-context document analysis, Tool use, Agentic workflows, OpenAI-compatible chat completions
- License
- Apache 2.0
- Model string
sarvam-105b
Benchmarks
{
"source": "https://www.sarvam.ai/blogs/sarvam-30b-105b",
"math500": 98.6,
"mmlu_pro": 81.7,
"tau2_avg": 68.3,
"aime_2025": 88.3,
"browsecomp": 49.5,
"gpqa_diamond": 78.7,
"vendor_reported": true,
"live_code_bench_v6": 71.7,
"swe_bench_verified": 45,
"aime_2025_with_tools": 96.7,
"indian_language_win_rate_avg": "90%"
}Sarvam 105B model IDs and conversational variant
Use sarvam-105b for complex reasoning, coding, long-context analysis, and agentic tool use. Sarvam also documents sarvam-105b-conversations, a post-trained variant for real-time dialogue, voice agents, and customer-facing chat.
- Both variants use the OpenAI-compatible chat-completions request shape and share a 128K context window.
- The conversations variant is available through the v1 chat-completions endpoint; verify endpoint support before changing model IDs.
- Choose with representative native-script, romanized, and code-mixed conversations rather than English-only tests.
Sources: Sarvam 105B model documentation
Sarvam 105B deployment and evaluation
Sarvam documents the flagship model as a 105B+ Mixture-of-Experts model with 128 sparse experts, Multi-head Latent Attention, a 128K context window, streaming, and an Apache 2.0 license.
- Test all supported scripts and code-mixing patterns that appear in the product.
- Keep vendor-reported benchmark results separate from independent application evaluations.
- Compare Sarvam 30B when lower latency or lower cost matters more than maximum reasoning quality.
Sources: Sarvam 105B specifications, Building for Indian languages
Sarvam family lineup
Current models
Compare with
Sarvam 105B FAQ
What is Sarvam 105B?
Sarvam 105B is Sarvam AI's flagship 105B+ parameter Mixture-of-Experts reasoning model for Indian-language and English chat, complex reasoning, coding, long-context document analysis, and agentic tool-use workflows. Sarvam documents it as a 128K-context OpenAI-compatible chat model with Multi-head Latent Attention, 12...
When does Sarvam 105B fit best?
Indian-language enterprise assistants with native-script, romanized, and code-mixed input.
What should teams watch out for with Sarvam 105B?
Published benchmark results are primarily vendor-reported and should be validated with an independent task-specific eval.
Explore next
Models, tools, and comparisons that connect to this reference.