Sarvam Voice Cloning API Verified
Key insights
Concrete technical or product signals.
- Create the voice once. Later synthesis calls send voice_id, text, and language_code. They do not accept the reference audio again.
- Documented output languages include Assamese, Bengali, English (Indian), Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu.
- Voice count is capped by subscription tier. The create-voice page gives starter as an example cap of 3 voices and returns 403 TIER_LIMIT_REACHED when the cap is hit.
Use cases
Where this shines in production.
- Reusable Indic voiceovers
- Cross-lingual speech from one reference clip
- Content Studio voice library voices
Limitations & trade-offs
What to watch for.
- Texts longer than 1,000 characters are rejected. Split them in the client.
- No public per-character price was on the voice-cloning reference pages checked for this entry.
- This is separate from the Dubbing API, which clones speakers inside a dubbing job.
Models referenced
Declared model dependencies or integrations.
Sarvam voice cloning
Related prompts
Hand-picked or latest prompt templates.
Prompt
Prompt-injection resistant assistant
System prompt that treats retrieved docs, tickets, and web pages as untrusted data. Use in multi-tenant RAG and browsing agents. This is a defensive template, not an attack or bypass playbook.
Prompt
Computer-use agent with approval gates
System prompt for browser/desktop agents (Grok Bot, computer-use tools). Allows research and drafts, but pauses before send, purchase, delete, or production changes. Not a jailbreak or bypass guide.
Prompt
Transcript to summary and action items
Turn a speech transcript into decisions, owners, and follow-ups. Use after Gemini 3.5 Transcribe, Whisper, or MAI-Transcribe. Do not use on raw audio; transcribe first.
Prompt
Product image generation prompt
A reusable image prompt for catalog and hero shots: subject, lighting, lens, negative constraints, and text-in-image rules. Works with Firefly, Midjourney, MAI-Image, Qwen-Image, and SDXL. Not for photorealistic people or brand-logo cloning.
Prompt
LLM output evaluation rubric
Judge a model answer against a task, retrieved evidence, and pass/fail thresholds. Use for golden-set grading and regression evals. Not a substitute for human review on high-risk decisions.
Prompt
Strict JSON extraction prompt
Extract structured fields into JSON that matches a schema. Pair with provider structured-output / JSON-schema modes when available. Do not use when the output is meant to be prose.
Looking for a tighter match? Search the prompt library.
Sarvam Voice Cloning API FAQ
What is Sarvam Voice Cloning API?
The Sarvam Voice Cloning API creates a reusable voice from a short reference clip and then synthesizes speech in that voice. POST /voices/create accepts the audio and returns a voice_id immediately. POST /voices/clone speaks text in that voice and is cross-lingual across the documented Indian language codes. Sarvam's September 26, 2026 changelog marks the API as available. Recommended reference audio is 10 to 15 seconds, single speaker, up to 50 MB. Synthesis text is limited to 1,000 characters.
When should teams use Sarvam Voice Cloning API?
Reusable Indic voiceovers
What should teams watch out for with Sarvam Voice Cloning API?
Texts longer than 1,000 characters are rejected. Split them in the client.
Related
Comparisons, platforms, and models teams often view next.