GenAIWiki
Voice and speech

Sarvam Voice Cloning API Verified

The Sarvam Voice Cloning API creates a reusable voice from a short reference clip and then synthesizes speech in that voice. POST /voices/create accepts the audio and returns a voice_id immediately. POST /voices/clone speaks text in that voice and is cross-lingual across the documented Indian language codes. Sarvam's September 26, 2026 changelog marks the API as available. Recommended reference audio is 10 to 15 seconds, single speaker, up to 50 MB. Synthesis text is limited to 1,000 characters.
API availableSarvam API subscription. The voice-cloning reference pages checked on September 30, 2026 do not publish a per-character rate.sarvamvoicecloningttsindic
FeaturedUpdated 4 days agoLast verified: September 2026

Key insights

Concrete technical or product signals.

  • Create the voice once. Later synthesis calls send voice_id, text, and language_code. They do not accept the reference audio again.
  • Documented output languages include Assamese, Bengali, English (Indian), Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu.
  • Voice count is capped by subscription tier. The create-voice page gives starter as an example cap of 3 voices and returns 403 TIER_LIMIT_REACHED when the cap is hit.

Use cases

Where this shines in production.

  • Reusable Indic voiceovers
  • Cross-lingual speech from one reference clip
  • Content Studio voice library voices

Limitations & trade-offs

What to watch for.

  • Texts longer than 1,000 characters are rejected. Split them in the client.
  • No public per-character price was on the voice-cloning reference pages checked for this entry.
  • This is separate from the Dubbing API, which clones speakers inside a dubbing job.

Models referenced

Declared model dependencies or integrations.

Sarvam voice cloning

Related prompts

Hand-picked or latest prompt templates.

Looking for a tighter match? Search the prompt library.

Sarvam Voice Cloning API FAQ

What is Sarvam Voice Cloning API?

The Sarvam Voice Cloning API creates a reusable voice from a short reference clip and then synthesizes speech in that voice. POST /voices/create accepts the audio and returns a voice_id immediately. POST /voices/clone speaks text in that voice and is cross-lingual across the documented Indian language codes. Sarvam's September 26, 2026 changelog marks the API as available. Recommended reference audio is 10 to 15 seconds, single speaker, up to 50 MB. Synthesis text is limited to 1,000 characters.

When should teams use Sarvam Voice Cloning API?

Reusable Indic voiceovers

What should teams watch out for with Sarvam Voice Cloning API?

Texts longer than 1,000 characters are rejected. Split them in the client.

Related

Comparisons, platforms, and models teams often view next.

This page is based on publicly available documentation, benchmarks, and real-world usage patterns. Last reviewed for accuracy recently.