GenAIWiki
Models

TTS

TTS, text-to-speech, synthesizes spoken audio from text, sometimes with voice cloning or style controls.

Expanded definition

Text-to-speech models turn text into waveforms. Production concerns include latency, voice identity, SSML or pronunciation hints, and watermarking. This is the inverse of speech-to-text. Voice cloning and celebrity likeness have consent and legal constraints. Evaluate on names, numbers, and your target language, not only on demo sentences.

Related terms

Explore adjacent ideas in the knowledge graph.

TTS FAQ

What is TTS?

TTS, text-to-speech, synthesizes spoken audio from text, sometimes with voice cloning or style controls.

How is TTS used in AI systems?

Text-to-speech models turn text into waveforms. Production concerns include latency, voice identity, SSML or pronunciation hints, and watermarking. This is the inverse of speech-to-text. Voice cloning and celebrity likeness have consent and legal constraints. Evaluate on names, numbers, and your target language, not only on demo sentences.

Related

Comparisons, tools, and models that connect to this idea.