Speech model · Fast-moving · Intermediate
Text-to-speech models
Neural models that render written text as natural-sounding speech, increasingly in real time.
What Text-to-speech models is
Current systems control prosody, emotion and pacing, support many languages, and can stream audio fast enough for live conversation.
How it works
Text is converted to acoustic features by a neural model and rendered to a waveform by a vocoder, or generated end to end. Voice cloning conditions the model on a speaker embedding.
Why it matters
Voice output quality is now good enough for commercial narration, which changed the economics of audiobooks, courses and IVR.
Common uses
- →Course and audiobook narration
- →Voice agents
- →Accessibility
- →Game dialogue
Strengths
- ✓Fast and inexpensive
- ✓Multilingual
Watch for
- ✓Consent requirements for cloned voices
- ✓Domain pronunciation errors
Continue exploring
More in this collection
Browse all AI Models