Moyan AI Training Institution LogoMoyan AI

Speech model · Fast-moving · Intermediate

Text-to-speech models

Neural models that render written text as natural-sounding speech, increasingly in real time.

What Text-to-speech models is

Current systems control prosody, emotion and pacing, support many languages, and can stream audio fast enough for live conversation.

How it works

Text is converted to acoustic features by a neural model and rendered to a waveform by a vocoder, or generated end to end. Voice cloning conditions the model on a speaker embedding.

Why it matters

Voice output quality is now good enough for commercial narration, which changed the economics of audiobooks, courses and IVR.

Common uses

  • Course and audiobook narration
  • Voice agents
  • Accessibility
  • Game dialogue

Strengths

  • Fast and inexpensive
  • Multilingual

Watch for

  • Consent requirements for cloned voices
  • Domain pronunciation errors

Continue exploring

More in this collection

Browse all AI Models