Generative AI

Text-to-Speech

Read written text aloud in a natural voice

Generation Audio Basic Mature
Capability type
Generation
Modality
Audio
Typical input
Written text (+ voice/style)
Typical output
Speech audio
Measured by
MOS (naturalness)

Overview

Text-to-speech (TTS) synthesises audio from text. Neural TTS models produce prosody — rhythm, stress and intonation — that older concatenative systems could not, which is what moved synthetic voice from robotic to broadcast-usable.

How it works

The text is normalised (numbers, dates, abbreviations expanded), converted to phonemes, and passed to an acoustic model that predicts a spectrogram. A vocoder turns that into a waveform. Style tokens or reference audio steer emotion and delivery.

Supported AI tools

Support level is recorded per tool, so an integration is never shown as a built-in feature.

Use cases

E-learning narration

Voice course modules and update them whenever the script changes.

Education

Accessibility playback

Offer an audio version of every article for readers who prefer listening.

Media

IVR and voice agents

Speak dynamic responses in a phone or in-app assistant.

Support

Audiobook production

Narrate long-form text at a fraction of studio cost.

Publishing

Benefits

  • Narration for video and courseware without a booth.
  • Accessibility for readers who prefer or need audio.
  • Consistent brand voice across every touchpoint.
  • Instant updates — change the script, regenerate the audio.

Limitations

  • Long-form delivery can flatten emotionally.
  • Uncommon names and technical terms are mispronounced without a lexicon.
  • Voice cloning raises consent and impersonation issues.
  • Real-time synthesis costs more latency than playing a recording.

What to look for when choosing a tool

  • Voice range and language coverage you need
  • Custom pronunciation lexicon
  • SSML or equivalent control over pace and emphasis
  • Streaming latency for interactive use
  • Commercial licence covering your distribution

FAQ

Does synthetic speech still sound robotic?
Not on modern neural systems for short and medium passages. Very long narration can still flatten in emotional range, which is why audiobook work often mixes synthetic and human takes.
Can I control pronunciation of brand names?
Yes, through a custom lexicon or inline phonetic markup. Any tool aimed at business use should support one or the other.
Do I need permission to use a particular voice?
For stock voices the provider licenses them to you. For a cloned voice you need documented consent from the person — see the Voice Cloning skill for the specifics.