Overview
Text-to-speech (TTS) synthesises audio from text. Neural TTS models produce prosody — rhythm, stress and intonation — that older concatenative systems could not, which is what moved synthetic voice from robotic to broadcast-usable.
How it works
The text is normalised (numbers, dates, abbreviations expanded), converted to phonemes, and passed to an acoustic model that predicts a spectrogram. A vocoder turns that into a waveform. Style tokens or reference audio steer emotion and delivery.
Supported AI tools
Support level is recorded per tool, so an integration is never shown as a built-in feature.
Use cases
E-learning narration
Voice course modules and update them whenever the script changes.
EducationAccessibility playback
Offer an audio version of every article for readers who prefer listening.
MediaIVR and voice agents
Speak dynamic responses in a phone or in-app assistant.
SupportAudiobook production
Narrate long-form text at a fraction of studio cost.
PublishingBenefits
- Narration for video and courseware without a booth.
- Accessibility for readers who prefer or need audio.
- Consistent brand voice across every touchpoint.
- Instant updates — change the script, regenerate the audio.
Limitations
- Long-form delivery can flatten emotionally.
- Uncommon names and technical terms are mispronounced without a lexicon.
- Voice cloning raises consent and impersonation issues.
- Real-time synthesis costs more latency than playing a recording.
What to look for when choosing a tool
- Voice range and language coverage you need
- Custom pronunciation lexicon
- SSML or equivalent control over pace and emphasis
- Streaming latency for interactive use
- Commercial licence covering your distribution