Overview
Voice cloning builds a reusable voice model from a sample of someone speaking, then synthesises new speech in that voice. It is the highest-consent-burden capability in this catalogue: the technology is mature enough that misuse is trivially easy.
How it works
A speaker-encoder extracts a voice embedding capturing timbre and speaking style from the sample. A TTS model conditioned on that embedding generates new speech. Instant cloning uses seconds of audio; high-fidelity cloning uses longer, studio-quality recordings.
Supported AI tools
Support level is recorded per tool, so an integration is never shown as a built-in feature.
Use cases
Narrator scaling
Produce a large audio library in one consistent, licensed voice.
PublishingMultilingual delivery
Speak the same message in several languages in the presenter's own voice.
CorporateVoice restoration
Give someone who has lost their voice a synthetic version of their own.
HealthcareLocalised game dialogue
Extend a cast across languages and late script changes.
GamingBenefits
- Scales one narrator across a whole content library.
- Lets a person keep their voice after illness or injury.
- Produces multilingual delivery in a speaker\'s own voice.
- Fixes a line without recalling the talent to a studio.
Limitations
- Impersonation and fraud risk is severe and well documented.
- Consent and likeness rights are legally live and jurisdiction-dependent.
- Emotional range in long-form output still trails a real performance.
- Many jurisdictions now require disclosure of synthetic voice.
What to look for when choosing a tool
- Documented consent workflow before any clone is created
- Watermarking or provenance signalling on output
- Clear ownership and deletion rights over the voice model
- Misuse detection and takedown process
- Disclosure requirements in every market you publish to