Generative AI

Voice Cloning

Reproduce a specific voice from a short sample

Generation Audio Advanced Growing
Capability type
Generation
Modality
Audio
Typical input
Voice sample + text
Typical output
Speech in that voice
Measured by
Speaker similarity / MOS

Overview

Voice cloning builds a reusable voice model from a sample of someone speaking, then synthesises new speech in that voice. It is the highest-consent-burden capability in this catalogue: the technology is mature enough that misuse is trivially easy.

How it works

A speaker-encoder extracts a voice embedding capturing timbre and speaking style from the sample. A TTS model conditioned on that embedding generates new speech. Instant cloning uses seconds of audio; high-fidelity cloning uses longer, studio-quality recordings.

Supported AI tools

Support level is recorded per tool, so an integration is never shown as a built-in feature.

Use cases

Narrator scaling

Produce a large audio library in one consistent, licensed voice.

Publishing

Multilingual delivery

Speak the same message in several languages in the presenter's own voice.

Corporate

Voice restoration

Give someone who has lost their voice a synthetic version of their own.

Healthcare

Localised game dialogue

Extend a cast across languages and late script changes.

Gaming

Benefits

  • Scales one narrator across a whole content library.
  • Lets a person keep their voice after illness or injury.
  • Produces multilingual delivery in a speaker\'s own voice.
  • Fixes a line without recalling the talent to a studio.

Limitations

  • Impersonation and fraud risk is severe and well documented.
  • Consent and likeness rights are legally live and jurisdiction-dependent.
  • Emotional range in long-form output still trails a real performance.
  • Many jurisdictions now require disclosure of synthetic voice.

What to look for when choosing a tool

  • Documented consent workflow before any clone is created
  • Watermarking or provenance signalling on output
  • Clear ownership and deletion rights over the voice model
  • Misuse detection and takedown process
  • Disclosure requirements in every market you publish to

FAQ

Is voice cloning legal?
Cloning your own voice, or someone's with documented consent, is generally lawful. Cloning a person without permission risks liability under likeness, publicity and fraud law, and rules differ sharply by jurisdiction.
How much audio does it need?
Instant cloning works from seconds and gives a recognisable approximation. Studio-grade results need longer, clean, consistently-recorded samples.
Must synthetic voice be disclosed?
Increasingly yes. Several jurisdictions and most major platforms now require disclosure, and responsible providers watermark output by default.