The Disclosure Rule Arrived Before the Quality Did
YouTube requires creators to disclose meaningfully altered or synthetic content at upload — and the policy explicitly names synthetic voice that makes it sound like a real person said something they did not. TikTok has the same obligation with its own AI-generated content label. Both rules landed while cloned voices were still obviously synthetic — and YouTube has kept expanding the surrounding likeness-detection tooling since, as tracked on the YouTube official blog. They are now not obviously synthetic, and the compliance question got harder, not easier.
We have cloned voices for client work across 10,000+ edited projects, 200M+ views and $10M+ in generated client revenue, and the single most useful thing we have learned is a boundary, not a technique: a cloned voice is a delivery tool, not a performance. It can read. It cannot act. Every failed synthetic narration we have shipped or rejected failed on that line, and almost none of them failed on audio fidelity.
This is a different problem from AI dubbing. Dubbing translates a performance that already exists — the timing, the emphasis and the emotional arc were performed by a human and the model reproduces them in another language. Synthetic narration has no performance underneath it. It is generating one from a script.
Where a Cloned Voice Actually Holds Up

The pattern across our roster is consistent enough to state as a rule. Cloned narration survives contact with an audience when the content is informational, densely written, and paced evenly. It falls apart when the content depends on timing.
| Content type | Cloned voice verdict | Why |
|---|---|---|
| Explainer / documentary VO | Works well | Even pacing, no comedic timing, script carries the value |
| Product walkthroughs, tutorials | Works well | Viewers are watching the screen, not listening for personality |
| Multi-language versions of existing videos | Works well | Reference performance exists to clone against |
| Pickup lines and script corrections | Works well | Seamless if the clone is trained on the same session |
| Comedy and reaction | Fails | Timing and surprise cannot be generated from text |
| Personal storytelling | Fails | The catch in a real voice is the whole point |
| Anything delivered as "you", on camera | Fails | Audience detects the mismatch even when they cannot name it |
The fourth row is the one most creators overlook and the one that pays for itself immediately. If you record a 20-minute video and find one wrong number or one mangled sentence in the edit, a clone trained on that session's audio can generate a four-second replacement that sits inside the original recording undetectably. We do this weekly. It is the least glamorous use of the technology and by a wide margin the most valuable.
Garbage In Is Still Garbage Out
The quality ceiling of a clone is set by the training audio, not by the model. Every provider — ElevenLabs, Descript, the platform-native tools — will happily build a voice from a noisy phone recording and hand you something that sounds like a compressed phone recording forever.

What we require before we will train a clone for a client:
- Signed written consent from the voice owner, naming the projects the clone may be used on. Not a Slack "yeah go ahead."
- 20–30 minutes of clean speech from a single session, one microphone, one room, one distance.
- No processing on the training audio. No compression, no de-noise, no EQ. The model will learn the artefacts as part of the voice.
- Varied delivery in the source — statements, questions, lists, at least one passage read fast and one slow. A monotone training set produces a monotone clone.
- A held-out reference clip the clone was not trained on, kept for A/B checking every future render.
Steps 2 and 3 are just a proper voice recording chain applied once, deliberately. If your normal capture is not good enough to clone from, it was never good enough to publish either.
Direct the Clone Like a Voice Actor
The mistake is pasting a script and accepting take one. Synthetic narration is directable, and the direction happens in the script rather than in the booth.
- Punctuate for breath, not for grammar. A comma is a short pause and a full stop is a long one. Split long sentences the way you would want them read, even where a copy editor would object.
- Spell out anything ambiguous. Write "twenty twenty-six" not "2026", "dot com" not ".com", "A P I" not "API" if you want it spelled aloud. Numbers and acronyms are where clones announce themselves.
- Generate in paragraph-length chunks, not whole scripts. Prosody drifts over long generations, and a bad phrase means regenerating three minutes instead of fifteen seconds.
- Regenerate rather than fix in post. Two or three seeds on the same line is seconds of work; pitch-editing a wrong reading never sounds right.
- Leave real room tone under it. Perfect digital silence between phrases is the single loudest tell. Our editors bed synthetic VO on a light noise floor and treat it exactly like a recorded track in the sound design pass.
Consent, Disclosure and the Part That Ends Channels
The legal exposure here is not copyright, which is where creators are used to looking. It is the right of publicity — a person's control over their own voice and likeness — and it does not require registration to exist, which makes it very different from the copyright and Content ID system most creators already navigate.

Practically, three rules cover almost everything:
Never clone a voice you do not have written permission to clone. Not a celebrity, not a politician, not a dead public figure, not your co-host on a handshake. "It's obviously a parody" is a defence you get to make after you have already been removed and demonetised.
Disclose in layers. Set the platform's altered-content flag at upload; add a line in the description; and if the entire narration is synthetic, say so once in the video. Layered disclosure is unglamorous and it is the thing that keeps a channel intact when a policy tightens retroactively.
Do not disclose your way out of a deception. Labelling does not make a synthetic endorsement acceptable. A cloned voice recommending a product the human never agreed to recommend is a problem no label solves.
The Pre-Flight Gate Before You Publish Synthetic Narration
- Written consent on file from the voice owner, scoped to this project.
- Platform synthetic-content disclosure set at upload.
- Description line stating narration is AI-generated, if it is.
- A/B checked against the held-out reference clip — not just against your memory of the voice.
- Every number, acronym and proper noun listened to individually.
- Room tone under the VO; no digital silence gaps.
- Captions written from the script, not auto-generated, so the accessibility layer matches what was actually said.
If you cannot tick line 1, nothing else on the list matters.
The Bottom Line
A cloned voice is a production tool for scale and repair, not a substitute for a performance — it reads brilliantly and it cannot act. Use it for explainers, localisation, high-volume faceless formats and invisible pickup lines; keep a human on anything where timing, humour or emotional truth is the product. And disclose in layers, every time, because the policies were written before the quality arrived and they will only get stricter.


