Inside our voice models: timbre, accent, and consent
A convincing cross-language voice preserves who the performer is, not just what they said. That means treating timbre, accent, and consent as three separate things you can actually control.
When people hear "voice cloning" they imagine a single dial. In practice a voice is at least three independent properties, and good localization means keeping some fixed while deliberately changing others.
Three properties of a voice
- Timbre, the fingerprint. The grain and color that make a voice recognizably that person.
- Accent, the geography. Where the speaker sounds like they're from, independent of who they are.
- Prosody, the performance. Rhythm, stress and melody that carry emotion.
Transferring timbre without the accent
The hard trick in dubbing is to keep an actor's timbre while swapping the language, and to do it without dragging their original accent into the new one. We separate the two so a French performer can speak natural Egyptian Arabic and still sound like themselves.
Accent as a dial, not a switch
Accent isn't binary. We can hold a target accent at, say, 80% strength, native enough to feel local, gentle enough to survive a character who is supposed to sound foreign in the story.
Consent, engineered in
None of this ships without permission. Consent isn't a checkbox at the end, it's a gate at every stage, with scope and expiry baked into the asset itself.
- Opt-in with scopeThe performer approves which titles, languages and time window their voice may be used for.
- Watermarked modelsEvery generated line carries an inaudible tag tracing back to its consent record.
- Revocable by designConsent can be withdrawn; the model and its outputs are retired.
The technology is only half the product. The other half is a chain of permission that a performer, and a broadcaster's legal team, can actually trust.






