Lip-sync that holds up: the metrics we obsess over
Phoneme alignment, viseme matching, and a handful of perceptual thresholds are what separate a convincing dub from an uncanny one. Here's what we actually measure.
Lip-sync is one of those things audiences only notice when it's wrong, and then they can't un-notice it. Getting it right is less about a single clever model and more about respecting a few stubborn numbers.
Phoneme alignment comes first
Before anything looks right, the sounds have to land in the right place in time. We align the dubbed phonemes to the original mouth movements and score the drift in milliseconds.
Visemes: matching the shapes
Timing isn't enough, the mouth shapes have to plausibly match. Bilabials like b, m and p close the lips; if the dub opens the mouth there, viewers feel the wrongness even if they can't name it.
Timing only
- Words land on time
- But shapes can conflict
- Closed-lip sounds leak
- Subtle "off" feeling
Timing + visemes
- Words land on time
- Lip shapes plausibly match
- Bilabials close correctly
- Reads as natural
The perceptual threshold
Ultimately the only judge is a human eye. We panel-test cuts and track the point where viewers stop reporting a "dubbed" feeling, our real pass mark.
The pipeline that holds it together
None of these numbers is glamorous. But respect all of them at once and you get the thing audiences never comment on, a dub that simply disappears into the story.






