How the lip sync score is measured
1,000 videos, identical source material, every vendor — and real viewers decide blind who wins.
videos in a fixed, versioned test corpus
- 1,000
independent raters per video
- 5
judgments per test round
- 5,000
Only the lip sync model differs
Every video is sent through every vendor with exactly the same input. Everything before and after the lip sync step stays constant, so the result can differ in one place only.
What the corpus covers
Different languages
Different resolutions
Different face sizes in frame
Different lighting
Different recording quality
Held constant
Source video
Translated audio track
Voice
Timing
The one variable
The lip sync model — one per vendor
One question, asked 5,000 times
For each video, all results are shown side by side — without vendor names, in random order. Five independent people rate each video.
- D
- A
- E
- B
- C
If two results really cannot be told apart, the rater may choose “no visible difference”. Both then receive the point.
5 raters per video
5,000 judgments per round
What a score of 78 means
The score is the percentage of all judgments in which a vendor was chosen as the best.
The scale
- 100
Chosen best in every one of the 5,000 judgments
- 80
Ahead in 4 of 5 judgments
- 50
Ahead in half of all judgments
- 20
Chance level with five vendors — no measurable advantage
- 0
Never chosen best
Worked example
Vendor A is named best in 3,900 of 5,000 judgments. Score: 78. Read it directly: 78% of viewers consider this result the best in the field.
Scores do not add up to 100
Every vendor is measured on its own. If two vendors are consistently equally good, both can reach 90.
Accurate to about ±3 points
At 5,000 judgments the statistical uncertainty is around three points. Gaps of less than three points are reported as a tie, not as a ranking.
Why the field sits close together
A low score is possible in this procedure: with five vendors, chance alone gives 20. That every vendor tested lands well above it is a result in itself — the market has tightened. The differences that matter no longer show in the average but in the hard cases, which is why two further readouts are reported.
Tie rate
The share of videos in which the panel saw no difference at all.
Score by difficulty
Reported separately for easy recordings — frontal, studio — and hard ones: profile, movement, poor light, a small face in the frame.
What the test can say, and what it cannot
Every result is rated blind. The raters know neither the vendor names nor who commissioned the test.
Named model versions are tested at a named date. Lip sync models change quickly, and a score is a snapshot of that date.
The test is commissioned by a market participant and run by an external agency. Precisely because of that, the test corpus, the rating protocol and the raw data are documented and available from our team on request.
The results hold for this corpus. Other material can produce a different order.
Instant results