Skip to main content
Benchmark method

How the lip sync score is measured

1,000 videos, identical source material, every vendor — and real viewers decide blind who wins.

videos in a fixed, versioned test corpus

1,000

independent raters per video

5

judgments per test round

5,000
The setup

Only the lip sync model differs

Every video is sent through every vendor with exactly the same input. Everything before and after the lip sync step stays constant, so the result can differ in one place only.

What the corpus covers

  • Different languages

  • Different resolutions

  • Different face sizes in frame

  • Different lighting

  • Different recording quality

Held constant

  • Source video

  • Translated audio track

  • Voice

  • Timing

The one variable

The lip sync model — one per vendor

The rating

One question, asked 5,000 times

For each video, all results are shown side by side — without vendor names, in random order. Five independent people rate each video.

  • D
  • A
  • E
  • B
  • C
Which video looks best?No visible difference

If two results really cannot be told apart, the rater may choose “no visible difference”. Both then receive the point.

  • 5 raters per video

  • 5,000 judgments per round

The score

What a score of 78 means

The score is the percentage of all judgments in which a vendor was chosen as the best.

The scale

  • 100

    Chosen best in every one of the 5,000 judgments

  • 80

    Ahead in 4 of 5 judgments

  • 50

    Ahead in half of all judgments

  • 20

    Chance level with five vendors — no measurable advantage

  • 0

    Never chosen best

Worked example

Vendor A is named best in 3,900 of 5,000 judgments. Score: 78. Read it directly: 78% of viewers consider this result the best in the field.

Scores do not add up to 100

Every vendor is measured on its own. If two vendors are consistently equally good, both can reach 90.

Accurate to about ±3 points

At 5,000 judgments the statistical uncertainty is around three points. Gaps of less than three points are reported as a tie, not as a ranking.

Reading the results

Why the field sits close together

A low score is possible in this procedure: with five vendors, chance alone gives 20. That every vendor tested lands well above it is a result in itself — the market has tightened. The differences that matter no longer show in the average but in the hard cases, which is why two further readouts are reported.

Tie rate

The share of videos in which the panel saw no difference at all.

Score by difficulty

Reported separately for easy recordings — frontal, studio — and hard ones: profile, movement, poor light, a small face in the frame.

Fairness and limits

What the test can say, and what it cannot

  • Every result is rated blind. The raters know neither the vendor names nor who commissioned the test.

  • Named model versions are tested at a named date. Lip sync models change quickly, and a score is a snapshot of that date.

  • The test is commissioned by a market participant and run by an external agency. Precisely because of that, the test corpus, the rating protocol and the raw data are documented and available from our team on request.

  • The results hold for this corpus. Other material can produce a different order.