Skip to main content

Lip-sync: The basics

Learn what Lip-Sync does, when to use it for on-camera speakers, and the resolution, face size, and footage requirements for the best results.

Lip-Sync is Dubly.AI's feature that re-animates the speaker's mouth to match the translated audio. The result: a dubbed video that looks like it was filmed in the target language.

Dub page with the Show lip-synced version toggle on the player

What Lip-Sync syncs to

Lip-Sync always matches the mouth to the dubbed audio Dubly generated for that subdub. Two things follow from that:

  • You can't supply your own audio in the app. The standard workflow has no upload for a voice-over or dubbing track (MP3, WAV, M4A) that Lip-Sync would match the video to. If your project comes with an existing audio track, ask Dubby in the chat: for individual projects our team checks on request what we can make possible for you.

  • There is no Lip-Sync-only mode. Lip-Sync is a step inside a dub and runs only after dubbing has finished, so a translation into a target language always happens first.

What this means in practice — and why merging your own audio into the video beforehand does not help either — is explained in Lip-sync limitations and workarounds.

When to use Lip-Sync

Lip-Sync is recommended when:

  • The speaker is visible on camera (interview, selfie video, presenter shot)

  • The mouth is clearly visible for most of the shot

  • You want the highest-quality, "native-feeling" result

  • The video contains at least 10 seconds of clear, unobstructed footage of the speaker's face for the AI to process the lip-sync accurately.

Lip-Sync is not needed for voice-over style content where the speaker is off-screen (screencasts, product tours, animations).

Related articles

Did this answer your question?