What AI lip sync changes in a video
Traditional dubbing changes only the sound. The picture stays exactly as filmed, and skilled adapters write translations whose mouth shapes roughly fit what the actor's lips are doing. AI lip sync takes the opposite route: it keeps the new audio as it is and changes the picture to fit, regenerating the lower face frame by frame.
That distinction matters. A lip-sync model is a video generation system, a close relative of face-swap and talking-head technology. The output is a new video in which a real person's face has been edited, even if the edit covers only the mouth and jaw.
How visual lip-sync models work
Implementations vary, but most follow a similar sequence.
- Detect and track the face in every frame, estimating landmarks such as the lip corners, jaw line and nose.
- Crop and align the lower face so the model sees a stable region regardless of head movement.
- Convert the new audio into features, typically a mel spectrogram over a short window around each frame, which describe the sounds being made.
- Generate a new mouth region conditioned on those audio features and on reference frames of the same person, so identity, skin tone and lighting are preserved.
- Blend the generated region back into the original frame and smooth it across time to avoid flicker.
Wav2Lip, published in 2020, became a widely cited open research example. Its key idea was training against a separate pretrained lip-sync "expert" network that judges whether mouth movement matches audio, which pushed the generator toward accurate sync. More recent approaches use diffusion models or full talking-head generation for sharper results, at greater computational cost.
Lip sync versus timing sync
There are two very different meanings of "in sync" for a dubbed video, and tools do not always make clear which they offer.
- Timing sync
- The translated speech starts and stops where the original speaker did, so pauses, cuts and gestures line up with the voice. The picture is untouched.
- Lip sync
- The mouth shapes on screen match the sounds in the new audio. Requires either human adaptation of the script or AI modification of the picture.
Timing sync is the baseline every decent dub needs. Lip sync is an additional layer that matters mainly for faces in close-up.
When AI lip sync helps
There are real situations where visual lip sync earns its complexity.
- Close-up presenter videos and advertisements, where the speaker's mouth fills much of the frame.
- Short promotional clips with one on-camera speaker facing the lens, the easiest case for current models.
- Content for audiences used to polished lip-sync dubbing, who may find mismatched mouths distracting.
- Corporate messages where the speaker has explicitly approved a lip-synced translated version of themselves.
Common artifacts and limitations of AI lip sync
Even good models have characteristic failure modes, and viewers notice them quickly on large screens.
- Blur or softness: the regenerated mouth is often less sharp than the rest of the face, especially in high-resolution footage.
- Teeth and tongue: interiors of the mouth are hard to generate, leading to smeared teeth or a dark, featureless gap.
- Jitter and flicker: frame-to-frame inconsistency produces a shimmering mouth.
- Difficult angles: profile views, fast head turns and tilted faces break tracking.
- Occlusions: hands, microphones, beards and glasses across the mouth confuse the model.
- Multiple faces: shots with several people require knowing who is speaking at every moment.
- Identity drift: subtle changes to the person's face that make them look slightly unlike themselves.
- Mismatched expressions: the mouth says one thing while eyes and cheeks still carry the original emotion and rhythm.
Suppose an 8-minute product video has about 6 minutes of screen recording, a minute and a half of the presenter at medium distance, and 30 seconds of close-up talking to camera. Without lip sync, only that 30-second close-up shows an obvious mismatch. Re-editing those moments to cover them with product footage often solves the problem at no risk to the presenter's likeness, whereas lip-syncing the whole video means scrutinizing every on-camera frame for artifacts.
Ethical concerns: consent, deepfakes and disclosure
Lip-sync technology makes a real person appear to say words they never spoke. In a translation context those words are meant to be a faithful rendering, but nothing in the technology enforces that, and the same tools can put fabricated statements in someone's mouth.
- Consent: the person on screen should approve a lip-synced version of themselves, ideally in writing, and know which languages it will appear in.
- Accuracy: an error in translation now looks like the speaker's own words, which raises the stakes on review.
- Disclosure: many platforms and some jurisdictions expect altered or synthetic media to be labeled. Rules change, so check the current policies of the platforms you publish on.
- Trust: audiences who discover undisclosed face manipulation may doubt your other content.
Our article on AI voice ethics covers the related questions for synthetic voices.
Why mydubly leaves the picture untouched
mydubly does not alter the picture. It provides timing sync only: each translated line is placed where the original was spoken, allowed to start up to 0.3 seconds early or end up to 0.6 seconds late, and sped up only slightly when the translation is longer. The new AAC audio track, the translated voice over the kept background, is muxed with your original video stream in your browser without re-encoding it, so every frame is exactly what you filmed.
That design avoids the artifacts and the likeness questions described above, at the cost of visible lip mismatch in close-ups. It works best for screen recordings, tutorials, lectures, narrated explainers and presenters seen at a distance. If your project truly needs lip sync, mydubly can still supply the pieces you would feed into a separate lip-sync tool: the translated audio file, subtitles and transcripts. The AI dubbing page lists what a run produces, and dubbing vs voice-over explains how this replacement-track format compares with traditional ones.
Making a dub without lip sync look natural
Viewers accept unsynced mouths far more readily than most creators expect, especially when the edit helps.
- Identify the close-ups: scrub the video and mark every stretch where the speaker's mouth is large and clearly visible.
- Cover what you can with B-roll, screen recordings, slides or product shots, keeping the translated voice running underneath.
- Prefer medium or wide shots for on-camera segments in future recordings meant for translation.
- Add subtitles in the target language, which draw the eye down from the mouth; mydubly's SRT and VTT files are ready for this.
- Watch the final cut at normal speed on the device your audience uses, not frame by frame in the editor.
Where to go next
Decide first whether your video actually needs lip sync: count the close-up seconds, and you may find the answer is no. For most explainers and tutorials, timing sync plus thoughtful editing is enough. Try that approach on a short clip with AI dubbing; an 8-minute video costs 400 credits ($0.40). For more on preparing footage, read making videos translation-ready.
Frequently asked questions
Can AI lip sync work on any video?
It works best on a single, well-lit face looking roughly toward the camera at decent resolution. Profile shots, fast motion, beards, hands near the mouth and group shots all reduce quality. Many tools let you preview a short section, which is worth doing before committing a whole video.
Does AI lip sync require the speaker's permission?
Legally, requirements depend on where you are and how the video is used, and they are changing. Ethically, altering a real person's face to say new words calls for their informed consent. Get it in writing for anything published, and label the altered version where platforms expect it.
Is lip-synced dubbing more accurate than a regular dub?
No. Lip sync changes how the speaker looks, not what the translation says. Accuracy still depends on recognition, translation and review. If anything, a lip-synced error is more damaging, because it appears to come from the speaker's own mouth.
Why do AI lip-sync mouths sometimes look blurry?
The model regenerates only a small region of the face, often at lower resolution than the source footage, and then blends it back. Teeth and the inside of the mouth are especially hard to reproduce, so detail softens. Higher-quality models reduce this, usually at the cost of much more computation.
How do I get a translated track ready for a separate lip-sync tool?
Run your video with full translation output on and download the translated audio file and subtitles. The audio is already timed to the original lines, which gives a lip-sync tool a sensible starting point. Check that tool's requirements and terms before uploading someone's likeness.