Video engineering & media pipelines

How Players Keep Sound and Picture Together: Master Clocks, Timestamps and Frame Drops

Players keep sound and picture together by choosing one master clock, almost always the audio, and then showing each video frame when its timestamp matches that clock. If the picture falls behind, frames are dropped; if it runs ahead, a frame stays on screen longer. Sync therefore depends on two things: correct timestamps in the file and a player that measures its audio clock honestly, which is why a replacement soundtrack has to start on the same timeline as the original.

8 min read · Updated

Why sync needs managing at all

In a perfect world, a player would start audio and video at the same instant and both would run at exactly the rate the file says. In practice, three clocks disagree. The sound card plays samples at a rate set by its own crystal, which may run very slightly fast or slow. The display refreshes at its own rate, commonly 60 times a second, regardless of the video's frame rate. And decoding takes a variable amount of time: one frame may decode in a few milliseconds, the next much longer.

Over a two-hour film, even a tiny difference between the sound card's rate and the computer's system clock adds up. If audio and video were started together and left to run independently, they would drift. A player has to keep correcting, and the correction strategy is built around a single reference.

The master clock, usually audio

A player picks one clock as the master and slaves everything else to it. The usual choice is the audio clock: the playback position of the sound actually reaching the speakers, worked out from how many samples have been handed to the audio device minus those still waiting in its buffer.

Audio is chosen because the ear is unforgiving. Skipping or repeating a few milliseconds of sound produces an audible click or stutter, and stretching it changes pitch unless done carefully. Showing a frame slightly longer, or skipping one, is far less noticeable. So the picture adapts and the sound flows uninterrupted.

Audio master
The default in most players. Video frames are timed against the sound card's playback position
Video master
Audio is resampled or adjusted to follow the video. Used in some editing and broadcast contexts
External clock
Both follow a system or network clock. Used for live streams and when there is no audio track
ffplay example
ffmpeg's simple player exposes all three through its -sync option, which makes it a handy way to experiment

When a file has no audio, or the audio ends before the video, players fall back on the system clock for the remaining picture.

Timestamps tell the player where everything belongs

The master clock answers "what time is it now?" The file's timestamps answer "when should this frame or sound appear?" Every decoded frame carries a presentation timestamp, and every audio packet does too. The player's job is to compare them.

Those timestamps are not always zero-based and are not always in the same unit for each stream; the mechanics of presentation and decode timestamps, timebases and edit lists are explained in PTS and DTS. For sync, what matters is that audio and video timestamps sit on one shared timeline. An audio packet stamped 10.000 s and a frame stamped 10.000 s are meant to be experienced together.

Dropping and repeating frames

With the audio clock as the reference, the player runs a simple loop for each decoded frame:

  1. Read the current audio clock.
  2. Compare it with the frame's presentation timestamp.
  3. If the frame is early, wait until its time comes, then display it at the next screen refresh.
  4. If it is slightly late, display it immediately.
  5. If it is badly late, discard it and move to the next frame so the picture catches up.

The opposite case, video running ahead, resolves itself: the current frame simply stays on screen until the clock reaches the next frame's time. That is effectively a repeated frame.

There is a second layer of repetition that has nothing to do with drift. A 24 fps film on a 60 Hz screen cannot show each frame for an equal number of refreshes, because 60 is not a multiple of 24. Players hold frames for alternately three and two refreshes, a pattern called 3:2 cadence, which produces a slight unevenness in slow pans known as judder. Displays that can switch to a matching refresh rate avoid it.

Example: a busy laptop

A hypothetical 30 fps lecture plays on a laptop that is also exporting a project. For a moment, decoding falls 80 ms behind the audio. The player drops two frames, the picture jumps forward slightly, and the voice never stutters. A viewer sees a tiny hitch in motion but no lasting offset between lips and words.

Output latency: the delay after the player

The audio clock is only honest if the player knows how long sound takes to reach the listener. Wired speakers add almost nothing. Bluetooth headphones add a noticeable delay because audio is buffered, compressed and transmitted. TVs add video latency of their own while processing the picture, and sound bars may add audio latency.

Operating systems report output latency to players, and good players subtract it from the audio clock so the picture is timed against what you actually hear. When that reporting is wrong or missing, sync is off for every file on that device even though the files are fine. Diagnosing that and other causes, step by step, is the subject of audio and video out of sync.

How much offset viewers notice

There is no single number for perceptible sync error, and anyone quoting one precisely is simplifying. A few points are widely agreed:

  • Errors of a few milliseconds are invisible. Offsets in the tens of milliseconds can become noticeable on close-up talking faces, and larger offsets are distracting to most viewers.
  • People tolerate sound that lags the picture better than sound that leads it, plausibly because in daily life light arrives before sound from any distant source.
  • Broadcast recommendations such as ITU-R BT.1359 set separate limits for sound leading and sound lagging, with the tighter limit for sound arriving early. Check the current text of the relevant standard for exact figures.
  • Tolerance depends on content. A wide shot of a speaker, a voice-over or narration over slides is far more forgiving than a close-up of someone talking or a drummer hitting a snare.

This matters for any work that changes audio. Keeping the replacement within a few tens of milliseconds of where it should be is good practice; letting a whole track sit a quarter of a second early is not.

Why a replacement audio track must share the original start time

When an audio track is replaced, the player still does exactly what it always did: match timestamps against the audio clock. It has no idea the new track is a translation, a cleaned-up mix or a voice-over. If the new track's timeline is shifted relative to the picture, every sound is shifted with it.

  • The replacement must start at the same presentation time as the audio it was timed against. If the original audio began 0.3 s after the first frame and the new track begins at zero, every line plays 0.3 s early.
  • Its sample rate and timestamps must be read correctly, or the offset grows through the video instead of staying constant.
  • The video stream should keep its own timestamps. Re-encoding it at a forced frame rate can move frames relative to a soundtrack timed against the original.
  • If the replacement is shorter or longer than the video, players handle the difference by playing silence or continuing the picture on the system clock; neither causes sync errors inside the overlap.

Timing the content inside the new track, such as where each translated sentence starts, is a separate problem from the track-level alignment described here. That is covered in syncing translated audio with video.

Common sync misconceptions and limits

  • Lip sync and AV sync are not the same thing. AV sync means the soundtrack is aligned in time; matching mouth movements to different words in another language is a separate technique, explained in AI lip sync.
  • A file that plays in sync in one player and not another is not necessarily broken. Players differ in how they treat start offsets, edit lists and device latency.
  • Variable frame rate does not by itself cause sync errors in a player, because timestamps carry the true times. Problems start when an editor or converter assumes a constant rate.
  • Dropping frames to stay in sync is normal behavior, not a fault. Constant drops, though, point to a device that cannot decode the file smoothly.

mydubly and audio sync

A dubbed video from mydubly is an audio replacement in exactly the sense above. The original picture is copied packet by packet into a new file, without re-encoding, and a dubbed AAC track takes the place of the original audio. That track keeps the original music and sound effects, separated from the original speech and mixed under the translated voice, on the same timeline. The output is MP4, or MKV when the video codec cannot go in MP4.

Inside the voice track, each translated sentence is placed against the timing of the original speech, which the speech recognizer reports on the original audio's timeline. A line may start up to 0.3 s early or end up to 0.6 s late so that translations of different lengths fit. mydubly does not do lip-sync, so mouths will not match the new words; viewers generally accept this for narration, training and presentation content more readily than for close-up drama. If that trade-off matters for your video, our guide to subtitles versus dubbing helps you decide, since subtitles in SRT or VTT are timed to the same timeline and leave the original voices intact.

Next step: check sync on more than one device

Play your source video in two different players, one with wired audio, and watch a close-up of someone speaking. If both look right, the file's timestamps are sound and it is a good candidate for AI dubbing or the mydubly video translator. If one player is off and the other is not, the device or player is the suspect, not the file.

Frequently asked questions

Why do players use audio as the master clock?

Because gaps or repeats in sound are much more noticeable than a frame shown slightly longer or skipped. Using audio as the reference lets the picture absorb small timing errors invisibly. Some professional and live systems use video or an external clock instead, for their own reasons.

How much audio delay is noticeable?

It depends on the viewer and the content, so there is no single precise figure. Offsets of a few milliseconds are invisible, while offsets in the tens of milliseconds can become noticeable on close-ups of speech. Viewers are generally more forgiving of sound that arrives late than sound that arrives early.

Why does my video stay in sync even though frames are dropped?

Dropping frames is how the player keeps sync. When decoding falls behind the audio clock, late frames are skipped so the picture catches up while the sound continues. You may see a small hitch in motion, but lips and words stay together.

Do Bluetooth headphones break audio sync?

They add delay, and whether you notice depends on whether the operating system and player compensate for it. Many modern systems report Bluetooth latency so players can delay the picture to match. Where compensation is missing, every video looks out of sync on that device only.

Is a dubbed video lip-synced?

Not necessarily. Track-level sync only means the new audio is aligned with the picture's timeline. Matching mouth movements to translated words is a separate technique, and many dubbing workflows, including mydubly's, do not do it.

What happens if a replacement audio track is shorter than the video?

Players continue showing the picture after the audio ends, usually switching to the system clock for the remaining frames. Sync within the part where both exist is unaffected. The viewer simply hears silence at the end.