Video engineering & media pipelines

PTS and DTS: How Video Files Record When Each Frame Is Decoded and Shown

Every packet in a video file carries a presentation timestamp (PTS), which says when its frame or audio should be shown or heard, and usually a decode timestamp (DTS), which says when it must be decoded. The two differ only when frames are stored out of display order, which happens whenever a codec uses B-frames. Both are counted in a timebase, a fraction of a second chosen per stream, and getting them right is what keeps a replacement audio track lined up with the picture.

9 min read · Updated

Two clocks for every packet

A video player has to answer two questions for each compressed frame: when should it be decoded, and when should it be displayed? For audio the answers are almost always the same, so audio packets usually have identical PTS and DTS. For video they can differ.

  • PTS, the presentation timestamp, is the moment the decoded frame should appear on screen, or the moment the first sample of an audio packet should be heard.
  • DTS, the decode timestamp, is the order and time in which packets must be fed to the decoder.

Two rules hold in any valid stream. DTS values only ever increase from one packet to the next, because a decoder takes packets in sequence. And a frame's PTS is never earlier than its DTS, because a frame cannot be shown before it has been decoded.

Timebases: the unit timestamps are counted in

Timestamps are stored as whole numbers, not as seconds. Each stream declares a timebase, a fraction such as 1/90000 or 1/48000, and a timestamp of 180000 in a 1/90000 timebase means 2 seconds.

1/90000
The classic MPEG clock, used for video in transport streams and common in many tools
1/48000
Typical for 48 kHz audio tracks in MP4, so one tick equals one sample
1/30000
A common choice for 29.97 fps video, where each frame lasts 1001 ticks
1/1000
Millisecond precision, the default in Matroska files
1/1000000
Microseconds, used by browser APIs such as WebCodecs

Different streams in the same file can, and often do, use different timebases. Tools convert between them constantly, and every conversion can round. A frame rate like 29.97, which is really 30000/1001, cannot be represented exactly in milliseconds, so the frame durations in a Matroska file vary between 33 and 34 ms. Players handle this without trouble, but conversion code that rounds carelessly can accumulate drift over a long file. The background on these odd frame rates is in video frame rates explained.

Why B-frames make PTS and DTS differ

Modern codecs compress most frames by referring to others. An I-frame stands alone, a P-frame refers to earlier frames, and a B-frame can refer to frames both before and after it. A B-frame cannot be decoded until the later frame it depends on has been decoded, so the encoder stores that later frame first.

Display order
I1 B2 B3 P4 B5 B6 P7
Decode order
I1 P4 B2 B3 P7 B5 B6
PTS of P4
Frame 4 in display time
DTS of P4
Second packet in the file, decoded before B2 and B3

The decoder receives frames in decode order, holds P4 in a buffer, decodes B2 and B3, and then the player shows them in PTS order: I1, B2, B3, P4. The DTS sequence rises steadily; the PTS sequence, read in file order, jumps forward and back. Files without B-frames, such as many hardware-encoded camera files or streams encoded for very low latency, have PTS equal to DTS for every packet. The frame types themselves are covered in our guide to keyframes and GOP structure.

Start offsets: not every stream starts at zero

Nothing requires a file's first timestamp to be zero. A transport stream from a broadcast or a recorder may begin at whatever its clock happened to read. A clip cut from a longer recording may keep its original times. And audio and video in the same file often start a few milliseconds apart because the camera's encoders started at slightly different moments.

What matters is the relationship between streams. If the video's first PTS is 1.000 s and the audio's is 1.021 s, the audio is meant to begin 21 ms after the first frame. A player honors that. A tool that strips the offsets and starts both at zero shifts the sound earlier by 21 ms, which nobody will notice; the same mistake with a 600 ms offset is very noticeable. Running ffprobe with -show_streams reports each stream's start_time, which is the fastest way to spot this.

Negative timestamps and edit lists

B-frames create a small puzzle at the start of a file. If the first displayed frame should have PTS 0, the first decoded frame needs a DTS below 0, because decoding runs ahead of display. Containers deal with this in different ways.

  • MP4 stores DTS per sample and a composition offset that gives PTS. Newer versions of the format allow negative composition offsets, so the first frame can be shown at 0.
  • Older MP4 files, and many encoders, shift everything forward instead, so the first frame is shown slightly after zero, and then use an edit list to say where presentation should really begin.
  • Matroska and transport streams simply allow the timestamps they need, including small negative values in some tools.

An edit list, the elst box in MP4 and MOV, maps the media timeline onto the presentation timeline. It can say "skip the first 0.067 s of this track" or "start this track 0.5 s into the presentation". Editors use edit lists to trim without rewriting media, and encoders use them to hide the short run of priming samples that AAC and some other audio encoders add at the start.

The catch is that not every tool reads edit lists the same way. One player honors them and plays the file in sync; another ignores them and plays the audio a fraction of a second early. Remuxing can drop an edit list or bake it into the timestamps. When a file plays correctly in one app and slightly out of step in another, an edit list or start offset is a good first suspect. Full diagnosis is the subject of audio and video out of sync.

Example: a trimmed interview clip

A hypothetical 12-minute clip was trimmed in an editor that exported it without re-encoding, using an edit list that starts the video 0.4 s into its first GOP. One player shows it perfectly. An older tool ignores the edit list and also decodes those 0.4 s of hidden frames, so every spoken word appears to land 0.4 s early. Re-exporting the clip so that both streams genuinely start at zero removes the ambiguity.

Why timestamps matter when an audio track is replaced

Swapping a soundtrack sounds simple: drop the old audio, put new audio in. In timestamp terms, it means the new track must be laid onto the same presentation timeline the old one used.

  1. The new audio's first sample should line up with the moment the old audio started, not just with time zero. If the original audio began 0.25 s after the first frame, so should anything timed against it.
  2. Any timing inside the new track, such as a translated line placed to match a speaker, inherits the clock of the audio it was timed against. If that clock was shifted, every line shifts by the same amount.
  3. The video packets should keep their own timestamps. Re-timing them, for example by forcing a constant frame rate on a variable-frame-rate recording, moves frames relative to everything timed against the original.
  4. The container must still carry any start offset or edit list the video relies on, or the picture itself starts at a different moment.

A constant offset across the whole video usually means one of these start times was lost. Drift that grows over time usually means a frame rate or timebase was misread, which our guide to variable frame rate discusses.

Common timestamp mistakes and pitfalls

  • Treating frame numbers as time. Frame 1,800 is 60 s at 30 fps but 60.06 s at 29.97 fps. Work in timestamps, not counts.
  • Concatenating files without adjusting timestamps, so the second file's times restart at zero and the muxer reports non-monotonous DTS.
  • Assuming the first timestamp is zero and subtracting nothing, which fails on broadcast captures and many cut clips.
  • Forcing a new frame rate during a copy, which rewrites PTS and quietly breaks sync with existing audio.
  • Ignoring audio priming. Re-encoding audio without accounting for the encoder delay can add a few tens of milliseconds of offset each generation.

mydubly and timestamps

mydubly works on a single timeline: the original audio. In the browser, the default audio track is decoded and downmixed to 16 kHz mono, then cut into chunks of about 30 seconds. Speech recognition skips silence with a voice activity filter, but the timestamps it returns still refer to the original audio, so subtitle cues and the timestamped transcript line up with the recording as it plays.

For a dubbed video, each translated line is fitted to the timing of the original speech: it may start up to 0.3 s early or end up to 0.6 s late, and speech may be sped up or slowed slightly. The details are in syncing translated audio with video. The finished track, with the translated voice mixed over the original background, is then combined with your original picture in the browser, with the video packets copied one by one rather than re-encoded, and the original audio track left out.

If your source file has an unusual start offset or relies on an edit list, the safest preparation is the same as for any audio replacement: export it from your editor so that picture and sound both start at zero. That removes any doubt about which clock the dub should follow. You can then run the file through the mydubly video translator or follow the steps in how to dub a video with AI.

Next step: inspect your timestamps

Run ffprobe -show_streams on a file you plan to work with and compare the start_time of its audio and video streams. If they match or differ by a few milliseconds, nothing needs fixing. If they differ by tenths of a second, or ffprobe reports an edit list you did not expect, re-export the file before replacing its audio or creating a version with AI dubbing.

Frequently asked questions

Are PTS and DTS the same for audio?

Almost always. Common audio codecs do not reorder packets, so each packet is decoded and played in the same order and the two timestamps match. Differences you see in audio usually come from start offsets or priming, not reordering.

What does non-monotonous DTS mean in ffmpeg?

It means a packet arrived with a decode timestamp that is not greater than the previous one, which a valid stream should never have. It often appears when files are concatenated without adjusting their timestamps, or when a damaged or live-captured stream has resets. ffmpeg usually adjusts the values and continues, but the output deserves a playback check.

Why does my file start with negative timestamps?

Usually because the video uses B-frames, so the first frames must be decoded slightly before the first one is displayed. Some tools represent that lead with negative decode timestamps rather than shifting everything forward. It is normal and players handle it.

Do edit lists cause audio sync problems?

They can, when one tool honors an edit list and another ignores it. The file is technically correct, but the second tool shows hidden frames or plays priming audio and shifts everything slightly. Re-exporting so both streams start at zero avoids relying on the edit list.

Does remuxing change PTS and DTS?

A plain stream copy keeps the relative timestamps of each packet, though the muxer may convert them to a different timebase or shift the whole file so it starts at zero. Problems arise when options such as forced frame rates rewrite the times, or when an edit list is dropped. Comparing ffprobe output before and after is a quick check.

What timebase should I use?

Usually the one your tool picks. Encoders and muxers choose timebases that can represent the frame rate and sample rate exactly, such as 1/30000 for 29.97 fps or 1/48000 for 48 kHz audio. Only change it if you have a specific reason, such as a container that requires a fixed timebase.