How it's built

From Recognition Segments to Subtitle Cues That Stay in Sync

Subtitle timestamp alignment is the process of giving every subtitle cue a start and end time that matches when the words are spoken. In an automatic pipeline those times come from the speech recognizer's segment timestamps, which are relative to the chunk of audio that was recognized and must be shifted into the full video timeline. When the offset is exact and the source file is unchanged, subtitles stay in sync from the first cue to the last; drift almost always means the timeline changed somewhere along the way.

9 min read · Updated

What a subtitle timestamp actually encodes

A subtitle file is a list of cues. Each cue has a start time, an end time and the text to show between them. The player compares those times with its playback clock and displays whichever cue is active. Nothing in an SRT or VTT file refers to video frames, so the file only works if its clock and the video's clock agree.

The two formats differ in small but strict ways:

SRT cue
A sequence number, a line such as 00:01:01,605 --> 00:01:04,245, then one or more lines of text and a blank line
VTT cue
An optional identifier, the same time line with a period instead of a comma before the milliseconds, then the text
VTT header
The file must start with the word WEBVTT on its first line
Precision
Both formats store hours, minutes, seconds and milliseconds
Styling and position
VTT supports cue settings for position and alignment; plain SRT has none

The choice between them is mostly about where the file will be used, which our guide to SRT vs VTT covers. For alignment, they behave identically: a cue is shown when the playback clock passes its start time.

Where the times come from in speech recognition

A recognizer such as Whisper does not output words with a stopwatch attached. It predicts text tokens, and among them special timestamp tokens that mark where a segment begins and ends. In Whisper those timestamp tokens step in 20-millisecond increments within the 30-second window the model hears at once. A segment is typically a phrase or a sentence, often a few seconds long.

That design has consequences for subtitles:

  • Times are attached to segments, not individual words, unless a separate word-level alignment step is added.
  • A segment boundary can include a little of the silence around the speech, so a cue may appear slightly before the first word or linger after the last.
  • Timing is only as reliable as the model's judgement of where speech starts and ends; in music or heavy noise that judgement gets worse.

Some pipelines refine the times with forced alignment, which matches the recognized text against the audio phoneme by phoneme to get word-level boundaries. It is more precise, adds a second model and a second failure point, and is mostly worth it for karaoke-style highlighting or very tight broadcast work.

Offsetting chunk-relative times into the full timeline

Long recordings are recognized in chunks of roughly 30 seconds, and the recognizer reports each segment relative to the start of its own chunk. A segment at 6.48 seconds means 6.48 seconds into that chunk, not into the video. Before the segments can become cues, every start and end time must have the chunk's own start time added to it.

This is simple arithmetic, but only if the chunk start times are exact. That is one reason the cut points matter: when chunks are non-overlapping and contiguous, each chunk's start is a single known number and the offset introduces no error at all. Overlapping chunks need an extra decision about which copy of a repeated segment to keep. How cut points are chosen is explained in audio chunking for speech recognition.

Worked example: offsetting one segment

Suppose a recording was cut into windows starting at 0, 27.425 and 55.125 seconds. The third chunk's recognizer reports a segment from 6.48 to 9.12 seconds. Adding the chunk start gives 61.605 to 64.245 seconds, which is written in SRT as 00:01:01,605 --> 00:01:04,245 and in VTT as 00:01:01.605 --> 00:01:04.245. Forgetting the offset would put that line at six seconds into the video, almost a minute early.

The last step is formatting. Times are rounded to the nearest millisecond, split into hours, minutes, seconds and milliseconds, and zero-padded. Rounding to milliseconds is far finer than anyone can perceive, so it never causes visible drift on its own.

How mydubly builds its SRT and VTT files

mydubly follows exactly this offset-and-format approach. The browser splits the audio into pause-aligned windows of about 30 seconds and uploads them; Whisper, with large-v3-turbo as the default model, returns timed segments for each one; and the server shifts every segment by its window's start time before writing the files. The details:

  • One cue per recognized segment, numbered in order. Segments with no text are skipped, so numbering never has gaps.
  • The SRT and VTT files contain the same cues with the same times; only the formatting differs.
  • Translated subtitles keep the source segment's start and end times. Each segment is translated as a unit, so the translated line appears exactly when the original was spoken.
  • In full translation output, which includes a voice track, short fragments are merged into whole sentences before the voice is generated, and the subtitle cues follow those same sentence groups, so a cue can cover what were two very short segments.
  • In transcript mode, the timestamped transcript lists each cue's start time in minutes and seconds before its text, which is handy for quoting or navigation; see the timestamped transcript format.

Subtitles are delivered as SRT and VTT files, not burned into the picture, so you can correct a cue's time in any subtitle editor before publishing. The subtitle generator produces both formats from a video or audio file.

Cue length, line length and reading speed

Correct timing is necessary but not sufficient. A cue that is perfectly in sync can still be unreadable if it holds too much text for its duration. Reading speed is usually measured in characters per second: count the characters in a cue, including spaces, and divide by its duration.

Broadcaster and streaming style guides commonly target somewhere between the mid-teens and about 20 characters per second for adult viewers, lower for children's content, with lines of roughly 37 to 42 characters and no more than two lines per cue. Exact limits differ by organization and language, so check the guide that applies to your platform rather than treating any single number as a rule.

Translation makes this harder. A translated line keeps the original line's time slot but may need more characters to say the same thing. A cue lasting 2.6 seconds with 52 characters reads at about 20 characters per second; if its translation runs to 68 characters in the same slot, that is about 26 per second, which many viewers will struggle to finish. Why translations grow or shrink by language pair is covered in text expansion in translation.

Diagnosing subtitle drift by its pattern

When subtitles are out of sync, the shape of the error usually tells you its cause. Watch three points: the first minute, the middle and the last minute.

Same offset everywhere
The video was trimmed or padded at the start after the subtitles were made, or the subtitles were made from a different cut
Error grows steadily over time
A frame-rate or speed change between versions, such as a 23.976 fps source and a 25 fps broadcast copy
Sudden jump at one point
A scene was cut or inserted in the edit, or two files were joined at that point
Individual cues early or late
Normal segment-level imprecision, or speech over music where the recognizer misjudged the boundaries
Errors only at certain intervals
A chunk offset bug in a pipeline, where some chunks were shifted by the wrong start time

Steady growth is the easiest to recognize numerically. If a film at 23.976 frames per second is sped up to 25 for broadcast, everything runs faster by a ratio of 25 to 23.976, about 1.043. Over an hour of the original, subtitles timed to the original would fall behind by roughly two and a half minutes on the faster copy. A constant shift cannot fix that; the timeline has to be stretched.

Fixing out-of-sync subtitles step by step

  1. Confirm the subtitles were generated from the exact file you are playing. Re-exports, trimmed versions and platform-processed copies are the usual culprits.
  2. Measure the error at the start and at the end. Note how many seconds early or late a clearly spoken line appears in each place.
  3. If both errors are the same, apply a constant shift in your subtitle editor or player.
  4. If the error grows, use the editor's stretch or two-point sync feature: pin the first and last lines to their correct times and let it scale everything in between.
  5. If the error jumps at one point, split the file there and shift each part separately, or regenerate subtitles from the final edit.
  6. Spot-check a few cues in the middle afterwards, since a correct start and end do not prove the middle is right.

Regenerating from the final file is often faster than repairing. With machine subtitles at 1 credit per minute, re-running a 30-minute video costs 30 credits, which is 3 cents.

When automatic timestamps work well

  • Single-speaker content with clear pauses, such as lectures, tutorials and narration.
  • Recordings where the audio is the original and the video has not been re-edited since.
  • Workflows that need subtitles fast and can tolerate a few hundred milliseconds of looseness on some cues.
  • Translated subtitles, because reusing the source segment times keeps every translated line anchored to the moment it was said.

Limits of machine-generated subtitle timing

  • Cues follow the recognizer's segments. mydubly does not re-split cues by reading speed, so a long segment becomes a long cue; reflow long cues in an editor if your platform has strict limits.
  • Segment boundaries can absorb some surrounding silence, so a cue may appear a moment early or stay up slightly after the line ends.
  • Overlapping speakers and crosstalk produce segments that mix both voices, and cues are not labeled by speaker.
  • Speech over loud music or effects gives the recognizer less certainty about where words begin and end, which shows up as looser timing.
  • Translated cues inherit the source timing, so a much longer translation stays in the same slot and can exceed comfortable reading speed.
  • If you edit the video after generating subtitles, the files will not follow the edit; timing only stays correct for the file that was processed.

Generate subtitles that match your final cut

The most reliable way to avoid drift is to create subtitles from the exact file you will publish, after the edit is locked. Upload that file to the subtitle generator, review a handful of cues at the start, middle and end, then attach the SRT or VTT to your player or platform; the guide on adding subtitles to a video walks through that last step.

Frequently asked questions

Why do my automatic subtitles appear slightly before the speaker starts?

Recognition segments often include a little of the silence before the first word, and the cue inherits that start time. A lead of a fraction of a second is common and generally comfortable to read. If a particular cue is noticeably early, nudge its start time in a subtitle editor.

Can I convert an SRT file to VTT without changing the timing?

Yes. The cue times are identical; VTT uses a period instead of a comma before the milliseconds and needs a WEBVTT header line. Most subtitle editors convert in one step, and mydubly already provides both files with the same cues, as described on the VTT generator page.

My subtitles are fine at the start but late by the end. What happened?

A steadily growing delay almost always means the video you are playing runs at a different speed or frame rate from the one the subtitles were made for. Shifting every cue by the same amount will not fix it. Use a two-point sync to stretch the timeline, or regenerate the subtitles from the file you are actually publishing.

Do translated subtitles use different timestamps from the original language?

In mydubly, translated cues reuse the start and end times of the source segments, so each translated line appears when the original was spoken. The trade-off is that a longer translation must be read in the same time window.

How precise are Whisper's timestamps?

Whisper marks segment boundaries with timestamp tokens in 20-millisecond steps, but the practical accuracy depends on the audio. On clear speech, segment boundaries usually land close to the spoken words; with music, noise or crosstalk they can be off by more. Word-level precision requires an additional alignment step.