Subtitles & captions in depth

Professional Subtitle Spotting: Durations, Gaps and Shot Changes

Subtitle timing, often called spotting, decides exactly when each cue appears and disappears. Professional style guides converge on a few rules: keep every cue on screen long enough to register but not so long that it lingers, leave a short gap between consecutive cues, avoid crossing shot changes, enter with the speech and leave shortly after it, and chain cues together when speech is continuous. Exact numbers vary by broadcaster, platform and language, so treat the figures here as common ranges and follow your client's guide where one exists.

8 min read · Updated

What spotting means and why it matters

Spotting is the job of setting in and out times for every subtitle. A file can contain a perfect translation and still feel amateurish if cues flicker on for a few frames, pile up against each other or straddle a cut. Good spotting is mostly invisible: the viewer's eye moves to the bottom of the screen as someone speaks, reads, and returns to the picture without noticing any mechanics.

Spotting sits between two other tasks. Before it comes segmentation, the decision about how text is divided into cues, which is covered in subtitle segmentation. After it, or alongside it, comes the reading-speed check: whether a cue's text can be read in its duration. Reading speed, along with how recognition timestamps become cues and why files drift out of sync, is explained in subtitle timestamp alignment. This article focuses on the timing conventions themselves.

Time-based and frame-based timing

SRT and VTT store times in milliseconds, but much professional guidance is written in frames, because subtitles are judged against picture. Translating between the two is simple arithmetic: one frame lasts one second divided by the frame rate.

23.976 or 24 fps
About 41.7 ms per frame
25 fps
40 ms per frame
29.97 or 30 fps
About 33.3 ms per frame
Two frames at 25 fps
80 ms
Two frames at 24 fps
About 83 ms

When a guide says "two frames", it means a gap that depends on the frame rate of the video. Working in a subtitle editor that knows the frame rate lets it round in and out times to whole frames, so cues never start halfway through one.

Minimum and maximum duration

A subtitle that appears too briefly is seen as a flash rather than read, even if it holds a single word. Many style guides set a minimum duration somewhere between about five-sixths of a second and one second, with short interjections such as "Yes" or "Hey" given the minimum rather than the length of the spoken word.

At the other end, a cue that stays on screen long after its text has been read makes viewers read it again, and it can hide the start of the next line. Commonly cited maximums sit around six to seven seconds for a full two-line subtitle. If speech runs longer than that without a pause, split the text into two cues at a natural break rather than stretching one.

  • Never let a cue's duration fall below your minimum just to match a very short utterance; extend the out time instead, provided it does not collide with the next cue.
  • If an utterance is shorter than the minimum and the next cue starts immediately, consider merging the two.
  • Long held cues are usually a sign that text was not split, not that the speaker was slow.

Gaps between consecutive cues

When one subtitle is replaced by another with no break, viewers can miss that the text changed, especially if the two cues have similar shapes. Professional guides therefore require a small minimum gap between consecutive cues, frequently stated as two frames. That gap is short enough not to be noticed as an empty screen, but long enough for the eye to register a new subtitle.

The opposite problem is a gap that is just slightly too long. A blank of a few hundred milliseconds between two cues produces a distracting flicker. The fix is chaining, described below: either close the gap down to the minimum or leave a clear pause.

Lead-in and lag-out

A subtitle should appear when the speaker starts talking. Some guides allow it to come up a few frames before speech, since people perceive a cue that arrives late more readily than one that arrives slightly early. Starting a cue well before speech is a problem, though, because it gives away a line before it is spoken, which can spoil a reveal or a joke.

After speech ends, letting the cue stay a little longer, the lag-out, gives slower readers time to finish. Many guides allow up to roughly half a second to a second of lag-out, depending on the gap to the next cue and whether there is a shot change nearby. Lag-out is where most cues gain the extra time needed to meet a reading-speed limit.

Hypothetical example: timing one exchange

In an interview, the guest says "We started in a garage." from 12.40 to 13.70 seconds, pauses, and continues "Three of us, no money." from 14.05 to 15.60 seconds. A first pass sets the cues exactly to speech, leaving a 350 ms blank in the middle that flickers. The subtitler extends the first cue's out time to 13.97 seconds, which leaves two frames at 25 fps before the second cue enters at 14.05. The second cue gets half a second of lag-out, ending at 16.10, because the next line does not start until 17.30.

Shot changes

A cut makes the eye re-scan the picture. If a subtitle disappears or appears a few frames away from a cut, the viewer sees two visual events in quick succession and often perceives the subtitle as flickering. For that reason most professional guides ask subtitlers to respect shot changes:

  • If a cue would end within a short window before a cut, often around half a second, end it on the cut or a couple of frames before it.
  • If a cue would start just after a cut, move its in time to the cut.
  • Let a cue cross a cut only when the speech genuinely continues over it, and then keep it on screen for a clear stretch on each side, not a few frames.
  • In fast montages, accept that not every cut can be respected; prioritize readable durations and synchronization with speech.

Shot changes are invisible to anything that works from audio alone, so they must be checked against the picture. Many subtitle editors can detect shot changes from the video and show them on the timeline, then snap in and out times to them.

Chaining cues through continuous speech

When someone talks without pausing, the subtitles should feel like a continuous stream rather than a series of separate blinks. Chaining means closing short gaps between cues down to the minimum gap so the bottom of the screen is never blank for a fraction of a second.

A common working rule is to chain whenever the natural gap between two cues is shorter than some threshold, often in the range of half a second to a second, and to leave longer gaps as real pauses. The threshold matters less than consistency: once you pick one, apply it across the file. Subtitle editors usually offer a gap-bridging or continuity tool that does this in one pass, which is far less error-prone than adjusting hundreds of cues by hand.

Common timing mistakes and their limits

  • Timing each cue strictly to the audio and leaving hundreds of short, flickering gaps.
  • Cues that run into or overlap the next one, which some players show as two stacked subtitles.
  • Ignoring shot changes, so subtitles flash off a few frames before a cut and the next appears just after it.
  • Holding cues through long silences, which leaves stale text over a new scene.
  • Applying frame-based rules from one guide to video at a different frame rate without converting.
  • Assuming any timing rule fixes text that is simply too long; if the reading speed is too high after extending lag-out, the text needs condensing or splitting.

Rules also have limits of their own. Fast cross-talk, songs and rapid montages can make it impossible to satisfy every rule at once; in those passages, accurate synchronization with speech usually wins over tidy gaps.

How mydubly's timings relate to these rules

mydubly creates SRT and VTT files from speech recognition. Each cue is one recognized segment, on a single line, numbered in order, and its in and out times come from the recognizer's segment timestamps. Silence is skipped by a voice activity filter during recognition while timestamps still refer to the original audio, so cues line up with the speech in your file. When you choose a target language, the subtitle file is in that language; in a dubbed job, the target-language subtitles follow the dubbed lines.

What the files do not do is apply spotting conventions. Timing is derived from audio only, so mydubly knows nothing about shot changes, and it does not enforce minimum durations, two-frame gaps, lead-in, lag-out or chaining thresholds. Treat the output as accurately placed raw timing and do the spotting pass in a subtitle editor. The subtitle generator handles files up to 2 hours, and a 30-minute video costs 30 credits (3¢) for subtitles and transcripts. Mechanics such as shifting or splitting cues are covered in how to edit an SRT file.

Next step: run a spotting pass on one file

Take a finished subtitle file and open it in a subtitle editor with the video and a waveform visible. Set your frame rate, minimum duration, minimum gap and chaining threshold, then apply the editor's gap-bridging tool and step through every shot change. If you are starting from scratch, generate an SRT file from the final cut first, then spend your time on the spotting rather than on typing.

Frequently asked questions

What is the minimum duration for a subtitle?

Many professional style guides set it somewhere between about five-sixths of a second and one second, even for a single word. Shorter cues tend to be seen as a flash rather than read. Check the guide for your platform or client, since exact values differ.

How long should the gap between two subtitles be?

A common professional minimum is two frames, which is 80 ms at 25 fps. Gaps slightly longer than that but shorter than a natural pause, roughly up to half a second or so, are usually closed to the minimum so the screen does not flicker. Longer gaps can be left as genuine pauses.

Should subtitles appear before the speaker starts talking?

They should appear with the speech, and some guides allow a few frames early because late subtitles are noticed more than slightly early ones. Starting a cue well ahead of the speech gives the line away before it is said. It can also spoil reveals and jokes.

Can a subtitle stay on screen across a shot change?

Yes, when the speech continues across the cut, but it should then stay for a clear stretch on both sides. What guides discourage is a cue that ends or starts a few frames away from a cut, which looks like flicker. Snapping in and out times to the cut usually solves it.

What does chaining subtitles mean?

Chaining is closing short gaps between consecutive cues so that continuous speech produces a continuous run of subtitles. You choose a threshold, and any gap shorter than it is reduced to the minimum gap. Most subtitle editors have a tool that bridges gaps across the whole file at once.

Do automatic subtitles follow professional timing rules?

Usually not fully. Automatic timings come from where speech is detected, so they tend to be well placed relative to the audio but ignore shot changes, minimum gaps and chaining. A spotting pass in a subtitle editor brings them up to professional conventions.