How it's built

Keeping a Translated Voice in Time With the Original Speaker

To sync translated audio with video, each translated line has to start roughly when the original speaker starts and finish before the next line is due, even though the translation is almost never the same length as the original. Good pipelines solve this line by line: they anchor every clip to the source timestamp, borrow a few tenths of a second of nearby silence, and only then speed speech up slightly. Done well, the voice track stays aligned for the whole video without anyone hearing the adjustments.

9 min read · Updated

Why translated speech never fits the original slots

A speaker's timing is set by their language, their pace and their pauses. A translation inherits none of that. The same sentence can need noticeably more syllables in one language than another, a point explored in text expansion in translation. A synthetic voice also has its own natural pace, which may be faster or slower than the person on screen. And a single fluent sentence in the source can come back from translation with a different word order, so the emphasis lands somewhere else.

The upshot is that every translated clip arrives with a duration that is close to, but rarely equal to, the slot it should occupy. Some are shorter, some are longer, and a few are much longer. Syncing is the job of deciding where each clip goes and how much, if at all, to change its speed.

Three naive approaches and how each one fails

It helps to look at what does not work.

  • Play the clips back to back. Every clip that runs long pushes all later clips later. After a few minutes the voice is describing a slide that disappeared long ago, and the error never recovers.
  • Start every clip exactly at its source timestamp and ignore length. Sync is perfect at each start, but long clips overlap the next line or get cut off, and two voices talking at once is worse than being late.
  • Time-stretch every clip to exactly fill its slot. Lines that are much longer than their slot get compressed hard and sound rushed or robotic, while short lines get dragged out and sound sleepy.

The working approach sits between these. Starts stay anchored to the original, so error cannot accumulate. Lengths are adjusted, but in a fixed order of preference that keeps speed changes as small as possible.

The tools a timing planner has to work with

A planner has only a few levers, each with a cost:

  • Tempo change. Modern time-stretching changes speed without changing pitch. Small changes are effectively inaudible; larger ones start to sound hurried, and the threshold depends on the voice and the listener.
  • Starting early. If the time just before a line is silent, the line can begin slightly before the speaker's mouth moves. Viewers tolerate a small lead far better than a large lag.
  • Running late. A line can spill a little into the next line's time if the next pause will absorb it, so the next line still starts close to its speaker.
  • Slowing down. A clip that is much shorter than its slot can be stretched a little so it ends closer to when the speaker does, which keeps the voice from finishing while the person is visibly still talking.

Anchors come from the speech recognizer's segment timestamps, the same ones used for subtitles; subtitle timestamp alignment describes how they are produced.

How mydubly places each translated line

mydubly's voice track is assembled on the server, one roughly 30-second chunk at a time, using the chunk's recognized segments as anchors. Before any voice is generated, short fragments are merged into whole sentences of up to about 220 characters, because the default voice engine, Chatterbox Multilingual, sounds more natural reading complete sentences. A fragment is judged by how long it takes to say rather than by its character count, so languages written compactly, such as Chinese or Japanese, are not misjudged.

Then each line is placed with this order of preference:

  1. Anchor the line at the moment the original speaker started that segment.
  2. If the clip is too long for its slot, first speed it up by up to about 1.08 times, a change treated as inaudible.
  3. If that is not enough, start it up to 0.3 seconds early, provided the time before it is free.
  4. Next, let it run up to 0.6 seconds into the following line's time, pushing that line slightly late.
  5. Only then raise the tempo further, up to 1.15 times by default. If a whole chunk still cannot fit inside its window, the ceiling is raised in small steps, with a firm upper limit.
  6. If the clip is shorter than the original speech, slow it slightly, never below 0.9 times, so it ends with the speaker.

The planner also looks ahead. Working backward from the end of the chunk, it calculates how late each line may finish so that every later line can still start within its own allowance. A comfortable line early in a chunk may therefore be kept a little tighter to leave room for a long one that follows. A tenth of a second of space is kept between consecutive lines so they never run together.

Anchor
The start time of the original spoken segment
Inaudible speed-up
Up to about 1.08 times
Early start
Up to 0.3 s before the speaker, if the preceding time is free
Late finish
Up to 0.6 s into the next line's time
Default tempo ceiling
1.15 times, raised in small steps only when a chunk cannot otherwise fit
Slowest stretch
0.9 times, for lines shorter than the original speech
Gap between lines
0.1 s

A worked example: three lines in one chunk

The numbers below follow the placement rules exactly, for a fictional 12-second chunk of a product tutorial.

Line A runs long

The speaker talks from 2.0 to 5.0 s and the next line starts at 5.6 s, but the translated clip lasts 3.9 s. Played at 1.08 times it takes 3.61 s, so the planner starts it 0.11 s early, at 1.89 s, and it ends at 5.5 s, a tenth of a second before line B. No further speed-up is needed.

Line B makes room

The speaker talks from 5.6 to 8.0 s and the clip lasts 2.0 s, so it fits easily. Because the planner already knows line C will be tight, it plays B at about 0.98 times and ends it at 7.64 s instead of stretching it all the way to 8.0 s.

Line C hits the ceiling

The speaker talks from 8.4 to 11.0 s, the chunk ends at 12.0 s, and the clip lasts 4.6 s. It starts the full 0.3 s early, at 8.1 s. It cannot run late, because every chunk must end exactly on its boundary, so it needs about 1.18 times to finish in time. That exceeds the 1.15 default, so this is one of the rare chunks where the ceiling steps up to fit.

The listener hears line A arrive a hair early, line B at its natural pace and line C slightly brisk. Nothing overlaps and the next chunk starts exactly on time.

Matching loudness and joining chunks without gaps

Timing is half of sync; the other half is making the track sound continuous. Each generated clip is first normalized: converted to mono at 24 kHz, trimmed of leading and trailing silence so its measured length reflects actual speech, and loudness-normalized to the same target so lines and chunks sit at a consistent level. Each placed clip gets a fade of a few milliseconds at both ends to prevent clicks, and the mixed chunk passes through a limiter so overlapping edges cannot clip.

Every chunk is then padded with silence or trimmed to exactly the length of its window. Because the windows were contiguous slices of the original audio, the chunks add up to precisely the original duration, and drift cannot build up from one chunk to the next.

Finally the chunks are joined into one AAC track that matches the video length. Joining separately encoded AAC files naively tends to leave tiny gaps, because each encoder run adds a short priming delay at its start. mydubly instead joins the uncompressed audio first and encodes the continuous track. To use several processor cores, it encodes frame-aligned sections in parallel, giving each a few frames of the preceding audio to warm up the encoder and then discarding them, so the joined stream decodes as if it had been encoded in one pass. The browser then swaps this track into the video, as described in browser video muxing.

Where elastic timing works well

  • Narration, lectures and tutorials, where one person speaks at a steady pace with regular pauses.
  • Language pairs where translations come out at broadly similar lengths, so most lines need little or no adjustment.
  • Content with natural breathing room, since every pause is time the planner can borrow.
  • Long videos, because anchoring each line to the original prevents error from accumulating over an hour.
  • Viewers who need to follow along with the picture, such as screen recordings, where a voice that drifts minutes behind would be useless.

What timing fit cannot fix

  • Lip movements. mydubly aligns timing only and never alters the picture, so close-ups will show mouths that do not match the translated words; the AI lip sync article explains what visual lip-sync models do instead.
  • Very dense speech. When a fast speaker leaves no pauses and the translation is much longer, lines reach the tempo ceiling and sound noticeably brisk.
  • A perfect soundtrack. The translated MP4 keeps the original music and effects under the new voice, but separating them from the original speech is not perfect, so faint traces of the original voice can remain in dense mixes; keeping background music when translating video covers the details.
  • Multiple voices. One chosen voice reads every line, so timing is preserved but the sense of different speakers is not.
  • Vocal reactions. Laughter or sighs from the original speakers are not voiced by the new voice; where the original had them, only the background continues.
  • Small local lag. A line can land up to 0.6 seconds behind its speaker, which is rarely noticeable in narration but can show in fast back-and-forth dialogue.

Checking sync on your own dubbed video

  1. Watch the first two minutes with sound on and note any line that starts clearly before or after the speaker.
  2. Jump to the middle and the last two minutes. If those sound as well aligned as the start, the track is not drifting.
  3. Listen for brisk passages. They mark places where the translation ran long; consider shortening those sentences in the source script if you plan to re-record.
  4. Compare with the translated SRT file. Its cues use the original timing, so a voice line that trails its cue by a lot is worth a second look.
  5. For content with frequent on-camera close-ups, decide whether translated subtitles would suit viewers better than a voice, using the guide to subtitles vs dubbing.

Make a timed dub of your own video

The quickest way to judge the result is to dub a short, representative clip with mydubly's AI dubbing and listen at the start, middle and end. A translated voice costs 50 credits per minute with a 2-minute minimum per file, so a 12-minute tutorial costs 600 credits, which is 60 cents. The step-by-step dubbing guide covers choosing a voice and reviewing the output.

Frequently asked questions

Why does my dubbed audio drift further out of sync as the video goes on?

Progressive drift usually means clips were placed back to back, so every long line pushed the rest later. Anchoring each line to its original timestamp, and making every chunk exactly as long as the audio it replaces, keeps error from accumulating. If drift appears only after editing, the voice track was made for a different cut of the video.

How much can speech be sped up before listeners notice?

There is no universal threshold; it depends on the voice, the language and how closely someone listens. Changes of a few percent are generally hard to hear, which is why mydubly treats up to about 1.08 times as inaudible and prefers borrowing silence before going faster.

Does speeding up a translated line change the pitch of the voice?

Not with modern time-stretching, which changes duration while keeping pitch roughly constant. Very large changes can introduce a slightly processed sound, which is another reason good pipelines keep tempo adjustments small and use them last.

Why not shorten the translation instead of speeding up the voice?

Shortening keeps the voice at a natural pace and is what human dubbing adaptors do. It needs a rewrite that preserves meaning, which is harder to automate reliably than a small tempo change. If a passage sounds rushed, editing the source script to be more concise before the next version is a reliable fix.

Can I adjust the timing of individual lines in a mydubly dub?

Not inside mydubly. You can download the translated audio file separately and edit it in any audio or video editor, moving or trimming individual lines against the picture.