AI video translation

Inside an AI video translation pipeline, stage by stage

AI video translation works as a chain of specialized stages: the audio is pulled out of the video, cut into short pieces, transcribed by a speech recognition model, translated by a machine translation engine, read aloud by a text-to-speech model, stretched or squeezed to fit the original timing, and finally combined with the original picture. Each stage hands text or audio to the next, so understanding the chain tells you where mistakes come from and what you can do about them. This article walks through each stage as mydubly implements it.

8 min read · Updated

The pipeline at a glance

Most production systems today are cascaded: separate models for recognition, translation and speech, rather than a single model that hears one language and speaks another. Cascades are easier to inspect, because every intermediate result is plain text or plain audio that you can read or listen to.

Extraction
Decode the soundtrack from the video container into raw audio the models can use.
Chunking
Split the audio into windows of roughly 30 seconds, cut at pauses.
Recognition
Convert each chunk to text segments with start and end times.
Translation
Translate each chunk's segments into the target language.
Synthesis
Generate speech for the translated text with a chosen voice.
Timing fit
Place each synthesized line against the original timing and join everything into one track, mixed over the original music and effects.
Reassembly
Put the new audio track back with the untouched video stream.

Stage 1: extracting audio without uploading the video

A video file is a container holding separately encoded streams: usually one video stream and one or more audio streams. The first job is to decode the audio stream into raw samples. In mydubly this happens in your browser, using a WebAssembly build of ffmpeg that decodes the soundtrack once to 16 kHz mono 16-bit PCM, the format speech recognition models such as Whisper expect. For a dub it decodes at 24 kHz instead, because the same audio also supplies the music and effects kept under the new voice.

Browsers have tight memory limits, so the source file is mounted for ffmpeg to read rather than copied into WebAssembly memory. That is why multi-gigabyte files work. The decoded audio is much smaller than the video: two hours of 16 kHz mono PCM is about 230 MB, compared with several gigabytes for the original footage.

What can go wrong: an unusual or damaged audio codec may fail to decode, and a file with several audio tracks may have the wrong one picked up. If extraction fails, re-exporting from your editor with a standard AAC soundtrack usually solves it.

Stage 2: cutting the audio into speech-friendly chunks

Recognition models work on short windows (Whisper was trained on 30-second segments), and short pieces can be processed in parallel. The naive approach cuts every 30 seconds exactly, which slices words in half and leaves both chunks with a garbled fragment. mydubly instead looks at the last 6 seconds before each 30-second mark, finds the quietest 50 ms frame, and cuts there, so cuts land in pauses between words or sentences. A short leftover tail is folded into the final window rather than sent as a tiny chunk of its own.

Each chunk is then compressed with the browser's built-in WebCodecs encoder to Ogg Opus at 32 kb/s, about 120 KB per 30 seconds and roughly eight times smaller than the raw PCM. Where a browser can't encode Opus, the chunk falls back to WAV. Several chunks upload in parallel. The article on audio chunking goes deeper on the boundary trade-offs.

What can go wrong: a speaker who talks for 30 seconds without any pause leaves no clean place to cut, so the quietest frame may still fall inside a phrase.

Stage 3: speech recognition with Whisper

mydubly's transcription runs on OpenAI's open-source Whisper model, by default the large-v3-turbo variant. Whisper turns each chunk into a log-Mel spectrogram, a picture of how energy is spread across frequencies over time, and an encoder-decoder transformer predicts text tokens from it. The model also identifies the spoken language, which is why you never have to choose a source language. For each chunk it returns text segments with start and end timestamps, which are then offset by the chunk's position to place them on the full video timeline.

What can go wrong: names and jargon get replaced with common words that sound similar; heavy music or noise lowers accuracy; overlapping speakers usually produce only the dominant voice; and Whisper-family models can occasionally invent text in long silences. See Whisper hallucinations for the failure modes in detail.

Stage 4: machine translation, chunk by chunk

Each chunk's segments go to a neural machine translation engine, which produces target-language text segment by segment. Because the translation is done per chunk, the engine sees roughly 30 seconds of context at a time. That is plenty for most sentences, but it means a term introduced in minute three is not guaranteed to be rendered identically in minute forty.

What can go wrong: any recognition error is translated faithfully; idioms come out literally; ambiguous pronouns or formal and informal address may be guessed wrong; and some language pairs expand noticeably, producing longer lines that the timing stage then has to absorb.

Stage 5: synthesizing the translated speech

mydubly's default voice engine is Chatterbox Multilingual, an open-source text-to-speech model from Resemble AI. Some of the preset voices are produced by conditioning the model on a reference recording, so it reproduces that voice's character in the target language. Before synthesis, short recognition fragments are merged into whole sentences of up to about 220 characters, because the model's intonation is noticeably better on complete sentences than on clipped pieces.

What can go wrong: unfamiliar names may be mispronounced, numbers and abbreviations can be read awkwardly, and emphasis that the original speaker carried with their voice may flatten.

Stage 6: fitting the new speech into the original timing

Translated speech rarely matches the original line's duration. On the server, each synthesized clip is first normalized: converted to mono at 24 kHz, trimmed of silence at the edges, and brought to a consistent loudness so chunks don't jump in volume. Then lines are placed elastically. A line may start up to 0.3 seconds early or run up to 0.6 seconds late. Speeding up to about 1.08 times is treated as inaudible; beyond that the tempo rises up to 1.15 times by default, with a little more allowed only when a chunk otherwise cannot fit. Lines that come out too short are slowed slightly, never below 0.9 times. Finally the chunks are joined gaplessly into one AAC track exactly as long as the video. That track also carries the original background: AI vocal separation removes the original speech from the uploaded audio, and the remaining music, ambience and effects are mixed under the new voice and lowered automatically while it speaks. The output is mono.

Illustration: a German line that runs long

Suppose an English line occupies 3.5 seconds and its German translation is synthesized at 4.6 seconds. If there is quiet on both sides, starting 0.3 seconds early and ending 0.6 seconds late gives a 4.4-second window. Fitting 4.6 seconds into it needs a speed-up of about 1.05 times, inside the range treated as inaudible, so the listener hears natural pacing and the line still ends close to where the original did.

What can go wrong: in fast, dense speech there is no slack around lines, so the tempo has to rise toward its limit and the voice can sound hurried. The timing article covers this stage in more depth.

Stage 7: reassembly in the browser, and what mydubly keeps

The finished AAC track comes back to your browser, where it is muxed with the original video stream using the mediabunny library. Muxing means packaging encoded streams into a container without decoding them, so the picture is copied bit for bit and never re-encoded, and there is no quality loss or long render. If the original video codec can't go into an MP4 container, the output is MKV instead. Large outputs can be written straight to disk rather than held in memory.

That split between device and server is deliberate. The video file never leaves your device: your browser uploads compressed audio chunks, the server returns text and the translated audio, and your browser assembles the final video. Uploaded audio and results are deleted within 30 minutes of a job finishing, and unfinished jobs expire after 24 hours. The private video translation page explains that design from a privacy angle.

Trade-offs of a cascaded design

The cascade's main weakness is compounding: an error at recognition flows unchanged into translation and then into the voice. Its main strength is transparency. Because the pipeline produces transcripts in both languages, you can find exactly where a mistake entered, whether the recognizer misheard or the translator misread a correct transcript, and fix the right thing. Direct speech-to-speech models avoid some compounding but give you nothing to inspect in between.

A second trade-off is context. Chunking keeps memory use low and allows parallel work, but each recognition and translation step sees only its own window, which is why consistent terminology across a long video needs a human check.

Try the pipeline on your own file

The quickest way to see every stage's output is to run a two-minute clip through the video translator with full output on, then compare the source transcript, the translated transcript and the dubbed MP4 side by side. For the practical step-by-step, follow how to translate a video.

Frequently asked questions

Why does the audio get split into 30-second pieces?

Whisper-style recognition models were trained on windows of about 30 seconds, and short pieces can be uploaded and processed in parallel. Cutting at the quietest point near each boundary keeps words and sentences intact.

Is the video re-encoded when the translated audio is added?

Not in mydubly. The new audio track is muxed with the original video stream, which copies the picture data without decoding it. That avoids quality loss and keeps the final step fast.

Why is the output sometimes MKV instead of MP4?

Some video codecs aren't allowed inside an MP4 container. When the original stream can't go into MP4 unchanged, mydubly packages it in MKV rather than re-encoding the picture.

How does the translated voice stay in sync with the speaker?

Each line is placed against the original line's timing. It may start slightly early or end slightly late, and its tempo can be nudged up or down within limits. It is timing sync, not lip-sync; the mouth movements are not changed.

Which stage causes most visible mistakes?

Usually recognition. A misheard name or term is carried through translation and voiced with confidence. Reading the source-language transcript first tells you whether a translation error started in the audio or in the translation step.