Why dubbing is harder than subtitling
A subtitle can be slightly too long and the viewer simply reads a little faster. A voice track has no such slack: every translated line has to occupy real seconds, start close to where the original speaker started, finish before the next line, and still sound like one continuous performance at a believable pace.
Dubbing is also the last stage of a cascade. Speech recognition turns audio into text, machine translation turns that text into another language, text-to-speech voices it and a timing step fits it to the picture. An error introduced early travels through every later stage, and the voice delivers it with complete confidence. A misheard word in a transcript looks like a typo; the same word spoken fluently in a dub sounds like a fact. The full pipeline is laid out in how AI video translation works.
Timing and text expansion
Translations rarely match the source in length. English into Spanish, French or German often comes out longer, while other pairs come out shorter, and what matters for dubbing is spoken duration rather than characters on a page. Text expansion in translation explains why this varies by language pair.
Engineers handle it in layers. The first is elastic placement: a line does not have to start and stop exactly where the original did. The second is gentle tempo adjustment, because small speed-ups are close to imperceptible while large ones sound rushed. The third, used in research and by human adaptation writers, is producing a shorter translation in the first place, sometimes called length-controlled or isochronous translation.
Suppose the English line "Click save before you close the window" occupies 2.6 seconds, and the German version comes back from the voice at 3.6 seconds. If the line may start 0.3 s early and run 0.6 s late, its slot becomes 3.5 seconds, so the clip needs only about 1.03× speed, which nobody will notice. If the German clip were 4.2 seconds instead, it would need 1.2× in that slot, past a 1.15× cap, and the line would sound noticeably hurried.
Emotion and delivery
Text strips out most of what makes speech expressive. Sarcasm, hesitation, excitement and warmth live in pitch, timing and voice quality, and almost none of that survives transcription and translation. A text-to-speech model has to infer delivery from the words alone, so it tends toward a pleasant, neutral reading.
The responses so far are partial. Voices are designed with different default styles, from calm to energetic. Some models expose an intensity control; Chatterbox, for example, offers an emotion exaggeration setting. Research on direct speech-to-speech translation tries to transfer the original speaker's prosody into the new language, but that approach is not yet a standard part of production dubbing. Choosing a voice whose overall style matches the video does more than any single trick.
Multiple speakers and overlapping voices
To give each person their own voice, a system first has to work out who is speaking when, a task called speaker diarization. Diarization struggles with short turns, similar-sounding voices and people talking over each other, and every mistake hands a line to the wrong voice. Overlapping speech is harder still, because recognition tends to capture only the dominant speaker.
Some tools assign a voice per detected speaker; others, mydubly included, use one voice for the whole video. A single voice avoids misattributed lines but makes conversations harder to follow, which is why subtitles often serve interviews and panels better. Translating videos with multiple speakers goes through the options.
Names, brands and jargon
Proper nouns fail at every stage. Recognition spells an unfamiliar name the way it sounds. Translation may then translate it, turning a surname like Rose or a product called Notion into an ordinary word. Finally the voice pronounces whatever spelling arrives using the target language's rules, so an English brand name can come out sounding like a native word.
The usual engineering answers are glossaries or do-not-translate lists in translation systems, vocabulary hints for recognition, and pronunciation dictionaries for synthesis. Where a tool exposes none of these, the practical answer is review: search the translated transcript for every name and term before trusting the dub. Translating names and technical terms covers building a glossary.
Music beds and sound effects
In most finished videos, speech is mixed with music and effects into a single track. A dub needs to replace the voice and keep everything else, which requires either the original stems or source separation, a family of models that estimate the vocal part of a mix and the rest. Separation has improved a lot but can leave watery artifacts, ghost syllables of the original voice, or music that pumps where speech was removed. Music also unsettles recognition, sometimes producing invented text during instrumental passages.
mydubly does this separation automatically: AI vocal separation removes the original speech, and the remaining music and effects are mixed under the translated voice and lowered while it speaks. The limits above still apply. In dense or loud mixes, faint traces of the original voice can remain and the background can sound slightly thinner than the original, and sung vocals are removed along with the speech. If you have the edit project, dubbing a dialogue-only export and mixing the returned translated voice with your music stem in an editor still gives the cleanest result; keeping background music when translating video shows how.
Cultural fit
A literally correct translation can still feel foreign. Idioms, jokes, levels of formality, measurement units, and references to local holidays or celebrities all need adapting, and some languages force choices English avoids, such as formal or informal address. Machine translation leans literal, and a fluent voice makes a literal line sound deliberate.
Engineering helps at the margins, with context-aware translation and formality settings in some engines, but adaptation is still mostly human work. For content where tone carries the message, plan a review by someone from the target market. Cultural adaptation in video covers what to look for.
What engineering still cannot fix
- Longer translations
- Response: elastic placement and capped speed-up. Remaining gap: very dense speech still sounds brisk.
- Flat emotion
- Response: voice styles and intensity controls. Remaining gap: sarcasm, comedy and grief rarely land.
- Several speakers
- Response: diarization and per-speaker voices. Remaining gap: crosstalk and misattributed turns.
- Names and terms
- Response: glossaries and pronunciation lexicons. Remaining gap: anything nobody listed.
- Music beds
- Response: stems or source separation. Remaining gap: separation artifacts when no stems exist.
- Cultural fit
- Response: context-aware translation. Remaining gap: humor and local references need a person.
Some limits sit outside the audio entirely. Lips on camera will not match the new language unless a separate model re-renders the mouth, which brings its own artifacts and ethical questions covered in AI lip sync. Text that appears in the picture, such as slides and captions burned into the footage, stays in the original language.
How mydubly handles these challenges
mydubly's pipeline makes specific choices against each problem. Audio is split into windows of about 30 seconds, with each cut placed at the quietest 50 ms frame in the last 6 seconds before the mark, so words and sentences are not sliced in half. Recognition runs on Whisper large-v3-turbo, translation on a neural machine translation engine, and the voice on Chatterbox Multilingual, with short fragments merged into whole sentences of up to about 220 characters before synthesis.
For timing, each line may start up to 0.3 s early or run up to 0.6 s late; speed-ups up to about 1.08× are treated as inaudible, tempo rises to 1.15× by default beyond that, and short lines are slowed slightly but never below 0.9×. Clips are loudness-normalized and joined, then mixed over the separated original background, which is lowered while the voice speaks, into one mono AAC track that is muxed with your original video in the browser without re-encoding the picture. What it deliberately does not do: lip-sync, multiple voices, or translating on-screen text.
You can reduce most of these challenges before uploading:
- Prefer videos with one main speaker and little crosstalk.
- If you edit the video yourself and the music bed is loud, consider exporting a dialogue-only version for dubbing, since music under speech makes recognition and separation harder. The translated audio then comes back as essentially the voice alone, and you mix your music stem under it afterwards.
- Leave short pauses between sentences when recording, which gives the timing step room to work.
- Keep a list of names and terms so you can check them in the translated transcript.
- Dub a two-minute test first, which costs 100 credits, before committing a long video.
Where to go next
If your video is a good fit, with one speaker, clear audio and content where meaning matters more than performance, try it with mydubly's AI dubbing. Before publishing, run through the AI dubbing quality checklist so each challenge above gets a deliberate check rather than a hopeful one.
Frequently asked questions
What is the hardest part of AI dubbing to get right?
Timing and emotion compete with each other. Speeding a voice up to fit a longer translation makes it sound rushed, and leaving it at natural speed makes it drift away from the picture. Expressive delivery is the part current systems handle least reliably.
Why do some translated lines sound rushed in a dub?
Because the translation took longer to say than the original line, and the system sped the voice up to keep it in sync. Small speed-ups are inaudible, but dense passages can push tempo toward the cap. Slower, clearer source speech leaves more room.
Can AI dubbing separate the voice from background music?
Some tools use source separation models to estimate the vocal part of a mix, with mixed results when no clean stems exist. mydubly separates automatically: it removes the original speech and keeps the music and effects under the translated voice. Faint traces of the original voice can remain in dense mixes and sung vocals are removed too, so if you have the edit project, dubbing a dialogue-only export and mixing the translated voice with your music stem in an editor still gives the cleanest result.
Do these challenges affect audio-only content like podcasts too?
Mostly, yes. Timing matters less without a picture, but multiple speakers, names, music and cultural fit all apply. For podcasts, the audio translator produces the translated audio and transcripts without the video step.
Will AI dubbing get past these challenges soon?
Some gaps are narrowing, especially naturalness and timing. Others, like humor, cultural references and overlapping speakers, depend on understanding intent and context, which is why hybrid workflows with human review remain common for important content.