Three ways to replace speech in another language
When people say "dub" or "voice-over" loosely, they can mean very different products. The industry distinguishes at least three formats, and knowing which one you need saves confusion when comparing quotes or tools.
- Lip-sync dubbing
- The original voices are removed and replaced with actors speaking the target language, adapted to match mouth movements. Common for feature films, series and animation in many markets.
- UN-style voice-over
- The original speaker is heard for a moment, then lowered, and a translated voice speaks over the top, finishing slightly after. Common in news, documentaries and interviews.
- Lectoring
- A single reader, often one voice for every character, reads a translation over the lowered original soundtrack. Widely associated with Polish television, where the "lektor" is a familiar feature.
- Narration
- An off-screen voice that describes, explains or introduces, written for the target audience rather than translating any on-screen speaker line for line.
Lip-sync dubbing: full replacement
In lip-sync dubbing, the viewer should forget that the film was ever in another language. The translation is adapted so that open-mouth vowels and closed-lip consonants such as b, m and p fall roughly where the actor's mouth does, and line lengths match the on-screen speech. Actors perform against the picture, and the new dialogue is mixed with the original music and effects track, which productions usually deliver separately for this purpose.
It is the most immersive format and also the most demanding. Adaptation can change the wording substantially, casting is per character, and the work only succeeds if the music and effects stems are available.
UN-style voice-over: two voices at once
UN-style voice-over takes its name from how speeches are often presented in international broadcasting. The original speaker is audible at full level for a second or two, establishing who is talking and how they sound, then the original is ducked to a low level and a translated voice reads over it. The translation usually ends a beat after the original so the speaker's voice can come back briefly.
This format signals authenticity: the audience can hear the real person, their tone and their emotion, while understanding the words through the translation. It needs no lip-matching and much less adaptation, which is why it is the norm for interviews in news and documentaries.
Narration and off-screen voice
Narration does not replace a speaker at all. A documentary narrator, a product video voice or an explainer's off-screen guide speaks directly to the audience. When such a video is localized, the narration script is translated and re-recorded, and because nobody is seen speaking it, there are no lips to match and timing is loose. Narration-led videos are therefore among the easiest to localize with any method.
Strengths and trade-offs of each format
Every format buys one quality at the cost of another.
- Lip-sync dubbing is immersive but expensive, slow, and dependent on adaptation that can drift from the literal meaning.
- UN-style voice-over keeps authenticity but makes viewers process two voices, and overlapping audio can be tiring over long runs.
- Lectoring is cheap and fast but feels distant to audiences not used to it, since one flat voice reads every role.
- Narration is easy to localize but only applies when the original was narrated in the first place.
- Subtitles, the alternative to all of these, keep the original voice untouched but require reading; the subtitles vs dubbing guide weighs that choice.
What mydubly produces
mydubly's output does not fit neatly into one traditional box, so it is worth stating exactly. The translated MP4 carries a replacement voice track: a synthetic voice reads the translation, and each line is placed where the original line was spoken, starting up to 0.3 seconds early or running up to 0.6 seconds late, with gentle speed adjustment when needed. The original voice is removed with AI vocal separation, while the original music and effects stay underneath, lowered automatically while the new voice speaks.
That makes it like dubbing in that only the new language is heard, but without lip-sync, since the picture is never altered. It differs from UN-style voice-over because, although the music and effects remain, the original voice is not audible underneath. And one chosen voice reads every speaker, like a lector, so it suits narration and single-presenter content best. Alongside the MP4, you also get the translated audio as a separate file, SRT and VTT subtitles in the target language, and transcripts in both languages, which is what makes the do-it-yourself variations below possible. The voices and outputs are listed on the AI dubbing page.
Suppose you have a 6-minute interview with a German engineer and want an English version that keeps her presence. mydubly's translated MP4 gives English speech over the original background sound, with her own voice removed, which works for viewers who only want the content. For a documentary-style feel, you would instead use the translated audio file and build a UN-style mix in your editor. The full output for 6 minutes costs 300 credits ($0.30), and both versions come from the same run.
Building a UN-style voice-over from mydubly's files
If you want the original voice audible underneath, you can assemble it yourself in any video editor that supports multiple audio tracks.
- Run the video with full translation output on and download the translated audio file along with the MP4.
- In your editor, place the original video with its own audio on the timeline.
- Add the translated audio on a second track and shift it later by around half a second to a second, so the original speaker is heard first.
- Lower the original track to a quiet bed under the translated voice, raising it briefly at the start of each answer if your editor allows keyframes. The translated audio already carries the original music and effects, so the background is present on both tracks; keeping the original track low stops that doubling from getting loud or muddy.
- Check the ending, since delaying the voice track pushes the last line slightly past the original.
- Export and listen on headphones and on a phone speaker to make sure the translated voice stays clearly on top.
Choosing a format for your video
Match the format to the content. Narrated explainers, tutorials and lectures suit mydubly's replacement track as it comes. Interviews and documentaries often benefit from the UN-style mix described above or from subtitles. Character-driven fiction and close-up advertising still call for lip-sync dubbing with human actors; AI lip sync explains the visual alternative and its caveats. When you are ready to try the replacement-track approach, start with a short clip on AI dubbing.
Frequently asked questions
Is voice-over cheaper than dubbing?
Generally yes, because voice-over needs no lip-matched adaptation and often uses fewer voices. Lip-sync dubbing involves adapters, casting per character, direction and a full mix. Exact prices depend on the studio, language and length, so compare quotes for your specific project.
Why do documentaries use voice-over instead of dubbing?
Hearing the real person, even faintly, signals that the interview is genuine and preserves their emotion. Documentaries also contain many short interview clips with different people, which would be expensive to cast and lip-match. Voice-over gives a credible, efficient result.
Does mydubly keep the original voice quietly in the background?
No. The original voice is removed with AI vocal separation; the original music and effects stay underneath the AI voice, but the original speech is not heard. You can recreate a voice-over mix yourself by layering the downloadable translated audio file over the original in a video editor, as described above; keep the original low, since the background is present in both.
What is a lektor in voice-over?
A lektor is the single reader used in a style of voice-over common on Polish television, who reads the translation of all characters' lines over the lowered original soundtrack. The delivery is usually deliberately neutral. Audiences raised on it often prefer it, while others find it distancing.
Can I add my own music under a mydubly voice track?
Yes, but not on top of a normal dub: the translated audio already keeps the original background, so added music would double it. If you have the edit project, export a dialogue-only version and dub that; the result is essentially the translated voice alone, which you can mix with your own music stem in an editor for the cleanest, fully controllable result, keeping the voice clearly on top. Our article on keeping background music when translating video covers practical mixing workflows.