How AI dubbing works
Behind the single 'translate' button are five steps:
- Speech recognition turns the original audio into text with timings.
- Machine translation converts each segment into the target language.
- Text-to-speech reads each translated segment in a synthetic voice.
- Timing fit places each voiced segment where the original was spoken, speeding it up slightly when the translation is longer.
- Vocal separation removes the original speech, and the new voice is mixed over the remaining music and effects, which dip automatically while it speaks.
In mydubly, timing fit speeds speech up by at most 1.15× and keeps the pitch natural, so voices don't sound chipmunked. The finished audio is then combined with your original picture on your own device.
What AI dubbing is good at
- Explainers, tutorials, lectures, training and other single-narrator content.
- Making long videos watchable for people who won't read subtitles for 30 minutes.
- Testing a new market cheaply before paying for professional voice-over.
Where it falls short
- Several speakers share one voice — conversations lose who's speaking.
- Emotion and comic timing are flatter than a human actor's.
- There's no lip-sync; on-camera close-ups look like a traditional voice-over.
- Music and effects are kept, but separation isn't perfect: faint traces of the original voice can remain in dense mixes, and sung lyrics are removed with the speech.
If those matter more than speed and cost, subtitles may be the better choice — or use AI dubbing as a draft for human voice actors.
How to dub a video, step by step
- Pick a video with one main speaker and clear audio.
- Open the AI dubbing tool, drop in the file and choose the target language.
- Preview the 8 voices and pick one that matches the tone — a narrator voice for explainers, energetic for promos.
- Switch Full translation output on and start the job.
- Download the dubbed MP4 and the translated transcript, and spot-check a few minutes.
Dubbing is 50 credits per minute with a 2-minute minimum: a 5-minute video costs 250 credits (25¢), a 30-minute lecture 1,500 credits ($1.50).
Tips for a better dub
- Speak clearly and avoid talking over music in future recordings meant for dubbing — it hurts recognition and makes separating your voice from the music harder.
- Shorter sentences translate and time better than long, nested ones.
- Your music and effects stay automatically; for the cleanest result, dub a dialogue-only export of your edit project and mix the translated voice with your clean music stem in an editor.
Pre-flight: is this video a good dubbing candidate?
Five minutes with the source file predicts most of what you will think of the dub. Play it through once and answer these questions before spending credits:
- One main voice?
- Yes: dub it. If several people talk in turn, one stock voice will read all of them, which feels like a narrated documentary rather than a conversation.
- Speech over music?
- The original speech is separated out and the music stays under the translated voice, lowered while it speaks. Loud music under speech makes recognition and separation harder, and songs lose their sung vocals.
- Speaker in close-up on camera?
- There is no lip-sync, so mouths won't match the words. Wide shots, slides and screen recordings hide this; tight face shots expose it.
- Fast, dense delivery?
- When a translation runs longer than the original, lines are shortened to fit the timing, so rapid talkers lose more detail than relaxed ones.
- Jokes, slogans, wordplay?
- Expect literal results. Flag those lines for a human rewrite or keep them as subtitles only.
Choosing one of the eight voices
- Match the role rather than the person. One voice reads the entire dubbed track, so pick what suits the narration instead of chasing the original speaker's sound.
- Preview candidates with the content's register in mind: a calm, even voice for compliance or medical explainers, a brighter one for a product teaser.
- Keep the same voice across a series so returning viewers hear continuity, and write down which one you used so colleagues pick the same one.
- If a familiar presenter is on screen, consider a line in the description saying the translation uses an AI voice; a stock voice coming out of a known face can surprise people.
More on matching voice to content in choosing an AI voice.
Reviewing the dub in a single pass
- Play the preview at 1.5× from start to finish and note the timecode of every gap, abrupt cut or rushed passage.
- Go back to each note at 0.75× and read the original and translated text side by side on the result screen.
- Open the downloaded SRT next to the video. It reflects what the voice says, so comparing it with the full translated text shows where lines were shortened.
- Ask a speaker of the target language to watch the first two minutes plus any section containing numbers, prices or instructions.
- Finally, play the translated audio file on its own, without the picture. Gaps, sudden level changes and clipped word endings are easier to hear when nothing on screen pulls your attention.
The AI dubbing quality checklist turns these steps into a sign-off sheet you can reuse.
Symptoms after download, and what to do
- Faint original voice, or the music sounds thinner
- Vocal separation isn't perfect in dense or loud mixes. If you have your edit project, dub a dialogue-only export and mix the translated voice with your clean music stem in an editor, as described in background music in translated videos.
- The file is an MKV
- Your source codec can't be stored in an MP4, so the browser merged into MKV. Re-export from an editor if the destination requires MP4.
- A name is mispronounced every time
- Repairs happen after download: record a short pickup of the correct pronunciation over those words in your editor, or let the subtitles carry the correct spelling.
- One key sentence lost its meaning
- Check the original-language transcript first. If recognition misheard a word, every later step inherited the mistake.
- Long passages sound crammed
- The translated text ran longer than the source. For future recordings, shorter sentences and small pauses give the voice room; see text expansion in translation.
Worked example: a 7-minute onboarding video into French
An HR team has a 7-minute onboarding video: one presenter at mid-distance from the camera, quiet music underneath, and slides with headings. Dub mode into French costs 350 credits (35¢). They download the translated video, the M4A and the French SRT. The quiet music stays under the French voice, so in their editor they only rebuild the slide headings in French, because text in the picture isn't translated. A 90-second welcome clip from the same series is billed at the 2-minute minimum: 100 credits (10¢). The French video translator page covers language-specific checks, and training video translation shows how this fits a wider course rollout.
Frequently asked questions
Is AI dubbing the same as voice cloning?
No. Voice cloning recreates a specific person's voice. mydubly uses stock voices and doesn't clone anyone's voice.
How long does dubbing take?
It depends on video length and how busy the service is; audio is processed in 30-second chunks, and progress is shown as it goes.
Can each speaker get a different voice?
No. One stock voice reads the whole dubbed track. For interviews and panels, translated subtitles keep the speakers distinct while the original voices stay audible.
Can I download the dubbed audio without the video?
Yes. The translated audio download is an M4A file, which is what you need for editing in a video editor or for platforms that accept an extra audio track. It already includes the original background, so to mix in your own music, dub a dialogue-only export first.
Do the subtitles from a dub match what the voice says?
Yes. In dub mode the SRT and VTT follow the spoken translation, including lines that were shortened to fit the timing, so viewers who turn them on read what they hear.