What you start with and what you are building
A mydubly dub gives you two files: a translated MP4 and the same audio as a separate M4A. That audio is already a mix: AI vocal separation removes the original speech, and the original music, effects and room tone stay under the AI voice, lowered automatically while it speaks. Separation is not perfect, though. In dense or loud mixes, faint traces of the original voice can remain and the background can sound slightly thinner, and sung vocals are removed along with the speech. Keeping background music in translated videos explains those limits. If you want your own music levels, stereo music or a different track, the cleanest route is to export a dialogue-only version from your project and dub that, so the translated audio comes back as essentially the voice alone, then build the music under it here. Laying a music stem under the audio of a normal dub would play the music twice. This article picks up from there: you have the translated audio, a music and effects stem if you dubbed a dialogue-only export, and an editor.
The finished product is a mix in which the translated voice sits clearly on top, the music breathes in the gaps, and nothing sounds like it was pasted in. The steps are the same in every editor; only the button names differ.
Setting up the project so nothing drifts
Most sync trouble in editors comes from project setup rather than the audio itself.
- Use the translated MP4 as the picture, not your original export. Its video frames are identical to the original, because mydubly never re-encodes the picture, and its timeline already carries the voice in the right place.
- Match the project frame rate to the clip. Let the editor create the sequence from the clip rather than choosing a preset.
- Watch for variable frame rate footage from phones and screen recorders. Some editors drift with it over long timelines; if the voice slowly slides out of sync in the editor but not in a media player, read variable frame rate explained and convert the picture to a constant rate first.
- Place every audio file at the very start of the timeline, at frame zero. The translated audio and an M&E stem are both laid out on the original timeline, so they line up without nudging.
- Keep the editor at 48 kHz audio, the usual video standard; editors resample imported audio automatically.
A track layout that stays manageable
Give each kind of sound its own track and name it. A simple layout for a dub:
- Track A1
- Translated audio: the AI voice, over the kept background if you dubbed the full mix. The track you will fix and protect.
- Track A2
- Music bed, from your stem or a replacement library track. Only when A1 came from a dialogue-only export; otherwise the music doubles.
- Track A3
- Sound effects, if you have them separately.
- Track A4
- Spare track for replacement lines or pickups.
Mute the embedded audio of the video clip, or unlink and delete it, so the voice is not doubled. If you work on several languages, duplicate the finished sequence for each one and swap only A1; the music, effects and ducking choices carry over, although each language needs its own listening pass because the speech falls in different places.
Ducking music under the voice, editor by editor
Ducking lowers the music while someone speaks and lets it rise in the gaps. If you dubbed the full mix, mydubly has already ducked the kept background; this section is for music you add under a voice from a dialogue-only export. You can automate it or draw it by hand.
- Premiere Pro: in the Essential Sound panel, tag the voice clips as Dialogue and the music as Music, then enable ducking on the music and generate keyframes. Adjust the duck amount and fade length, and tidy individual keyframes where the music dips too late.
- DaVinci Resolve: on the Fairlight page, sidechain the music track's compressor or ducker from the dialogue track, so speech drives the reduction. Resolve's free version includes the Fairlight page.
- Final Cut Pro: select a range of the music clip over each line and lower its volume, or add volume keyframes; with many lines, apply the change to one clip and copy its attributes.
- Audacity: import the voice and music on separate tracks, select the music, and use the Auto Duck effect with the voice track as the control, then export a mixed audio file for your video editor.
- Shotcut and Kdenlive: add volume or gain keyframes on the music track around each line.
Menus and feature names change between versions, so check your editor's current documentation if a panel is not where described. Whatever tool you use, judge the duck by ear: lower the music until every translated word is easy to follow, then check that the music does not pump up and down distractingly in short gaps between lines.
Fixing individual lines
Most dubs need only a handful of repairs. Common ones and what to do about each:
- A line starts too early over a cut or a sound effect: select the clip region on A1, blade it at the silence before the line, and slip it a few frames later. Check that it doesn't now collide with the next line.
- A mispronounced name or acronym: replacement options include a human pickup recorded on A4, a short re-run of that section, or covering the moment with music and a subtitle. The causes are explained in why AI voices mispronounce names.
- A click, breath or odd artifact between words: cut it out and add a short crossfade of a few milliseconds on either side so the edit is silent.
- A line that sounds rushed: a gentle time-stretch with pitch preserved can add a little room, but more than a small amount sounds processed. It is often better to accept the pace or trim a pause elsewhere.
- A wrong word that changes the meaning: this cannot be fixed by editing audio. Replace the line, or correct the problem upstream and re-run.
To re-run a short section, cut the original video at clean pauses around the problem line, dub just that part with the same voice, and drop the new clip over the old one. Note the exact start time of the section you cut, because that is where the new clip goes on the timeline. Dubbing has a 2-minute minimum per file, so a short section costs 100 credits (10¢). The new take is synthesized fresh, so its delivery can differ slightly from the surrounding lines; listen across the joins.
Suppose a cooking channel exports a dialogue-only copy of a 7-minute video and dubs it into French for 350 credits (35¢). In Resolve, the editor drops the translated MP4 on a new timeline, mutes its embedded audio, adds the translated M4A, essentially the French voice alone, on A1 and the music stem from the original project on A2, and sidechains the music from the voice. Two fixes remain: one line lands on a sizzling-pan sound effect and is slipped back by eight frames, and the brand name of a spice blend is read oddly, so the editor records the French name on A4 with a friend and crossfades it in. The export is audio only, then remuxed with the original picture.
Exporting without re-encoding the picture
When an editor exports a video, it usually re-encodes the picture, which costs time and some quality. Since only the audio changed, a cleaner path is to export only the mixed audio, as WAV or high-bitrate AAC at 48 kHz, and then put it back with the untouched picture using a muxing tool:
ffmpeg -i translated.mp4 -i final-mix.wav -map 0:v:0 -map 1:a:0 -c:v copy -c:a aac -b:a 192k dubbed-final.mp4
This copies the video stream bit for bit and encodes only the new audio. If you need both languages in one file, add the original audio as a second track using the method in making a video with two audio tracks. If you would rather export from the editor, choose a high-quality setting or a passthrough option where your editor offers one.
Before the final export, check loudness. The AI voice is loudness-normalized, but adding a music stem raises the overall level, so measure the integrated loudness of the finished mix and adjust it toward your platform's published target. Most editors include a loudness meter, and platforms publish their own recommendations, which change from time to time.
Steps from download to finished dub
- Download the translated MP4 and the translated M4A as soon as the job finishes.
- Create a timeline from the translated MP4 and mute its embedded audio.
- Place the translated audio on A1 and, if you dubbed a dialogue-only export, your music and effects stems on A2 and A3, all at frame zero.
- Set up ducking, automatic or by keyframes, and listen through once at normal speed.
- Note every problem line with its timecode, then fix them in order of how noticeable they are.
- Check loudness on the whole mix.
- Export audio only and remux with the picture, or export from the editor at high quality.
- Watch the final file from start to finish on the device type your audience uses.
Limits of editing a dub after the fact
- An editor can move, trim and stretch lines, but it cannot change the words, the voice or the emotional delivery.
- One AI voice reads the whole video; giving different speakers different voices means separate runs and careful editing, as described in translating videos with multiple speakers.
- Replacement lines from a human or a fresh run rarely match perfectly in tone. Fewer, well-placed fixes sound better than many.
- Every language multiplies the work. Five languages means five mixes, each needing a full listening pass.
- There is no lip-sync. Editing can improve timing around cuts, but mouths on close-up shots will not match the new language.
Where mydubly fits and where your editor takes over
mydubly handles recognition, translation, the voice and the timing fit: each translated line is placed near where the original was spoken, sped up gently if needed, and loudness-normalized. It also keeps the original music and effects under the voice, lowered automatically while it speaks. It hands you the MP4 and the same mix as a separate M4A, with the picture never re-encoded and the video file never leaving your device. Everything after that, from your own music levels and line replacement to mastering and multi-track export, happens in your editor. mydubly has no manual mixing controls, no way to upload music for it to mix in, and no per-line regeneration.
Next step
Dub a short video with AI dubbing, set up the four-track layout above, and time how long the mix and fixes take for one minute of video. That number tells you how much editing to budget when you scale up to longer videos or more languages. If you have not produced a dub before, the overall process is laid out in how to dub a video with AI.
Frequently asked questions
Which free editor is easiest for finishing a dub?
DaVinci Resolve's free version has a full audio page with sidechain ducking, but it has a learning curve. Shotcut and Kdenlive are lighter and handle keyframed music levels well. If you only need to mix voice and music, Audacity's Auto Duck effect is quick, and you can remux the mixed audio with the picture afterward.
Can I keep the original voice faintly underneath, like a voice-over?
Yes, in your editor. Put the original audio on its own track at a much lower level under the translated voice, which is the voice-over style common in news and documentaries. The translated audio already carries the original music and effects, so the background is present on both tracks; keeping the original track low stops that doubling from getting loud or muddy. mydubly's output does not keep the original voice itself. The differences between this style and dubbing are covered in dubbing vs voice-over.
How do I change the speed of the AI voice in the editor?
Use your editor's speed or time-stretch control with pitch preservation on. Small changes are hard to hear; large ones sound processed or chipmunk-like. mydubly's timing fit already speeds lines up gently where needed, so extra stretching is best reserved for single lines that collide with a cut or an effect.
Can I use the edited mix as a YouTube dubbed audio track?
Yes. Export the finished mix as an audio file the length of the video and upload it as a language track where YouTube's multi-language audio feature is available to your channel. See YouTube multi-language audio for how those tracks work and where to check eligibility.
Will my edits change the subtitles?
No. Subtitle files are separate. If you slip a line by a few frames, the matching subtitle cue stays where it was, which is usually fine for small moves. If you replace a line with different wording, edit the corresponding cue in the SRT so the subtitles and voice say the same thing.