AI video translation

Background Music in Translated Videos: What Stays and How to Perfect It

When you translate a video with mydubly, your background music and sound effects stay in the result. AI vocal separation removes the original speech from the audio, and the translated voice is mixed over what remains, with the music dipping automatically while the new voice speaks. Separation has limits, though, and if you still have your project files, an optional route gives the cleanest result: dub a dialogue-only export, then mix the translated voice with your own music stem. Below is how the automatic mix works, where it falls short, and the dialogue-only workflow for when you want full control.

9 min read · Updated

How mydubly keeps your music and effects

An AI dubbing pipeline turns the original dialogue into text with speech recognition, translates the text, and has a synthetic voice read the translation. On its own, that chain produces a speech-only track, so the guitar under your intro or the whoosh on your title card would be lost.

mydubly adds one more step. AI vocal separation splits the uploaded audio into the original speech and everything else: music, ambience and sound effects. The original speech is discarded, the translated voice is mixed over the remaining background, and the background is automatically lowered (ducked) while the new voice speaks and comes back up in the gaps. The translated MP4 carries that mix with your original picture, and the translated audio download (M4A) is the same mix on its own. The output audio is mono.

The original speaker is not heard underneath the translation, so this is still dubbing rather than a voice-over, where the original voice stays audible at a lower level. The difference is covered in dubbing vs voice-over.

Why separating voice from music is hard

Once a video is exported, dialogue, music and effects are summed into one waveform per channel. There is no layer left to switch off. Getting the music back means estimating, moment by moment, which part of the signal belongs to which source. mydubly does this automatically for every dubbed video, but the task itself explains its limits.

Source separation models learn this from large collections of isolated stems. On a clean song the results can be impressive. Dialogue over a soundtrack is a tougher case:

  • Voices and many instruments overlap in frequency, so the model has to guess where one ends and the other begins.
  • Reverb belongs to the voice but smears across time, and often stays behind as a faint wash in the music output.
  • Sound effects are neither speech nor music, so they may land in either output or be split between both.
  • Artifacts are most audible where you need the bed to be cleanest: quiet passages directly under dialogue.

That last point matters more for translation than for anything else. If faint original-language speech survives in a separated music bed, viewers hear two languages at once, which undercuts the whole point of the dub.

What to expect from the automatic mix

For most videos with a music bed, ambience or sound design, the translated version sounds like your video in another language. Know the limits before you publish:

  • Separation is not perfect. In dense or loud mixes, faint traces of the original voice can remain, and the background can sound slightly thinner than in the original mix.
  • Singing counts as voice. Sung vocals and lyrics in songs are removed along with the speech, so a song keeps its instruments but loses its vocals.
  • The output audio is mono, so a wide stereo music bed is folded down.
  • Effects tied to on-screen action, like a door or a click, stay in sync. Effects that reacted to the original speech, such as a sting after a punchline, may no longer land exactly where the translated joke ends.

Listen to the quietest passages directly under dialogue first; that is where any leftover original speech is easiest to hear. Music under speech also makes recognition harder, so a loud bed can cost both transcription accuracy and separation quality.

For the best result: dub a dialogue-only export

If you edited the video yourself, you almost certainly still have the dialogue, the music and the effects on separate tracks. Broadcasters ask for exactly this split when they commission international versions: an M&E (music and effects) track, which is the full mix minus dialogue, so a new language can be laid on top. Your editing project already contains everything you need; you only need to export it.

This route is optional. The automatic mix is usable as it is; working from your own tracks simply gives you the fullest, fully controllable result, such as stereo music and your own levels. Don't layer a music or M&E stem over the downloaded audio of a normal dub, though: that file already contains the separated original background, so the music would play twice, and any faint original speech left in it would stay.

  1. Open the original project, mute every music and effects track, and export a copy of the video with only the dialogue audible, running the full length and starting at frame zero.
  2. Mute the dialogue tracks instead and export what remains, music and effects together, as one audio file of the same length.
  3. Store the M&E export next to the master file. It serves every language you translate into.

Mixing the translated voice with your music, step by step

  1. Run the dialogue-only copy through the AI dubbing tool and download the translated audio file. With no music in the upload, it comes back as essentially the translated voice alone.
  2. In your editor, create a project with the original video and mute its audio.
  3. Place the translated audio at the very start of the timeline. It is built to the full length of the video, so it lines up without nudging.
  4. Add the M&E export on another track, also from the start.
  5. Duck the music under speech, either with your editor's automatic ducking keyed to the translated audio or with volume keyframes around each line.
  6. Listen to the busiest and quietest passages, then normalize the finished mix to the loudness your publishing platform recommends.
  7. Export, and attach the translated SRT or VTT file separately if you want subtitles as well.

A practical way to set the ducking level is to lower the music until you can follow every translated word with your eyes closed, then let it swell a little in the gaps between lines. The translated voice is loudness-normalized across the whole video, so it sits at a steady level; the music setting that suited the original speaker often needs to come down slightly further.

When you only have the finished mix

Sometimes the project files are gone: an old webinar, an export from a former agency, a video you only kept as an upload. In that case the automatic separation is already doing the job a separation tool would, so the translated MP4 is usually the version to publish. If you are not happy with how the background sounds:

  • Check the quietest passages under dialogue. If faint original speech is only noticeable on headphones, most viewers will not hear it.
  • If the original language is clearly audible, captions over the original audio may serve that video better. The subtitle generator produces SRT and VTT files and leaves your mix exactly as it was.
  • For future videos, keep a dialogue-only export and a music-and-effects export of every project so the dialogue-only route is always open.
Worked example

Suppose a 12-minute product walkthrough has a light music bed throughout. Dubbing it into German costs 12 × 50 = 600 credits, or 60¢. Dubbing the finished video gives a translated MP4 with the music already under the German voice, dipping during each line. The editor, who still has the project, wants the stereo bed for the company website, so she dubs a dialogue-only copy instead, for the same 600 credits, lays her M&E stem under the returned German voice with automatic ducking and exports. Had the project been lost, dubbing the finished video and publishing the automatic mix would have been the way to go.

Trade-offs and when the dialogue-only route isn't worth it

  • Every language adds an editing pass. Five languages means five mixes to check by ear.
  • If the automatic mix already sounds right, the extra exports and mixing add work without a visible gain.
  • Only the dialogue-only route gives a clean custom mix: layering a music stem over the downloaded audio of a normal dub plays the music twice and does not remove any faint original speech left in it.
  • Music licenses sometimes cover a specific use, territory or number of versions, so check before publishing new language versions.

How mydubly fits into a music-preserving workflow

mydubly does the separation and the mix for you, and if you dub a dialogue-only export it gives you the translated voice for your own mix:

  • The translated MP4 contains your original video stream, not re-encoded, with the translated voice mixed over your original music and effects.
  • The translated audio file is that same mix on its own, matching the length of the video, so it drops straight onto an editor timeline.
  • Each translated line is placed elastically near where the original line was spoken, with only small tempo changes, so music cues written around the original speech still roughly fit.
  • Every full translation also includes SRT and VTT subtitles in the target language, and transcripts in both languages, which are handy for checking timing against the music.
  • One stock voice reads every line; there is no voice cloning, no lip sync, and the original speakers are not heard. There is also no way to upload a music file for mydubly to mix in.

The video file itself stays on your device and only its audio is sent for processing, so any custom mix is a local job in your own editor. Per-minute pricing for longer projects is laid out in the video translation cost guide.

Where to go next

Dub the video with mydubly's AI dubbing and listen to the result, especially the quiet passages under dialogue. If it sounds right, publish it. If you want stereo music or your own levels and you still have the project, dub a dialogue-only export and follow the steps above. For how the voice is fitted to the original timing, see syncing translated audio with video, and to make separation easier on future shoots, read making videos translation-ready.

Frequently asked questions

Can I get a translated MP4 that still has my original music in it?

Yes. AI vocal separation removes the original speech, and the translated voice is mixed over your original music and effects, which dip automatically while the new voice speaks. If you still have the project, dubbing a dialogue-only export and mixing the translated voice with your own music stem in an editor is an optional route for the cleanest result.

Is the downloaded translated audio file only the voice?

No. It is the same mix as the translated MP4's audio: the translated voice over your original music and effects, in mono. It runs the full length of the video from the first frame, so it lines up on an editor timeline without nudging. For a voice-only track, dub a dialogue-only export of your video.

Can traces of the original speech remain under the translation?

In dense or loud mixes, faint traces of the original voice can remain, especially reverb tails and quiet syllables. Listen to quiet passages under dialogue before publishing; if the original language is clearly audible, translated subtitles over the original audio may suit that video better.

What happens to songs with vocals?

Sung vocals count as voice, so the lyrics are removed along with the speech. The song's instruments stay in the mix.

How loud should music sit under a translated AI voice?

There is no single correct number, because it depends on the music and the platform. Set it by ear so every word is easy to follow, use ducking so the bed rises between lines, and finish by normalizing the mix to your platform's published loudness guidance.

If I change the music later, do I have to dub the video again?

It depends on what you dubbed. If you dubbed a dialogue-only export, you can mix the translated voice with a new music file as often as you like. The audio from a normal dub already contains the original background, so a new track would play on top of it rather than replace it; to swap the music cleanly, dub a dialogue-only export once. Changes to the spoken content always require a new translation.