Troubleshooting

Hearing Two Voices at Once: Fixing Doubled Audio in a Video

Hearing two voices or two soundtracks at once almost always means two sources were mixed when only one was wanted. The usual culprits are an original dialogue track left unmuted under a translated voice in an editor, a multi-track file whose tracks an editor plays together, a screen recorder capturing both system audio and a microphone that hears the speakers, or a voice added twice across successive exports. Check whether the extra audio is a separate track you can mute or already mixed into one track, because only the first is a quick fix.

8 min read · Updated

What doubled audio sounds like

The character of the doubling tells you a lot about where it came from.

  • A translated voice with the original speech faintly underneath. Usually an editing slip, though in voice-over it is deliberate, and in a dub that keeps the original background, faint traces can also be left by imperfect voice separation in dense or loud mixes.
  • The same voice twice, with a tiny delay. Sounds hollow, phasey or metallic when the delay is a few milliseconds, and like a distinct echo when it is longer.
  • Two different soundtracks. Music from the original video plus music you added, or two versions of the same mix slightly offset.
  • Your own voice echoing in a call recording. The other side's speakers fed your voice back into their microphone.
  • Game or app sound twice in a screen recording. Captured once digitally and once through the microphone from the speakers.

The usual causes

Original track left on
In an editor, the original dialogue stays attached to the video clip while the translated voice plays on another track
Multi-track source in an editor
Camera, screen recorder and dual-language files can hold several audio tracks, and many editors import and play all of them
System audio plus microphone
A screen recorder captures computer sound directly while the microphone also hears it from the speakers
Monitoring loops
A setting that plays the microphone back through the speakers, which the microphone then picks up again
Echo on calls
A participant on speakers without effective echo cancellation sends your voice back to you
Layered exports
A voice or music bed added in one export, then added again when that export is used as the source for the next

A media player normally plays one audio track at a time, so a file with two separate tracks sounds normal in a player and doubled only in software that plays every track. That is why the same file can seem fine and broken depending on where you open it. How players choose a track is covered in multiple audio tracks in video files.

Separate track or already mixed?

This is the key question, because it decides whether the fix takes seconds or is barely possible.

  1. Open the file in VLC and look at the audio track menu. If it lists more than one track, switch between them and listen to each on its own.
  2. If there is only one track and you still hear both sources, they are mixed together in that track.
  3. In your editor, solo each audio track in turn. If the doubling disappears when one track is muted, you have found the extra source.
  4. Check the clip's linked audio. Video clips often carry their own embedded audio, which may sit on a track you scrolled past.
  5. Look at the export or mixdown settings to see which tracks are included.

Fixes when the extra audio is a separate track

  • Mute or delete the original dialogue track under a translated voice. If the original audio is embedded in the video clip, unlink it first so you can mute it without removing the picture.
  • When a multi-track file is imported, disable the tracks you do not want in the clip or sequence settings, or in the editor's track list.
  • In a finished file with two tracks, remux it keeping only the track you want. With ffmpeg, this keeps the video and the second audio track, copying both without re-encoding:

ffmpeg -i input.mp4 -map 0:v:0 -map 0:a:1 -c copy single-track.mp4

Audio tracks are counted from zero, so a:0 is the first and a:1 the second. If you intended a file with two selectable language tracks, that is a different job, covered in making a video with two audio tracks.

Fixes when the sources are already mixed

Once two sources share one track, there is no clean way to remove one. The practical options, in order:

  1. Go back to the editing project or the original recordings and export again with only the source you want.
  2. Re-record if the material is short and the setup is simple, such as a narrated screen recording.
  3. Try a voice or music separation tool as a last resort. These can reduce one element, but they leave artefacts, and a voice separated from another voice is particularly hard.
  4. For short echoes in a single voice, echo reduction tools may help a little; they cannot remove a long, loud echo.

Preventing doubled audio in screen and call recordings

  • Choose deliberately: microphone only for narration, system audio only for capturing an app's sound, or both on separate tracks if the recorder allows it, so you can balance or mute each later.
  • Wear headphones while recording so the microphone cannot hear the speakers.
  • Turn off any setting that plays the microphone back through the speakers, and avoid loopback devices unless you need them.
  • For calls, ask everyone to use headphones and, for important interviews, to record their own microphone locally.

The guide to screen recording audio settings walks through these options on macOS and Windows.

A worked example: a dub with the original underneath

Course video, hypothetical

Suppose a course creator translates a 10-minute lesson into Spanish, downloads the translated audio as an M4A file, and brings the original video and the M4A into an editor to replace one mispronounced line. The export has the Spanish voice with the English narration clearly underneath. Soloing the tracks shows the English came from the video clip's embedded audio, which the editor placed on its own track. Unlinking the clip's audio and muting it fixes the export, and the final video has one voice over the original music.

The same slip happens in reverse when someone adds the translated M4A to a translated MP4 that already contains it, producing the same voice twice, sometimes slightly offset. Check what each file contains before layering it. The full editing workflow, including mixing music under a voice, is in editing dubbed audio in a video editor.

Pitfalls and limits

  • Deleting the wrong track. In multi-track files, check what each track contains before removing anything, and keep the original file.
  • Expecting separation tools to undo a mix cleanly. They reduce, they do not restore.
  • Confusing voice-over with a mistake. Some formats deliberately keep the original voice audible under the translation; if that is the goal, lower it consistently rather than leaving it at full level.
  • Fixing the export but not the template or preset that caused it, so the next project repeats the error.
  • Ignoring the source. If the original recording already has two voices mixed, such as an interpreter speaking over a speaker, any transcript or translation of it will contain both.

Where mydubly fits

A mydubly dub produces the translated audio as an M4A file and a translated video in which the dubbed track is the only audio: the picture is copied without re-encoding, AI vocal separation removes the original speech, and the original music and effects are kept under the translated voice. So the translated video itself contains one voice, although in dense or loud mixes faint traces of the original voice can remain. Clearly doubled audio appears when the files are combined again in an editor, as in the example above. If you want a cleaner music bed than separation gives, export a dialogue-only version from your project, dub that, and mix the translated voice with a separate music-and-effects source. Adding that source under the audio of a normal dub would double the music, and using the original mixed track would bring the original voice back too; keeping background music in translated videos explains why.

On the input side, mydubly uses only the default audio track of a multi-track file, with no track picker. If a screen recording puts the microphone on one track and system audio on another, make sure the speech is on the default track, or mix the tracks before uploading. If both voices are mixed into the source audio, speech recognition will hear both.

Next step: one voice, then publish

Mute or remove the extra source, export, and play the result in a normal player on headphones. If you are creating the translated voice in the first place, AI dubbing costs 50 credits per minute with a 2-minute minimum, so the 10-minute lesson above would cost 500 credits (50¢), and the guide on how to dub a video with AI covers the steps from upload to download.

Frequently asked questions

Why can I hear the original voice under my dubbed audio?

Most often the original dialogue is still playing on another track in your editor, frequently as audio embedded in the video clip. Unlink the clip's audio and mute or delete it before exporting. If the original voice is mixed into the same track as the dub, you need to go back to the separate source files.

Why does my video play two audio tracks in my editor but one in a player?

Media players usually play one audio track at a time, while many editors import every track in a multi-track file and play them together. Disable the tracks you do not need in the editor, or check which track holds what before you start cutting.

Why does my screen recording have echoey, doubled sound?

The recorder is probably capturing system audio directly while your microphone also picks up the same sound from the speakers, with a slight delay. Wear headphones, or record only one of the two sources, or put them on separate tracks so you can mute one later.

Can I remove one voice from a mixed audio track?

Not cleanly. Separation tools can reduce music or a voice, but they leave artefacts, and separating one voice from another is especially hard. Re-exporting from the original project or recordings is almost always the better route.

How do I keep only one audio track in a video file?

Remux the file keeping the video and the audio track you want, which copies both streams without re-encoding. In ffmpeg, the map option selects streams, for example the first video stream and the second audio stream, and the copy codec keeps them unchanged.

Is the original audio kept under a mydubly dub?

Partly. AI vocal separation removes the original speech, and the original music and effects are kept under the translated voice, lowered while it speaks. In dense or loud mixes, faint traces of the original voice can remain. For full control over the music, dub a dialogue-only export of your project and combine the translated voice with a clean music stem in an editor; adding a stem to the audio of a normal dub would double the music.