Audio & video formats

When a Video File Has More Than One Audio Track

A single video file can hold several separate audio tracks: a microphone and game audio from OBS, two camera inputs, or a stereo mix next to a music-and-effects stem. Players normally pick one, guided by flags and language tags in the file, and many tools, including mydubly, use only one. If the speech you need is not on that track, the transcript or translation comes out empty or wrong, so check the tracks and export a version with the speech as the only or main track.

7 min read · Updated

Tracks are not the same as channels

An audio track, also called an audio stream, is a complete, independent piece of audio inside the container, with its own codec, channel layout and length. A video file can hold one video track and several audio tracks side by side. Channels are the layers inside one track: a stereo track has a left and a right channel, a 5.1 track has six.

The distinction matters because tools treat them differently. When a stereo track is processed for speech recognition, its channels are mixed together, as explained in mono vs stereo for speech. Separate tracks are not mixed; a tool has to choose one, or the user has to combine them.

Where multi-track files come from

  • Streaming and game-capture software. OBS Studio can record up to six audio tracks into one file, so streamers often keep the microphone, game audio and music on separate tracks for later editing.
  • Professional and semi-professional cameras. Inputs such as an on-camera microphone and an external XLR or wireless microphone are often written as separate tracks or channels.
  • Screen and call recorders that can save the microphone and system audio separately.
  • Film and television masters, which may carry a full mix, a music-and-effects (M&E) track for dubbing and sometimes separate dialogue stems.
  • Distribution files and disc rips with several languages, commentary tracks or audio description.
  • Your own translated videos, if you have added a dubbed track next to the original; building such a file is covered in making a video with two audio tracks.

How players decide which track to play

Containers carry hints for players. Matroska (MKV) tracks can be marked default or forced and tagged with a language and a title. MP4 and MOV files have their own enabled flags, language codes and alternate-group settings for tracks that substitute for one another. A player typically plays the default or enabled track, may honor your preferred language if the tags allow it, and otherwise falls back to the first audio track in the file.

Behavior varies widely in practice:

  • Desktop players such as VLC and mpv let you switch tracks during playback, in VLC from the Audio menu under Audio Track.
  • Many simple players, phone galleries and messaging apps play one track and offer no way to switch.
  • Web browsers generally play a single audio track from a local file and provide no track menu.
  • Video editors usually import every track, sometimes as separate clips or as extra channels on one clip, so you may hear them all at once in a timeline.
  • Video platforms differ in whether they accept, ignore or reject extra tracks in an upload.

Because the same file can sound different in different apps, "it plays fine on my computer" does not tell you which track another tool will use.

Why mydubly uses only one track

mydubly reads one audio track from the file you choose. There is no track picker, and the tracks are not mixed, so the speech you want transcribed or translated needs to be on the only track, or on the main one. The safest preparation is to export a version of the file in which the speech is the single audio track, or at least the first and default one. That copy is quick to make and loses nothing, as the steps below show.

The translated video mydubly produces also has a single audio track: the original picture is copied without re-encoding, and the translated audio takes the place of the original audio, so additional tracks from the source are not carried over. Within that one track, the original speech is removed and the music and effects from the track that was read are kept under the translated voice. The mechanics are described in browser video muxing. If you want a final file with both the original and the translated audio for viewers to choose between, add the downloaded translated audio as a second track afterwards; for platforms that support selectable languages, see YouTube multi-language audio.

Example: an OBS recording with the voice on track 2

A streamer records a 90-minute tutorial in OBS with game audio on track 1 and the microphone on track 2, then uploads it for a transcript. The result is sparse and odd, mostly fragments of in-game dialogue, because the commentary never reached the recognizer. Opening the file in MediaInfo shows two audio tracks. A one-line ffmpeg command copies the video and track 2 into a new file without re-encoding, and the next transcript contains the full commentary. The transcript costs 90 credits (9¢).

Checking and selecting the right track

  1. Inspect the file. MediaInfo lists each track as Audio #1, Audio #2 and so on, with codec, channels, language, title and default flag. ffprobe gives the same information: ffprobe -v error -select_streams a -show_entries stream=index,codec_name,channels:stream_tags=language,title -of compact input.mkv
  2. Listen to each track. In VLC, switch tracks from the Audio menu and note which one carries the speech.
  3. Keep only the speech track without re-encoding. In ffmpeg, -map selects streams and -c copy copies them: ffmpeg -i input.mkv -map 0:v:0 -map 0:a:1 -c copy output.mkv keeps the first video track and the second audio track, since audio tracks are counted from zero.
  4. If the speech is split across tracks, such as a microphone on one and a guest call on another, mix them into one. Do this in a video editor, or with ffmpeg's amix filter, which re-encodes only the audio while the video is copied. Check the level afterwards, because mixing tools often lower each input to avoid clipping.
  5. In OBS itself, set the track assignments so that track 1 carries a full mix of microphone and desktop audio, keeping isolated tracks on 2 and above for editing. Recordings then work in single-track tools straight away.
  6. Play the new file in a basic player to confirm that the speech is what you hear by default, then upload that version.

Pitfalls when reducing a file to one track

  • Losing tracks you will want later. Make a new file for processing and keep the multi-track original for editing.
  • Assuming the first track is the main one. Recording software and cameras order tracks by input, not by importance.
  • Choosing a container that cannot hold the audio codec. Uncompressed PCM audio, common in camera files, is handled well by MOV and MKV but not by every MP4 tool; either keep MOV or MKV for the copy, or encode the speech track to AAC.
  • Mixing tracks with very different levels, which leaves one voice buried. Balance them by ear before export.
  • Forgetting language tags. If you build a multi-language file for viewers, set language and title tags so players can label the tracks.
  • Missing a commentary or description track that a viewer actually needs, when trimming distribution files.

Where multiple tracks are the right choice

Multi-track files are excellent working masters. Keeping voices, music and effects separate lets an editor rebalance, replace music, remove a cough from one microphone or produce an M&E track for dubbing later. Livestream recordings in particular benefit, which is why translating livestream recordings recommends separate tracks for voice and game audio. For distribution, multiple language tracks let one file serve several audiences in players that support switching. The constraint is only that single-track tools need a single-track copy.

Next step

Open one of your recordings in a media inspector and count its audio tracks. If there is more than one, make a speech-only or properly mixed copy before uploading it to the video translator or video to text. The MKV to text page has notes on OBS recordings, which are the most common source of multi-track files.

Frequently asked questions

How do I know if my video has multiple audio tracks?

Open it in a media inspector such as MediaInfo, which lists each audio track separately, or check the Audio Track submenu in VLC while the file plays. If you see more than one entry, the file has several tracks. Editors also reveal them on import, often as extra audio clips stacked under the video.

Can I choose which audio track mydubly uses?

No. mydubly reads one audio track and has no track selector. Prepare the file so that the speech is on the only audio track, or at least on the first, default one, by copying the right track into a new file with a tool such as ffmpeg or exporting a mixed version from your editor. This takes seconds and does not touch the picture.

Will removing an audio track reduce video quality?

Not if you use stream copy. Tools such as ffmpeg with the copy option, or editors that offer a passthrough export, copy the compressed video packets unchanged and simply leave out the tracks you do not select. Quality only drops if the tool re-encodes the video, which is unnecessary when you are only choosing audio tracks.

Why does my editor play all the audio tracks at once?

Editors usually import every audio track so you can mix them yourself, and by default they may all be active in the timeline. That is helpful for editing but confusing if you expected one track. Mute or delete the tracks you do not want, or export with only the speech track enabled, before creating a file for other tools.

Is a music-and-effects track the same as background music?

An M&E track contains everything in a mix except dialogue: music, sound effects and ambience. It is produced so that a dubbed voice can be laid on top. If your source video has one, it is still useful after dubbing. The translated audio already keeps the original background through automatic vocal separation, but dubbing a dialogue-only version and mixing the returned translated voice with a clean M&E track in an editor gives the cleanest, fully controllable result, without the limits of automatic separation. Don't add the M&E track to the audio of a normal dub, which already contains the background.