What a dual-audio file is, and what it is not
A video file is a container holding separate streams: usually one video stream and one audio stream, sometimes subtitles and chapter markers too. Nothing stops a container from holding two, three or ten audio streams. Each one is a complete soundtrack, the player plays one at a time, and the viewer switches between them from a menu. This is how DVDs and Blu-rays offer several languages on one disc, and how film masters carry separate stereo, surround and commentary tracks.
A dual-audio file is not a mix. The original speech and the translation are never heard together, which is what separates it from the voice-over style compared in dubbing vs voice-over. It is also not the same as YouTube's multi-language audio, where you upload each dubbed track separately in YouTube Studio rather than inside the video file; that platform route is covered in YouTube multi-language audio. A dual-audio file matters when the file itself is the deliverable: a training video on an intranet, a kiosk, a USB drive handed to a partner, a media server at home, or an archive master.
Why the translated MP4 has only one audio track
When mydubly produces a dubbed video, the browser swaps the original audio for the translated track, the new voice mixed over the original music and effects, and keeps the picture exactly as it was, without re-encoding. The result has one audio stream, the translation. That is deliberate: most players, phones and upload forms only ever play the first audio track, so a single-track file behaves predictably everywhere. The mechanics of that swap are explained in browser video muxing.
Alongside the video you can download the translated audio on its own as an M4A file. That separate file is the building block for a dual-audio master: you put it next to the original audio in a new container. Because the translated track is AAC audio inside M4A, it can be copied into MP4, MOV or MKV as-is, with no conversion step.
Containers, codecs and what fits together
Before you combine anything, check what each stream is, because the container must be able to hold every codec you put in it.
- MP4
- Holds H.264, HEVC and AV1 video with AAC audio comfortably. Support for other audio codecs such as Opus or PCM is uneven across players.
- MKV
- The most permissive container. Takes almost any video and audio codec, and stores track names and language tags reliably.
- MOV
- Apple's container. Common for camera and editing exports, often with PCM audio, and handles several audio tracks well in professional tools.
- WebM
- Limited to VP8, VP9 or AV1 video with Opus or Vorbis audio, so an AAC voice track has to be converted to go in.
The common snag is the original audio, not the translation. A MOV export from an editor often carries uncompressed PCM audio, which many MP4 players will not accept. In that case either keep the result as MOV or MKV, or convert only the original audio stream to AAC while copying the video untouched.
Building the file with ffmpeg
ffmpeg is free, runs on macOS, Windows and Linux, and can do the whole job with stream copy, meaning nothing is re-encoded. The commands below assume an original file called lesson.mp4 and a downloaded Spanish voice track called lesson-es.m4a.
- Inspect both files so you know which codecs you have: ffprobe -hide_banner lesson.mp4 lists the streams and their codecs.
- Combine them, keeping the original video and audio and adding the translation as a second audio track: ffmpeg -i lesson.mp4 -i lesson-es.m4a -map 0:v:0 -map 0:a:0 -map 1:a:0 -c copy -metadata:s:a:0 language=eng -metadata:s:a:1 language=spa -disposition:a:0 default -disposition:a:1 0 lesson-dual.mp4
- If the original audio is PCM and you want MP4, convert just that stream by replacing -c copy with -c:v copy -c:a:0 aac -b:a:0 192k -c:a:1 copy.
- For an MKV master with readable menu names, change the output to lesson-dual.mkv and add -metadata:s:a:0 title="English (original)" -metadata:s:a:1 title="Spanish (AI voice)".
- To make the translation the default instead, swap the two disposition values so -disposition:a:1 default and -disposition:a:0 0.
- Add more languages by adding more inputs and maps: a third input becomes -map 2:a:0, tagged with -metadata:s:a:2 language=fra.
- Check the result with ffprobe again, then open it in a player with an audio menu and listen to the first minute of each track.
The -map options are what make this work. Without them, ffmpeg picks one audio stream automatically and silently drops the other, which is the most common reason a "dual-audio" export turns out to have one track.
Doing it without a command line
MKVToolNix is a free graphical tool for building MKV files. You add the original video and the M4A file, untick anything you do not want, set each audio track's language and name in the properties panel, choose which one is the default, and start multiplexing. Nothing is re-encoded, and it is quick even for long files.
Video editors can also export multiple audio tracks, but behavior differs. Professional editors such as DaVinci Resolve and Premiere Pro can route timeline tracks to separate output tracks in certain formats, typically QuickTime or MXF, while consumer editors usually mix every track down to one. Editors also re-encode the picture on export unless they offer a passthrough or smart-render option. If you only need to add a track, a muxing tool is faster and keeps the original picture bit for bit. Check your editor's export documentation for the current options, as they change between versions.
Suppose a warehouse plays a 14-minute onboarding video from a shared drive, and one shift mostly speaks Spanish. Dubbing it costs 700 credits (70¢). The trainer downloads the translated M4A, runs the ffmpeg command above with Spanish as the second track, and names the file onboarding-en-es.mkv. On the break-room PC, VLC shows "English (original)" and "Spanish (AI voice)" in its audio menu. One file replaces two, so nobody plays an outdated copy.
Where players and platforms support switching
Selecting an audio track is a player feature, not a file feature, so test where the file will actually be watched.
- Desktop players such as VLC, mpv and IINA show every audio track in a menu and respect language tags.
- Media servers and many smart TV apps read language tags and remember a preferred language, though support varies by device.
- Phone photo galleries and simple built-in players often play only the default track with no way to switch.
- In a web page, a plain video element generally plays only the default audio track, and the browser API for switching tracks is not broadly enabled. Websites that need language switching usually stream with HLS or DASH, which support alternate audio renditions as separate files.
- Most video-sharing and social platforms keep one audio track from an uploaded file. Upload separate language versions, or use a platform's own feature for extra audio tracks where one exists.
Common mistakes and limits
- Forgetting the -map options, so the output has one audio track.
- Leaving tracks untagged, so viewers see "Track 1" and "Track 2" and guess.
- Marking both tracks as default, or neither, which makes players behave inconsistently.
- Expecting the translated track to match the original mix exactly. It keeps the original music and effects, but it is mono, sung vocals are removed along with the speech, and in dense mixes faint traces of the original voice can remain. If you have the edit project and want the cleanest result, dub a dialogue-only export and mix the translated voice with your music stem first, as described in keeping background music in translated videos, and add that mix as the second track.
- Mixing up files. If the original was trimmed after the translation was made, the voice will be offset; always translate the exact file you are combining.
- Using Opus or PCM audio in MP4 and assuming every player copes. Prefer AAC in MP4, or use MKV.
How mydubly fits this workflow
mydubly produces the ingredients but not the final multi-track file. A dubbing job gives you a translated MP4 whose only audio is the AI voice mixed over the original music and effects, and the same mix as a separate M4A download. The picture is never re-encoded, so the translated MP4 and your original share identical frames. If the video codec cannot go into MP4, the translated video comes out as MKV, and on devices without enough memory for a very large file you get the translated audio to download instead of the merged video, which is exactly the file this workflow needs anyway.
What mydubly does not do: it does not keep the original speech underneath the translation or add the original audio beside it as a second track, it does not write language tags for a second track, and each job produces one language. For three languages you run three jobs and add three tracks yourself. The video file never leaves your device; only the audio is sent for processing, and results are deleted within 30 minutes of completion, so download the M4A promptly.
Next step
Take one short video, produce a translated voice track with AI dubbing, and build a two-track MKV with the commands above. Play it in VLC and in whatever player your audience actually uses before making it the standard. If you are deciding whether a voice track is the right deliverable at all, compare it with subtitles in subtitles vs dubbing.
Frequently asked questions
Does adding a second audio track reduce video quality?
No, as long as you use stream copy. The ffmpeg commands with -c copy, and tools like MKVToolNix, place the existing compressed streams into a new container without decoding them, so every frame is identical to the source. Quality only changes if you re-encode, for example by exporting through a video editor without a passthrough option.
How much bigger does the file get?
Only by the size of the added audio stream, which is small next to the video. A speech track compressed with AAC typically adds roughly a megabyte or less per minute, depending on its bitrate, while the video stream of an HD file is many times that. A two-hour film master with three audio tracks is still dominated by its picture.
Can I include subtitles in the same file?
Yes. Add the SRT as another input and map it. In MKV the SRT can be copied as-is and tagged with a language like any audio track. In MP4, subtitles are converted to the mov_text format with -c:s mov_text, and player support for showing them is less consistent, so MKV is the safer choice for a file with several audio and subtitle tracks.
Why does my phone play only one language?
Many built-in phone players and photo galleries play the default audio track and offer no audio menu. The other tracks are still in the file. Install a player with track selection, such as VLC for mobile, or mark the language that phone viewers need as the default track before distributing the file.
Can I add a track to a video that is already on YouTube?
Not by re-uploading a multi-track file, because YouTube processes one audio track from an uploaded file. YouTube has its own multi-language audio feature for eligible channels, where each dubbed track is uploaded separately in YouTube Studio. Check YouTube Help for current availability before planning around it.