A video file is several streams in one wrapper
A typical video file is not one thing. It is a container holding several independently encoded streams: usually one video stream, one or more audio streams, and sometimes subtitle, chapter or timecode streams. Each stream was produced by its own encoder, so the picture might be H.264 while the sound is AAC. What the codecs themselves do is covered in our explainer on video codecs.
Inside the file, the data of those streams is chopped into small units called packets and laid out interleaved: a little video, a little audio, a little more video, ordered roughly by time. Interleaving means a player reading from the start of the file always has both picture and sound data close at hand, instead of having to read the whole video track before reaching the first second of audio.
The words stream and track are used almost interchangeably. Container specifications such as MP4 and Matroska formally speak of tracks; ffmpeg and most players speak of streams. Either way, each one has an index, a codec, a timebase for its timestamps and, often, a language tag.
What a demuxer actually does
A demuxer is the component that understands one container format and reverses the interleaving. Given a file, it does four jobs:
- Reads the container's headers and index to learn how many streams exist and what codec each one uses.
- Collects the configuration each decoder will need, such as the codec profile, sample rate or the parameter sets an H.264 decoder requires before its first frame.
- Walks through the file and hands out packets one at a time, each labeled with the stream it belongs to.
- Attaches timing to every packet: a presentation timestamp, often a decode timestamp, and a duration, all expressed in that stream's timebase.
What comes out is still compressed data. A video packet is a compressed frame or part of one; an audio packet holds a short run of compressed samples, often around 20 milliseconds or a few dozen milliseconds depending on the codec. The demuxer does not care what is inside, only where each packet starts, how long it is and when it should be presented. How those timestamps work is explained in our guide to PTS and DTS.
Demuxing versus decoding
These two steps are often blurred because a media player does both in one go, but they are different kinds of work.
- Demuxing
- Parses the container and splits it into per-stream packets. Cheap, mostly a matter of reading bytes and following an index
- Decoding
- Runs the codec to turn compressed packets into raw frames or raw audio samples. Expensive, especially for high-resolution video
- Output of demuxing
- Compressed packets plus timestamps and decoder configuration
- Output of decoding
- Uncompressed pictures or PCM samples
- Depends on
- Demuxing depends on the container; decoding depends on the codec
The cost difference is large. Decoding a minute of 4K HEVC takes real processing power and, in browsers and on phones, often dedicated hardware. Demuxing the same minute is closer to copying a file: the program reads headers, follows offsets and passes byte ranges along. That is why remuxing a two-hour film into another container can finish in seconds while re-encoding it can take far longer than the film itself.
Why audio can be pulled out without touching the picture
Because each stream is independent, a demuxer can simply ignore the streams nobody asked for. If you only want the soundtrack, it reads the audio packets and skips the video packets without decoding them. You then have two options:
- Keep the audio compressed and write it to an audio container. An AAC track from an MP4 goes into an .m4a file with no quality change, because nothing was re-encoded.
- Decode only the audio to raw samples, for example to resample it, mix it to mono or feed it to speech recognition. The picture is still never decoded.
With ffmpeg, the first case is a command along the lines of ffmpeg -i talk.mp4 -vn -c:a copy talk.m4a, where -vn drops the video and -c:a copy tells ffmpeg to pass the audio packets through untouched. If the audio codec does not fit the target container, the copy fails and you either choose another container or re-encode the audio. Step-by-step instructions for different tools are in how to extract audio from a video.
The one cost you cannot avoid is reading through the file. Packets are interleaved, so the audio for a long recording is spread across the whole file. Formats with a good index, such as MP4, let the demuxer jump to the byte offsets of the audio data, but it still has to visit most of the file to collect them.
A hypothetical 90-minute 1080p screen recording is about 4 GB, almost all of it video. Copying out its AAC audio stream with a demuxer produces a file of roughly 80 MB at 128 kb/s and takes about as long as reading 4 GB from disk. Decoding and re-encoding the whole video to get the same audio would spend most of its time on picture work that nobody needed.
Demuxers in ffmpeg and in the browser
ffmpeg ships a demuxer for almost every container in use: MP4 and MOV share one, Matroska and WebM share another, and there are separate ones for MPEG transport streams, AVI, Ogg, FLV and many more. Running ffprobe on a file shows what the demuxer found: each stream's index, codec, duration, language tag and which stream is marked as default. That listing is the quickest way to answer questions such as whether a file has two audio tracks or where its subtitles live.
In a web page, demuxing happens in three common ways:
- The video element demuxes internally. When you play an MP4 in a page, the browser's own media stack parses the container, but scripts get no access to the packets.
- Media Source Extensions let a page feed segments of fragmented MP4 or WebM into a player. The browser still does the demuxing; the page only supplies bytes.
- JavaScript or WebAssembly libraries do it in the page itself. Mediabunny and mp4box.js read containers in JavaScript, and ffmpeg.wasm compiles ffmpeg's own demuxers to WebAssembly.
The third route matters because the browser's WebCodecs API offers decoders and encoders but no container handling at all. A page that wants to decode frames or audio itself must first demux the file with a library and then pass the packets, along with the decoder configuration, to WebCodecs. Putting packets back into a container is the opposite job, muxing, which our article on browser video muxing covers.
Choosing a track when there are several
A demuxer exposes every stream; something else has to choose which ones to use. Players normally pick the stream flagged as default, then fall back on language preferences. Command-line tools have their own rules: ffmpeg, without explicit mapping, picks one audio stream by its own heuristics, which may not be the one you expect. Writing -map 0:a:1 selects the second audio stream explicitly.
This choice catches people out with camera files that carry separate microphone tracks, game captures with game and voice on different tracks, and film masters with several languages. Our guide to files with multiple audio tracks explains how to inspect them and set the right default before you hand the file to another tool.
Common demuxing problems and their limits
- Truncated files. If a recording stopped before the container's index was written, the demuxer may find no streams at all even though the media data is present.
- Unsupported containers. A demuxer is format-specific; a tool without the right one reports the file as unsupported even if it could decode the codecs inside.
- Mislabeled extensions. Demuxers usually probe the first bytes of a file instead of trusting its name, but some players and upload forms rely on the extension and refuse a correctly formed file.
- Odd timestamps. Files from broadcasts, edits or streaming captures can carry start offsets or gaps that the demuxer reports faithfully, leaving the player or the next tool to deal with them.
- Demuxing cannot repair damage inside a stream. If packets are corrupt, the demuxer delivers corrupt packets and the decoder shows the errors.
How mydubly reads the audio track
When you choose a video in mydubly, the work starts in your browser. ffmpeg.wasm opens the file, which is mounted rather than copied into the page's memory, demuxes it and decodes the audio, then downmixes it to mono at 16 kHz, the form the speech recognizer needs. If the file contains several audio tracks, only one is used: the default track. There is no track picker, so set the right default before you start.
The decoded audio is cut into chunks of roughly 30 seconds, compressed to Opus and sent over HTTPS for recognition. The video picture never leaves your device. This is what makes private video translation in the browser practical: only the audio stream, and only in compressed speech-sized chunks, is uploaded.
If you ask for a dubbed version, the opposite operation happens at the end. The original picture is copied packet by packet, without re-encoding, into a new file together with the dubbed audio track, which carries the translated voice over the original music and effects in place of the original audio. mydubly accepts MP4, MOV, WebM, MKV and M4V video up to 2 hours long, and does not import from links.
Next step: look inside one of your files
Run ffprobe or open a media inspector such as MediaInfo on a video you use often and note how many streams it has, which codecs they use and which audio stream is marked default. If the speech is on the default track, you can open the file in the mydubly video translator as it is; the MP4 to text page covers the most common case.
Frequently asked questions
What does demux mean in ffmpeg logs?
It refers to ffmpeg's demuxer stage, the part that reads the input container and splits it into streams. Messages about the demuxer usually concern the file's structure, such as a missing index or unusual timestamps, rather than the codecs. Decoder errors are reported separately and point to problems inside a stream.
Is demuxing lossless?
Yes. Demuxing only separates and labels compressed packets, so the data in each stream is unchanged. Quality is lost only if a stream is later decoded and re-encoded with a lossy codec. Copying a demuxed stream into another container keeps it bit for bit identical.
What is the difference between a demuxer and a codec?
A demuxer understands a container format such as MP4 or MKV and extracts packets from it. A codec understands a compression format such as H.264 or AAC and turns packets into frames or samples. A file plays only if the tool has both the right demuxer for the container and the right decoder for each stream it needs.
Can I demux a video on my phone?
Yes, with apps that offer audio extraction or stream copy, and in browser-based tools that run a demuxer in the page. Phones handle demuxing easily because it needs little processing power. Very large files can still be limited by storage or by how much memory a browser tab is allowed to use.
Why does extracting audio sometimes take as long as reading the whole file?
Audio and video packets are interleaved through the file, so collecting every audio packet means reading most of the file even when the video is skipped. The work is mostly disk or network reading, not computation. On a fast drive this is quick; over a slow network share it can take minutes.
Does demuxing change the timing of audio and video?
No. A demuxer passes along the timestamps stored in the file. Timing problems appear when a later step ignores those timestamps, re-encodes with different settings or joins streams that do not share a start time.