The stages at a glance
- Probe
- Inspect the file: container, streams, codecs, durations, sizes. Decide what the rest of the pipeline must do
- Demux
- Split the container into per-stream compressed packets with timestamps
- Decode
- Turn compressed packets into raw frames or audio samples
- Filter
- Change raw frames or samples: scale, crop, deinterlace, resample, mix, normalize, overlay
- Encode
- Compress raw frames or samples with a codec
- Mux
- Write encoded streams into a container with timing and metadata
- Package
- Prepare for delivery: fast start, streaming segments and manifests, upload, storage
Not every pipeline uses every stage. A remux uses probe, demux and mux only. An audio extraction for speech recognition might probe, demux, decode the audio, filter it and encode it in a speech-friendly codec, never touching the picture. Tools such as ffmpeg bundle all the stages in one program, which is convenient but hides where the work is happening.
Probe: decide before you do anything
Probing reads a file's headers and, sometimes, its first few seconds, to learn what it contains. ffprobe and MediaInfo are the familiar tools. A good pipeline probes first and makes decisions from what it finds:
- Is the container one we can read, and is each codec one we can decode or copy?
- How long is the file, and is it within the limits we support?
- Which audio stream is the default, and are there several?
- Does the video codec fit the output container, so it can be copied, or must it be re-encoded?
Rejecting an unsupported or oversized file at this stage costs milliseconds. Discovering the same problem after an hour of encoding costs an hour.
Demux and decode: separating and unpacking
Demuxing is cheap: it follows the container's index and hands out compressed packets. Our explainer on what demuxing is covers it in detail. Decoding is where real computation begins. A minute of high-resolution video can mean billions of pixel operations to reconstruct, which is why phones and computers carry dedicated hardware decoders. Audio decoding is light by comparison.
The rule at this stage is to decode only what you need. If the job concerns the soundtrack, the video packets can be skipped without decoding. If a stream will be copied unchanged into the output, it never needs decoding at all.
Filter: where the actual change happens
Filters work on raw frames and samples, which is why they sit between decode and encode. Typical video filters scale, crop, rotate, deinterlace, adjust color, change frame rate, or draw text and subtitles into the picture. Typical audio filters resample, change channel layout, adjust loudness, remove noise or mix tracks together.
Filters vary hugely in cost. Resampling audio or cropping video is quick; motion-compensated frame rate conversion or heavy noise reduction can be slower than the encode that follows. Order matters too: scaling down before an expensive filter means that filter processes fewer pixels.
Encode, mux and package
Encoding is usually the most expensive stage for video and the only stage where lossy quality loss is introduced. Encoder speed presets trade time for compression efficiency: a slower preset produces a smaller file at the same quality. Audio encoding is cheap.
Muxing interleaves the encoded streams into a container, writes their timestamps and builds the index. It is fast and lossless, though it can fail when a codec is not allowed in the chosen container. Our article on browser video muxing covers how this works inside a web page and what happens when MP4 cannot take a codec.
Packaging covers whatever the destination needs: moving an MP4's index to the front for web playback, cutting the file into segments with HLS or DASH manifests, adding metadata, or uploading to storage. It is mostly I/O.
Where time and quality are spent
For a typical video job, the rough order of cost from highest to lowest is video encoding, expensive filters, video decoding, audio work, then demuxing, muxing and packaging, which are bound mainly by how fast data can be read and written. Exact proportions depend on codecs, resolution, presets and hardware, so measure your own pipeline instead of trusting a rule of thumb.
Quality follows a simpler rule. Probing, demuxing, muxing and packaging never change the media. Decoding is exact. Loss enters at lossy encoding and at filters that resample, such as scaling. Every decode and re-encode cycle adds another generation of loss, so a file that passes through three tools that each re-encode has been degraded three times.
Avoiding unnecessary re-encodes
The question to ask for each stream is whether its content needs to change. If it does not, copy it.
- Changing container only, such as MOV to MP4: remux every stream. Nothing is decoded or encoded.
- Replacing or adding an audio track: copy the video, encode only the new audio.
- Trimming: copy streams and cut at keyframes, or re-encode only the partial group of pictures at each cut.
- Extracting audio: copy the audio stream into an audio container, or decode just the audio if it must be resampled.
- Burning subtitles into the picture: the video must be re-encoded, so do it once, at the end, from the best available source.
A hypothetical one-hour 1080p video needs a new voice track. A naive pipeline decodes the whole picture, re-encodes it with the new audio and takes a long time, leaving the picture slightly softer than before. A pipeline that copies the video packets and only writes the new audio finishes in roughly the time it takes to read and write the file, and the picture is bit-for-bit identical to the source.
Memory, chunking and parallel work
Media files are large, and a pipeline that loads a whole file into memory before working on it will fail on long recordings, especially in browsers and on phones. Well-built pipelines stream: they read packets incrementally, process them and write output as they go, so memory use stays roughly flat regardless of duration. The browser-specific techniques are described in processing large video files in the browser.
Long jobs are also split for parallelism. Video encoders can work on separate segments in parallel, and audio for speech recognition is commonly cut into chunks at quiet points so pieces can be processed independently. The trade-offs of choosing chunk boundaries are discussed in audio chunking for speech recognition.
Error handling and common pipeline mistakes
Media pipelines meet malformed files, interrupted uploads and unexpected codecs constantly. Habits that keep failures manageable:
- Validate at probe time and give a specific reason when rejecting a file, such as an unsupported codec or an over-length duration.
- Never modify the original. Write outputs to new files and keep the source until the job is confirmed.
- Decide per stage whether corrupt packets are skipped or fatal. Skipping a damaged frame in a long recording is usually better than failing the job.
- Retry network stages with limits, and make stages safe to repeat so a retry cannot produce duplicate output.
- Clean up partial files when a job fails or is cancelled.
- Log which stage failed and with what input, so a bug report says more than "processing failed".
The common mistakes are the mirror image: re-encoding by default, decoding streams nobody uses, assuming every file starts at time zero, ignoring extra audio tracks, and only discovering limits at the end of a long job.
mydubly's media steps as an example
mydubly's pipeline is split between your browser and a server, and the media stages are easy to map onto the model above. The speech recognition, translation and voice stages in the middle are explained in how AI video translation works.
- Accept. Video in MP4, MOV, WebM, MKV or M4V and audio in MP3, WAV, M4A, AAC, OGG or FLAC, up to 2 hours per file, chosen from your device.
- Demux and decode, in the browser. ffmpeg.wasm reads the file, which is mounted rather than copied into memory, and decodes the default audio track. Only one audio track is used.
- Filter. The audio is downmixed to mono at 16 kHz and cut into chunks of roughly 30 seconds, each cut placed at the quietest 50 ms in the last 6 seconds of the window.
- Encode and send. Each chunk is encoded to Opus at 32 kb/s with WebCodecs, with a WAV fallback, and sent over HTTPS. The picture never leaves your device.
- Assemble the voice. For a dub, the server builds the translated voice as an AAC track in an M4A file.
- Mux, in the browser. Mediabunny copies the original video packets without re-encoding and adds the AAC track in place of the original audio. The output is MP4, or MKV when the video codec cannot go in MP4. Devices with little memory may be offered the voice track alone.
Two error-handling rules follow from this design. Closing or leaving the page cancels the job, because the browser is part of the pipeline. And credits are reserved when a job starts and refunded if no result is produced. Uploaded chunks and results are deleted within 30 minutes of a job finishing. This split is also why private video translation is possible without uploading the picture.
Next step: map your own workflow
Write down the stages your current workflow runs for one typical file and mark each stream as copied or re-encoded. Any stream that is re-encoded without changing is time and quality you can recover. If your workflow ends in a transcript or a translated version, how to transcribe a video and the mydubly video translator show where mydubly can take over.
Frequently asked questions
What is the difference between transcoding and remuxing?
Transcoding decodes a stream and encodes it again, usually with a different codec or settings, which costs time and, with lossy codecs, quality. Remuxing moves already-encoded streams into a new container without touching their contents. Remuxing is fast and lossless but only possible when the codecs fit the new container.
Which stage of video processing takes the longest?
Usually video encoding, followed by expensive filters and video decoding. Demuxing, muxing and packaging are mostly limited by disk or network speed. The exact balance depends on resolution, codec, presets and whether hardware acceleration is used.
Does every pipeline need to decode the video?
No. If the picture is not being changed, it can be copied packet by packet. Jobs that only concern audio, such as transcription or replacing a soundtrack, never need to decode the video at all.
Why does my pipeline run out of memory on long files?
Most often because some stage loads the whole file or the whole decoded stream into memory instead of streaming it. Raw decoded video in particular is enormous. Processing packets incrementally and writing output as you go keeps memory roughly constant.
How do I know whether a tool re-encodes my video?
Check its options or logs for a copy mode; in ffmpeg, -c copy means no re-encoding. A fast job with an output roughly the same size as the input usually indicates copying, while a slow job with a different size usually indicates re-encoding. Comparing the codec settings in a media inspector before and after confirms it.
Where does quality loss happen in a media pipeline?
At lossy encoding stages and at filters that resample, such as scaling. Probing, demuxing, decoding, muxing and packaging do not change the media. Minimizing the number of lossy encodes is the most effective way to preserve quality.