What changed in the browser
For most of the web's history, a page that wanted to do anything with a video had one option: upload it. The browser could play media through the video element, but it could not hand a page the decoded frames or audio samples in a form useful for processing, and JavaScript was too slow to run a real decoder.
That changed in stages. WebAssembly made it practical to compile existing C and C++ media code to run in the page at close to native speed. WebCodecs then exposed the decoders and encoders the browser already ships for playback. Add the long-standing File API and background threads, and a web page can now do work that used to require desktop software or a server farm.
The File API: reading a file in place
When you pick a file with a file input or drag it onto a page, the browser gives the page a File object. That object is a handle, not a copy. Nothing is read until code asks for bytes, and code can ask for just a slice, for example the first few megabytes where a container keeps its index, or a range in the middle where a particular stretch of audio sits.
This is what makes multi-gigabyte inputs feasible. A page that reads a file slice by slice never needs the whole thing in memory at once. It also means a page can inspect a file, decide what it needs and leave the rest untouched on disk.
WebAssembly builds of ffmpeg
ffmpeg is the open-source workhorse behind a large share of the world's media tooling. Compiled to WebAssembly, usually with the Emscripten toolchain, it runs inside the page with the same demuxers, decoders and filters as the desktop version. Projects such as ffmpeg.wasm package this for the web.
The catch is the file system. Native ffmpeg expects to open files by path. Emscripten emulates a file system, and its default in-memory version means copying the input into WebAssembly memory first, which fails quickly for large videos. Emscripten also offers a worker file system that exposes a File object to the program without copying it, so ffmpeg reads bytes from the original file on demand. Mounting the file this way, rather than copying it, is the difference between handling a 300 MB clip and handling a 6 GB recording.
WebCodecs: the browser's own encoders and decoders
WebCodecs gives JavaScript direct access to the codecs the browser already uses for playback: VideoDecoder, VideoEncoder, AudioDecoder and AudioEncoder. These are native code, and video codecs are often hardware accelerated, so they are much faster than a WebAssembly equivalent.
WebCodecs deliberately does only one job. It turns encoded chunks into raw frames or samples and back again, but it does not read or write containers like MP4, MKV or WebM. For that you need a separate muxing and demuxing library. Codec availability also differs between browsers and versions, so careful code asks the browser whether a configuration is supported before relying on it and keeps a fallback ready.
Web Workers and keeping the page responsive
Decoding an hour of audio or remuxing a large file can take a while. Run on the main thread, that work would freeze scrolling, clicks and progress bars. Web Workers run scripts on background threads, and media pipelines typically put ffmpeg and heavy encoding inside a worker that reports progress back to the page.
Multithreaded WebAssembly goes further by letting one program use several cores, but it depends on SharedArrayBuffer, which browsers only enable for pages served with cross-origin isolation headers. Many tools therefore run ffmpeg single-threaded and get parallelism from doing several independent jobs at once instead.
Writing large results back to disk
Producing the output is the mirror image of reading the input. A page can assemble a result in memory and offer it as a download, which is fine for small files. For large ones, some browsers let a page write directly to a file the user chooses, streaming data out as it is produced so the full output never sits in memory. Where that is not available, tools fall back to building the file in memory, which caps the practical output size on that browser.
How mydubly uses these building blocks
mydubly is a concrete example of the split. When you choose a video in the video translator, the browser mounts the file and a WebAssembly build of ffmpeg decodes its audio once to 16 kHz mono 16-bit PCM, or 24 kHz for a dub. The audio is cut into windows of about 30 seconds, each cut placed at the quietest 50 ms frame within the last 6 seconds before the 30-second mark so sentences are not split mid-word.
Each window is encoded to Ogg Opus, at about 32 kb/s for a transcript or about 48 kb/s for a dub, using the browser's native WebCodecs encoder, falling back to WAV where Opus encoding is not supported, and several chunks upload in parallel. Speech recognition, translation and voice generation happen on mydubly's servers. When the AAC voice track comes back, the mediabunny library muxes it with the original video stream in the browser without re-encoding the picture, producing MP4, or MKV if the video codec cannot go in MP4. Large outputs can be written straight to disk.
Suppose the file is a 6 GB 4K MP4. The page reads only what ffmpeg needs to decode the audio, producing about 173 MB of PCM (16,000 samples per second, 2 bytes each, 5,400 seconds). That becomes about 180 Opus chunks of roughly 120 KB, around 21.6 MB uploaded in total. The 6 GB of video never leaves the laptop, and the final file reuses the original video stream untouched.
The private video translation page covers the same flow from the point of view of what data goes where.
Limits of processing video in a browser
Browser processing is powerful, but it is not a drop-in replacement for a workstation or a server.
- Memory ceilings: classic 32-bit WebAssembly can address at most 4 GB, and browsers often allow less, especially on phones. Anything that decodes full-resolution video frames into memory hits this fast, which is why audio-only decoding and stream copying are far more practical than re-encoding video.
- Speed: WebAssembly is fast but generally slower than the same code compiled natively, and a single-threaded build uses one core. Re-encoding long video in the page can take a long time.
- Codec coverage: WebCodecs support for particular codecs and encoders differs by browser, so a pipeline needs fallbacks, and a WebAssembly ffmpeg build may leave out codecs to keep its download size down.
- Tab lifecycle: mobile browsers may pause or discard background tabs, and closing the tab stops the work. Long jobs need the page to stay open.
- Device variety: the same page can run on a recent desktop and a five-year-old phone, so performance is unpredictable in a way server processing is not.
When browser processing is worth it
The approach shines when the expensive or sensitive part of a file is something the page can leave alone. Translation and transcription need only the speech, so a page can extract a small audio stream, send that and keep the footage local. It also saves bandwidth on slow connections and removes the wait for a giant upload. The trade-offs between this and full server processing are covered in client-side vs server-side video processing.
It is a poor fit for jobs that genuinely need heavy compute on every frame, such as re-rendering 4K video with effects, or running very large AI models that are impractical to download into a browser.
Where to go next
If you want to see the split in practice, open the private video translation page and run a short clip: watch the progress as the audio is extracted locally before anything is sent. For the privacy implications in detail, read on-device video processing and privacy.
Frequently asked questions
Can a website read my whole hard drive if I pick one video?
No. Choosing a file grants the page access to that one file only. It cannot browse other files or folders unless you explicitly pick them too.
Why not re-encode the video in the browser as well?
Re-encoding video means decoding and compressing every frame, which is slow in a page and heavy on memory. Copying the existing video stream into a new container avoids all of that and keeps the original picture quality.
Does browser-based processing work on phones?
Often, yes, but phones have less memory and may suspend background tabs. Keep the tab in the foreground for long files, and expect a desktop browser to handle very large inputs more comfortably.
What is the difference between ffmpeg.wasm and WebCodecs?
ffmpeg.wasm is a WebAssembly build of a complete media toolkit, including container handling and many codecs, running inside the page. WebCodecs is a browser API that exposes the browser's own native codecs but leaves container reading and writing to other libraries.
Do I need to install anything for in-browser processing?
No. Everything runs in a modern browser. The page may download a WebAssembly module the first time, which the browser can cache for later visits.