Why big video files break browser tabs
A browser tab is not a desktop application with free run of the machine's memory. Each tab works within limits set by the browser and the operating system, and those limits are lower and less predictable than most people expect.
- WebAssembly memory. Standard WebAssembly uses 32-bit addressing, so a module's linear memory can never exceed 4 GB, and browsers or particular builds may cap it lower. Anything a WebAssembly program processes in memory has to fit inside that space alongside the program itself.
- JavaScript buffers. Large ArrayBuffers are allowed, but allocating one the size of a long 4K recording often fails outright or pushes the tab into heavy memory pressure.
- Phones. Mobile browsers, iOS in particular, terminate or reload tabs that use too much memory, often without a clear error. The user simply sees the page restart.
- Shared machines. Other tabs, extensions and apps compete for the same memory, so a file that works on one laptop can fail on another.
Video files are routinely larger than all of these comfort zones. A 90-minute 1080p recording can easily be several gigabytes, and the 2-hour limit for a single file in mydubly allows recordings far larger still.
The copy trap: loading the whole file before doing anything
The most common way to crash a tab is also the most obvious implementation. Code reads the selected file into an ArrayBuffer, then hands those bytes to a WebAssembly build of ffmpeg by writing them into its virtual file system. The default virtual file system keeps file contents in memory, so at that point the video exists twice: once in the JavaScript buffer and once inside WebAssembly memory.
For a 300 MB clip that is wasteful but survivable. For a 3 GB recording it cannot work at all, because the copy inside WebAssembly alone would nearly fill the 4 GB address space before ffmpeg has allocated anything for decoding. The fix is not a bigger buffer; it is avoiding the copy.
Mounting the file instead of copying it
Emscripten, the toolchain used to compile ffmpeg to WebAssembly, offers a file system called WORKERFS for exactly this case. Instead of copying a file's contents, it exposes the browser's File object inside a web worker as if it were an ordinary file on disk. When ffmpeg asks to read bytes at some offset, WORKERFS reads just that slice from the original file on demand.
The effect is that ffmpeg can seek through and read a multi-gigabyte container while only the slices currently being processed occupy memory. Nothing is preloaded. This works because the browser's File API treats a selected file as a handle to data on disk, and any byte range can be read from it without loading the rest.
Mounting has requirements: it runs inside a worker, and it depends on the ffmpeg build exposing that file system. A robust tool therefore tries to mount first and only falls back to copying when mounting is unavailable, accepting that the fallback works only for smaller files.
Decode once into the smallest useful representation
Mounting solves reading. The next question is what to keep. For speech processing, the video frames are irrelevant and the audio can be reduced dramatically: mono instead of stereo, 16 kHz instead of 44.1 or 48 kHz, 16-bit samples. That format is all a speech recognizer such as Whisper needs.
Decoding the whole audio track to that format once, then working from the result, is far cheaper than going back to the container repeatedly. ffmpeg still has to read through the file, because audio packets are interleaved with video packets, but it skips the video stream entirely and decodes only the audio. The numbers follow directly from the format, 16,000 samples per second at 2 bytes each:
- One second of 16 kHz mono 16-bit audio
- 32,000 bytes
- 30 seconds
- 960 KB
- One hour
- about 115 MB
- Two hours
- about 230 MB
- 30 seconds as Ogg Opus at 32 kb/s
- about 120 KB, roughly 8 times smaller than the raw audio
- Two hours as Ogg Opus at 32 kb/s
- about 29 MB
A buffer of a couple of hundred megabytes is comfortable on most desktops and manageable on many phones, whatever size the original video was. Everything downstream, from finding pauses to cutting chunks, reads from that one buffer. Cutting a chunk can then be a view into the buffer rather than a copy, so slicing two hours into a few hundred pieces allocates almost nothing new.
How mydubly keeps a two-hour video inside browser limits
mydubly is built around these techniques because the video file stays on the user's device; only the audio is uploaded. As implemented, the browser side of a job runs like this:
- Mount the selected file into a WebAssembly build of ffmpeg rather than copying it into WebAssembly memory, so multi-gigabyte files work. If mounting is not possible, it falls back to copying, which suits smaller files.
- Decode the audio once to 16 kHz mono 16-bit PCM, about 230 MB for a two-hour file, then release the mount.
- Measure loudness in 50 ms frames across that buffer and plan windows of about 30 seconds, each ending at the quietest moment in the 6 seconds before its target, as described in audio chunking for speech recognition.
- For each window, take a view into the buffer and encode it to Ogg Opus at 32 kb/s with the browser's native WebCodecs encoder, falling back to WAV where Opus encoding is not supported. Several chunks encode and upload in parallel.
- When the server returns the finished voice track, merge it with the original video stream by copying packets, without re-encoding the picture, and stream the result to disk when the browser supports it.
The ffmpeg work runs in a web worker, so the page stays responsive while a long file is decoded.
Suppose a 90-minute 1080p talk recorded as a 6 GB MOV file. Mounted rather than copied, the file never sits in memory as a whole. Its audio decodes to about 173 MB of PCM. That becomes at least 180 windows of up to 30 seconds, which together upload as roughly 22 MB of Opus audio. If the browser can stream to disk and has space for about 6 GB of output, the merged MP4 is written there in pieces. A transcript of the talk costs 90 credits, which is 9 cents; a translated voice costs 4,500 credits, which is $4.50.
Streaming the output to disk instead of memory
Reading large inputs is only half of the problem; a merged video is as large as the original, so building it in memory brings back every limit described above. The origin private file system gives each website a sandboxed storage area, and in browsers that support writable streams there, a page can write a file progressively, piece by piece, while memory use stays flat.
mydubly checks how much storage the browser reports as available before choosing this route, and only uses it when the output plus a safety margin will fit. Otherwise it assembles smaller outputs in memory, with lower allowances on phones, and when neither fits it offers the translated audio file on its own. The merge itself is described in browser video muxing.
Practical steps for working with large files in a browser
If you are about to run a very large recording through any browser-based tool, these steps reduce the chance of a failed run:
- Use a desktop or laptop browser for files of several gigabytes. Current desktop versions of Chrome and Edge have the most complete support for streaming output to disk.
- Check free disk space first. A merged output needs roughly as much space as the original video, plus some margin.
- Close other memory-heavy tabs and applications, especially on machines with 8 GB of memory or less.
- Keep the tab open and visible until the job finishes. Some browsers slow down or suspend background tabs, particularly on phones.
- Plug in a laptop for long files. Decoding audio from a large container keeps the processor busy for a while.
- Split recordings longer than 2 hours into parts at a natural pause before uploading; the article on translating long videos covers planning those splits.
- On a phone, consider transcript output or downloading the translated audio rather than a merged video, as explained in translating video on a phone.
When in-browser processing of large files is the right choice
- Privacy-sensitive footage, where keeping the video on the device matters more than raw speed; see private video translation.
- Slow or metered connections. Uploading tens of megabytes of compressed speech is practical where uploading several gigabytes of video is not, a comparison worked through in reducing upload size.
- Recordings straight from cameras and screen recorders, which are large but only need their audio extracted.
- Workflows where the picture must remain untouched, since nothing in this pattern re-encodes the video.
Limits that remain, even with careful engineering
- Phones are the weak point. Memory allowances are small and unreported on some devices, so very large files may only produce transcripts or an audio download rather than a merged video.
- WebAssembly is slower than native code. Decoding audio from a long file in the browser takes longer than the same job in desktop ffmpeg.
- Storage quotas vary. Private or incognito windows often allow far less storage, which can block streaming output to disk.
- Browser support is uneven. Writable streams in the origin private file system, the WebCodecs Opus encoder and worker file systems are not available in every browser version, so tools need fallbacks and sometimes have to refuse a file.
- The tab must stay alive. Closing it, or letting a phone reclaim it, interrupts the local stages of the job.
- Reading speed still matters. Even with mounting, the whole container has to be read once to extract the audio, so a slow external drive makes that step slow.
Try a large file without uploading the video
The simplest test is your largest real recording. Open mydubly's private video translation on a desktop browser, choose the file and watch the first stage: the audio is extracted locally, and only compressed audio chunks are sent. If you only need text, the video to text tool uses the same on-device extraction and costs 1 credit per minute.
Frequently asked questions
Why does my browser crash when I open a large video in a web app?
Usually because the app reads the entire file into memory, sometimes twice, before processing it. Browsers cap memory per tab and WebAssembly memory has a hard 4 GB ceiling, so multi-gigabyte files exceed the limit. Apps that read the file lazily and keep only the audio avoid the problem.
What is WORKERFS in ffmpeg.wasm?
It is an Emscripten file system that makes browser File and Blob objects readable inside a web worker without copying them into WebAssembly memory. ffmpeg reads byte ranges from the original file as it needs them. That allows very large inputs to be processed with modest memory use.
How much memory does decoding a two-hour video's audio need?
At 16 kHz, mono and 16-bit, audio takes 32,000 bytes per second, so two hours comes to about 230 MB. The size of the video stream does not matter, because the video is skipped during decoding. Higher sample rates or stereo would multiply that figure.
Is there a maximum file size for mydubly?
The limit is duration, not size: up to 2 hours per file. Large files are mounted rather than copied, so multi-gigabyte videos are fine for transcripts and audio. Whether the merged video can be produced on the device depends on the browser's storage support and free space; if it cannot, you can still download the translated audio.
Does processing a large file in the browser upload the video anywhere?
Not in mydubly. The browser extracts the audio and uploads only compressed audio chunks; the video stays on your device and the final file is assembled there. Uploaded audio and results are deleted within 30 minutes of the job finishing.