Why a video upload is mostly picture
A video file holds at least two streams: the picture and the soundtrack. The picture is by far the larger. Even efficiently compressed 1080p video commonly runs at several megabits per second, and 4K, screen recordings with lots of motion or camera originals run much higher. A typical stereo soundtrack is a small fraction of that.
For translation and transcription, only one part of the soundtrack matters: the speech. Uploading the whole file to a server means sending gigabytes of picture that the speech models will never look at.
From gigabytes to megabytes, stage by stage
Here is how the data shrinks for one hour of footage, using a 1080p video at 8 Mb/s as the starting point. The video bitrate is an illustration; yours may be higher or lower, but the audio figures depend only on duration.
- Original video file, 60 minutes at 8 Mb/s
- About 3.6 GB
- Audio decoded to 16 kHz mono 16-bit PCM
- 32,000 bytes per second, about 115.2 MB per hour
- Audio as WAV chunks (fallback)
- Same size as PCM, about 960 KB per 30-second chunk
- Audio as Ogg Opus at 32 kb/s
- 4,000 bytes per second, about 120 KB per 30-second chunk and 14.4 MB per hour
- Dub audio, 24 kHz mono as Ogg Opus at 48 kb/s
- 6,000 bytes per second, about 180 KB per 30-second chunk and 21.6 MB per hour
- Number of 30-second chunks
- About 120 per hour, 240 for a 2-hour file
The two big steps are dropping the picture, and then compressing the speech. Going from raw PCM to Opus is roughly an 8-fold reduction on its own.
Why 16 kHz mono is enough for speech
Speech recognition models like OpenAI's Whisper work on audio at 16 kHz, so feeding them CD-quality 44.1 or 48 kHz audio adds data without adding anything the model uses. A 16 kHz sample rate captures frequencies up to 8 kHz, which covers the range that carries most of what makes speech intelligible.
Stereo is similarly unnecessary. Recognition works on a single channel, so the two channels are mixed into one. Together, the lower sample rate and single channel cut the raw audio to a fraction of the size of a typical stereo soundtrack before any compression is applied. Dubbing jobs use 24 kHz mono instead, because the same audio also supplies the music and effects that stay under the translated voice.
Why Opus at 32 kb/s
Opus is an open audio codec standardized by the IETF and designed for both speech and music, including at low bitrates used in voice calls. At 32 kb/s, mono speech sampled at 16 kHz sits comfortably inside the range the codec was built for. For dubs, mydubly uses about 48 kb/s at 24 kHz, giving the kept background more room. It is also widely available: modern browsers include an Opus encoder that a page can reach through the WebCodecs API, so compression happens natively and quickly on your own device.
Where a browser does not offer Opus encoding, the fallback is uncompressed WAV. That still avoids uploading the picture, but each chunk is about eight times larger.
What smaller uploads mean on slow or metered connections
Upload speeds are usually far below download speeds, so this is where the savings become tangible. The figures below are simple arithmetic, ignoring protocol overhead and congestion.
Uploading the 3.6 GB video would take about 4 hours (28,800 megabits divided by 2 megabits per second). The 14.4 MB of Opus audio takes about a minute, or about a minute and a half for a dub's 21.6 MB. If the browser had to fall back to WAV, the 115.2 MB would take a little under 8 minutes, still a fraction of the full file. On a 10 Mb/s uplink the video would take about 48 minutes and the Opus audio around 12 seconds.
The difference matters most in a few situations:
- Mobile hotspots and capped data plans, where 3.6 GB could consume a whole month's allowance.
- Shared office or school networks where large uploads are throttled or flagged.
- Rural and satellite connections with slow or unstable uplinks.
- Travel, where hotel and conference Wi-Fi is often congested.
Small chunks also behave better on unreliable links. Losing a 120 KB chunk to a dropped connection is far less painful than losing progress on a multi-gigabyte upload.
Less data sent, less data exposed
Data minimization, collecting and processing only the data a task needs, is a basic privacy principle. Here it also happens to be the engineering choice. Because the picture is not needed, it is not sent, and because it is not sent, it cannot be retained, logged or leaked by anyone downstream.
It is worth being clear about what does not shrink. Every second of speech is still uploaded, and whatever is said is still processed into text on a server. Compression reduces the size of the audio, not what it contains. The privacy side of that boundary is covered in on-device video privacy.
Trade-offs and limits of a smaller upload
Shrinking the upload is close to free for speech, but not entirely:
- Local work: your device has to decode the soundtrack and encode the chunks. On an older phone or a very long file this takes noticeable time and battery.
- Browser dependence: Opus encoding through WebCodecs is not available everywhere, and the WAV fallback is about eight times larger.
- Attended jobs: since the work happens in your browser, the tab has to stay open while it runs.
- No rescue for bad audio: low-bitrate compression is tuned to preserve speech, but it cannot add back clarity that a noisy or distant recording never had.
- The return trip: results still come back over the network as text and a compressed audio track, although the finished video itself is assembled locally rather than downloaded.
How mydubly keeps uploads small
mydubly applies each step above. In the browser, a WebAssembly build of ffmpeg mounts your file, rather than copying it into memory, and decodes the audio once to 16 kHz mono 16-bit PCM for transcripts, or 24 kHz for dubs; for a full 2-hour file that is about 230 MB at 16 kHz or about 345 MB at 24 kHz, held locally. The audio is split into windows of about 30 seconds, with each cut placed at the quietest 50 ms frame within the last 6 seconds before the 30-second mark. Each chunk is encoded to Ogg Opus with WebCodecs, at about 32 kb/s for transcripts or about 48 kb/s for dubs, or WAV where Opus is unsupported, and several chunks upload in parallel.
On the way back, the server returns text and one AAC voice track, and the browser muxes that track with your original video stream without re-encoding the picture. The private video translation page summarizes this flow. Transcript-only jobs on video to text upload the smaller 16 kHz, 32 kb/s chunks and simply return text and subtitle files, with no voice track to mux.
Trimming the upload further yourself
The pipeline already strips out the picture. You can cut the remaining audio down too:
- Trim dead air at the start and end, such as waiting rooms, countdowns and post-call chatter, in any video editor before choosing the file.
- Remove long breaks or off-topic sections you do not need translated.
- For recordings longer than the 2-hour limit, split them into parts at natural breaks; see translating long videos for how to plan the split.
- If you already have a separate audio export, such as an MP3 or WAV from your recorder, you can use it directly with the audio translator.
- Use a desktop browser for very long files, which tends to handle large local processing more comfortably than a phone.
Where to go next
If slow uploads have kept you from translating large recordings, try a short clip on the private video translation page and watch how quickly the audio goes up compared with your usual uploads. For the technology behind the local steps, see browser-based video processing.
Frequently asked questions
Does a 4K video upload more data than a 1080p one for translation?
No. Only the audio is uploaded, and its size depends on duration, not resolution. An hour of 4K and an hour of 720p both produce about 14.4 MB of Opus audio for a transcript, or about 21.6 MB for a dub.
Does compressing speech to 32 kb/s hurt transcription accuracy?
Opus at that bitrate is designed to keep speech intelligible, and recognition models work on 16 kHz audio anyway, so for clear recordings the effect is generally small. Poor source audio, on the other hand, stays poor however it is encoded.
How much data does a 2-hour video upload with mydubly?
About 28.8 MB as Opus audio for a transcript, or about 43.2 MB for a dub, sent as roughly 240 chunks of about 30 seconds each. If your browser falls back to WAV, it is closer to 230 MB, or about 345 MB for a dub.
Will a translated video download be as large as the original?
Roughly, because the finished file reuses your original video stream. But it is assembled in your browser, so the large file is written locally rather than transferred over the network.
Should I compress my video before translating to save upload time?
It won't help with upload size, since the picture is never uploaded. It would only reduce the quality of your final translated video, so keep the original.