The two approaches side by side
Before weighing them, it helps to see the main differences in one place. Neither column is better across the board; each wins different rows.
- Where the file goes
- Client-side: stays on the device. Server-side: a full copy is uploaded and stored, at least temporarily.
- Bandwidth
- Client-side: little or none for the media itself. Server-side: the whole file must be uploaded, and often the result downloaded.
- Compute available
- Client-side: whatever the user's laptop or phone has. Server-side: whatever the provider provisions.
- Predictability
- Client-side: varies widely by device and browser. Server-side: consistent hardware and software.
- Large AI models
- Client-side: impractical for big models in a browser. Server-side: routine.
- User's role
- Client-side: the page must stay open while it works. Server-side: the user can close the tab and come back.
Privacy: who ends up holding a copy
The privacy argument for client-side processing is about copies. Every upload creates at least one copy of the file on infrastructure the user does not control, and often more: storage buckets, processing caches, backups and logs. Each copy has its own retention rules and access controls, and each is something that could be exposed by a mistake.
When processing happens in the browser, those copies never exist. That is a structural guarantee rather than a policy promise, which is why it appeals to legal, HR and research teams. A server-side provider can delete files promptly and restrict access carefully, but the user has to trust that it does.
The protection only covers what actually stays local, though. If part of the job runs on a server, the data sent for that part needs the same scrutiny as any upload; on-device video privacy walks through that boundary in detail.
Bandwidth and upload time
Video is large, and upload links are usually much slower than download links. A server-side tool must wait for the entire file before it can finish, and on a slow or metered connection that wait can dominate the whole job. Client-side processing removes it, or shrinks it to whatever small derivative has to be sent. The arithmetic of that shrinkage is worked through in how to reduce upload size for video translation.
Device performance and predictability
Here the server pulls ahead. A provider chooses its hardware and knows how long a job will take on it. A browser-based tool runs on whatever the user brings: a recent desktop, a corporate laptop with aggressive security software, or an older phone with a fraction of the memory. WebAssembly code runs well but generally slower than the same code compiled natively, and browsers impose memory ceilings that native software does not have.
The practical consequence is that client-side steps should be the cheap ones: reading a file, decoding audio, copying streams between containers. Expensive per-frame work, like re-encoding long 4K video, is a poor fit for an arbitrary user's device.
Reliability and failure modes
Each approach fails differently.
- Server-side: a failed or interrupted upload of a large file may have to restart, depending on whether the service supports resumable uploads. Once the file is uploaded, the job usually continues even if the user's laptop goes to sleep.
- Client-side: the work stops if the tab is closed, the device sleeps or a mobile browser discards a background tab. But there is no giant upload to fail, and a pipeline that sends work in small pieces loses only the pieces in flight when a connection drops, provided it retries them.
- Both: browser and codec support must be tested. A server can install any codec it likes; a browser offers what its vendor ships.
Suppose a researcher has a 50-minute, 4 GB interview recording and a café connection that drops every few minutes. A server-side tool needs all 4 GB to arrive before work begins, and each drop risks restarting a long upload. A hybrid tool decodes the audio locally and sends it in 100 small chunks of about 30 seconds each, so a dropped connection affects only a chunk or two. The trade is that the researcher must keep the laptop awake and the tab open until the finished file is assembled.
Where client-side processing falls short
It would be misleading to present the browser as the answer to everything. Its limits are concrete:
- Large models: modern speech recognition, translation and voice models are far too big to download and run comfortably in a typical browser tab.
- Memory: 32-bit WebAssembly cannot address more than 4 GB, and browsers often allow less, especially on phones.
- Battery and heat: long local jobs drain laptops and phones.
- Consistency: a job that takes a moment on one device may struggle on another, which makes support and expectations harder to manage.
- Attended operation: the user's device has to participate until the end.
Hybrid designs, and how mydubly splits the work
The common answer is a hybrid: do locally what is cheap or sensitive, and send to a server only what genuinely needs server-scale models. mydubly is built this way.
- In the browser
- Decode the file's audio once with a WebAssembly build of ffmpeg, cut it into windows of about 30 seconds at quiet moments, encode each chunk to Opus with WebCodecs, upload chunks in parallel, and finally mux the returned voice track with the original video stream.
- On the server
- Speech recognition with OpenAI's open-source Whisper model, translation with a neural machine translation engine, voice generation with Chatterbox Multilingual by default, and timing fit that produces one AAC track matching the video length.
- Never uploaded
- The video picture and the original file.
- Kept on the server
- Uploaded audio and results only until they are deleted, within 30 minutes of the job finishing; unfinished jobs expire after 24 hours.
Because the picture is copied into the new file rather than re-encoded, the output keeps its original video quality, and because only audio travels, upload volume stays small. The cost is that the tab must stay open. The private video translation page describes this split from a user's point of view, and the video translator is where you run it.
Choosing an approach for your own project
If you are designing a media feature rather than choosing a tool, a few questions settle most decisions:
- List each processing step and mark whether it needs a large model or heavy per-frame compute. Those steps belong on a server.
- Identify which parts of the input are sensitive. If the sensitive parts are not needed by the server steps, keep them local.
- Estimate upload size and your users' typical connections. If uploads would dominate the job, move preparation work to the client.
- Check browser support for every codec and API you rely on, and design fallbacks.
- Decide whether users can be expected to keep a tab open. If not, server-side processing of the whole job may be the only workable choice.
Where to go next
To try a hybrid pipeline on your own footage, start a job on the private video translation page with a short clip. For the browser technologies underneath, see browser-based video processing.
Frequently asked questions
Is client-side processing always more private than server-side?
For the data that stays local, yes, because no copy is created elsewhere. But many client-side tools still send something to a server, so the privacy benefit is only as large as the portion that genuinely never leaves the device.
Why can't AI speech models simply run in the browser?
Smaller models can, with WebAssembly or WebGPU. Larger, more accurate recognition, translation and voice models are big downloads and need more memory and compute than a typical tab can offer, so most tools run them on servers.
Does server-side processing mean my file is kept forever?
Not necessarily; retention depends on the provider. Check how long uploaded files and results are kept, and whether deletion is automatic or needs a request.
Which approach is cheaper to run?
Client-side processing shifts compute and bandwidth costs to the user's device, which lowers the provider's costs. Server-side processing concentrates them on the provider but makes performance consistent.
Can a hybrid tool keep working if I close the tab?
Usually not for the browser-side steps. In mydubly's case the final video is assembled in your browser, so the tab needs to stay open until the download is ready.