What on-device means for a video file
"On-device" is used loosely in marketing, so it helps to be precise. It can mean everything happens locally, including the AI models. More often, as with browser tools, it means the file is opened and prepared on your device and only a derived piece of it is sent away for the heavy lifting.
For speech translation the derived piece is the audio. Recognition, translation and voice models are large, and running them well inside a typical browser is impractical today, so a hybrid design keeps what the models do not need and sends what they do.
What stays on your device
In a well-built hybrid tool, everything visual stays put. That is often more than people realize:
- Faces, bodies and locations of everyone on camera.
- Slides, documents, spreadsheets and dashboards shown on screen.
- Screen-recorded inboxes, chat windows, browser tabs and notifications that flashed up during a demo.
- Whiteboards, badges, license plates and anything else visible in the background.
- The original, full-quality file itself, so no complete copy of your footage exists on someone else's storage.
For many confidential recordings, the visual layer is where the accidental disclosures live. A presenter can be careful about what they say and still leave a payroll spreadsheet visible in a corner of the screen.
What leaves your device, and why
The audio leaves, because speech recognition needs it. In a typical pipeline the soundtrack is decoded, mixed down to mono and sent in chunks. Whatever is audible goes with it: the speaker, colleagues talking in the background, a phone on speaker, a name read aloud. Then the server produces text from that audio, which means a transcript of everything said exists on the server, at least while the job runs.
If you ask for translation and a dubbed voice, the server also produces translated text and a synthetic voice track. Those results travel back to your device, and copies sit on the server until they are deleted.
A simple threat model for sensitive recordings
A threat model just asks: who might see this, and how? For a recording you are about to translate, the useful questions are:
- Who could access the content in transit or on the provider's systems, and for how long?
- Which parts of the recording carry the sensitive information: the picture, the speech, or both?
- Who has access to your own device, browser and downloads folder?
- Who will receive the finished translated file, and where will they store it?
On-device processing directly addresses the first question for the picture, and shortens the list of copies. It says nothing about questions three and four, and only partly addresses the first for speech.
Where on-device processing genuinely helps
The benefit is real and easy to state: fewer copies of the footage, in fewer places. If the provider's systems were ever breached, misconfigured or accessed by someone who should not have access, there would be no video to find, because it was never there. That shrinks the impact of a whole class of incidents.
There are side benefits too. Upload volume drops from gigabytes to megabytes, which matters on hotel or office networks that log or throttle large transfers. And because the picture is never re-encoded by a third party, the final file keeps its original quality.
The limits: what on-device processing cannot protect
This is the part marketing tends to skip, so it is worth listing plainly.
- The spoken content. If the sensitive information is said aloud, such as a figure, a diagnosis or a person's name, it is in the uploaded audio and in the server-side transcript.
- The voices themselves. Voices can identify people, and they travel with the audio.
- A compromised device. Malware, a malicious browser extension or someone with access to your laptop can see the video regardless of where processing happens.
- Your outputs. The translated video, subtitles and transcripts land in your downloads, and from there they are only as private as wherever you put them.
- The fact of the job. Using a service at all creates account and billing records; read the provider's privacy policy for what is kept.
- Legal and policy obligations. A tool's design does not by itself satisfy your organization's data-handling rules, a consent form or a contract with a client.
Suppose a 25-minute internal town hall shows a slide with salary bands while the speaker says, "we are raising the lowest band to 52,000 this year." With on-device processing the slide never leaves the laptop. The spoken figure, however, is in the uploaded audio and will appear in the transcript and the translated subtitles. If that sentence is the sensitive part, keeping the video local did not protect it.
How mydubly draws the line
mydubly follows the hybrid pattern. The video file stays on your device; the browser extracts the audio and uploads only that audio, in chunks of about 30 seconds. The servers run speech recognition on OpenAI's open-source Whisper model, translate the text with a neural machine translation engine and generate the voice, then return text and the voice track. The final video is assembled in your browser by putting the new audio track alongside the original video stream.
Uploaded audio and results are deleted within 30 minutes of a job finishing, and unfinished jobs expire after 24 hours. mydubly does not claim any certification, so if your organization requires one, treat that as a gap to resolve before use. The private video translation page summarizes the flow, and the privacy policy is the authoritative source for account data and processors.
Practical steps before you process a sensitive video
- Decide whether the sensitive material is visual, spoken or both. If it is mainly spoken, on-device processing helps less than you might hope.
- Check consent: the people recorded should know their speech will be processed by an online service.
- Rename the file to something neutral before choosing it, since file names can reveal more than intended.
- Trim the recording to the part you actually need translated, so off-topic chatter is never sent.
- Read the provider's retention and privacy terms, and compare them with your own organization's rules.
- After downloading, store the outputs with the same care as the original, and delete local copies you no longer need.
If the recording is speech-only, consider audio to text or audio translation directly; the privacy picture is the same as the audio portion of a video job.
Where to go next
For a closer look at what the server side receives and how long it keeps it, read cloud speech processing privacy. If the trade-offs above fit your recording, the private video translation page is where to start a job.
Frequently asked questions
If the video stays on my device, is my recording fully private?
Not entirely. The audio is uploaded so speech can be recognized, translated and voiced, which means what people say, and their voices, are processed on a server. The picture and the original file are what stay local.
Does blurring faces matter if the video never leaves my device?
For the translation step, no: the picture is not sent. It still matters for wherever you publish or share the finished video, because the output contains the original picture.
Can someone tell what was on screen from the uploaded audio?
Only if it was spoken aloud. Text on slides and screens is not in the audio, so it is not transcribed or translated, which also means on-screen text stays in the original language.
Is on-device processing the same as offline processing?
No. Offline means no network at all. A hybrid tool processes the file locally but still needs a connection to send audio for recognition, translation and voice generation.
What should I do if the speech itself is too sensitive to upload?
Do not use a cloud-based speech service for it. Options include a fully offline speech tool run under your own control, or a vetted human transcriber bound by a confidentiality agreement.