What a cloud speech pipeline receives
Every cloud speech service needs the audio itself; there is no way to recognize speech without hearing it. What varies is how much else comes along. Some services take the entire media file, picture included. Others take an audio stream extracted on your device. Almost all also receive the settings needed to run the job, such as the target language.
In mydubly's case, the browser extracts the soundtrack, mixes it to mono at 16 kHz and splits it into chunks of about 30 seconds, cut at quiet moments. Each chunk is compressed to Opus at 32 kb/s and uploaded. The server never receives the picture, but it does receive everything audible in the soundtrack, including background voices.
What the pipeline creates from your audio
Uploading audio is only the start. Processing creates new data, and that derived data is often more sensitive than the audio because it is searchable and easy to copy.
- Transcript segments
- Text of what was said, with start and end timestamps, produced by speech recognition. mydubly uses OpenAI's open-source Whisper model.
- Translated segments
- The same content in the target language, produced by a neural machine translation engine.
- Synthetic voice clips
- Translated sentences spoken by the chosen stock voice, then fitted to the original timing.
- Joined voice track
- One AAC track matching the length of the video, sent back to your browser.
- Subtitle files and transcripts
- SRT and VTT subtitles in the target language, and transcripts in both languages, delivered for download.
A transcript turns a recording that someone would have to listen to into a document anyone can skim or search. That is the point of the service, and it is also why retention matters.
How long mydubly keeps audio and results
mydubly's retention is short and tied to the job's life cycle:
- Uploaded audio and results are deleted within 30 minutes of a job finishing.
- Unfinished jobs, for example ones abandoned when a tab was closed partway through, expire after 24 hours.
- The video file is never uploaded, so there is nothing of it to delete.
Suppose you start a translation of a 40-minute webinar recording and the job finishes at 14:10. By 14:40 at the latest, the uploaded audio chunks, transcripts, translations and voice track are deleted from the server. If you had instead closed your laptop halfway through a job started at 09:00, the unfinished job would expire by 09:00 the next day.
Your own downloads are a separate matter. Once the files are on your device, keeping or deleting them is up to you. For account records and the processors involved, the privacy policy is the authoritative source.
Why retention windows matter
Deletion is the strongest privacy control a cloud service can offer, because data that no longer exists cannot be leaked, subpoenaed or misused later. Short retention also limits how long any mistake, such as a misconfigured storage bucket, could expose your content.
A retention window is only meaningful if it covers everything. When you evaluate any provider, check whether its stated retention applies to the uploaded audio, the transcripts, the translations, generated audio and any logs or caches, not just the original upload.
Questions to ask any speech provider
These questions apply to mydubly and to any alternative you are considering. A provider that cannot answer them clearly has told you something useful.
- What exactly is uploaded: the whole file, the audio only, or something smaller?
- How long are uploads, derived text and generated audio kept, and is deletion automatic?
- Is customer audio or text used to train or improve models, and can you opt out?
- Which third parties process the data, and in which countries?
- Who inside the company can access customer content, and is access logged?
- Can you request deletion of your data, including account records?
- Does the provider hold the certifications or sign the agreements your organization requires?
- How and when would you be told about a security incident?
For mydubly, the first two are answered in this article and on the private video translation page. mydubly does not claim any security or compliance certification, so if question seven matters to your organization, raise it before uploading regulated material, and check the privacy policy for the rest.
When short-lived cloud processing is a reasonable choice
For a lot of recordings, cloud processing with tight retention is a sensible balance. Lectures, webinars, training videos, product demos and published talks are often semi-public or low-risk, and the time saved by machine transcription and translation is substantial. Keeping the picture local removes the visual layer from the risk entirely, and prompt deletion keeps the speech exposure brief.
It also beats some common alternatives. Emailing a recording to a freelancer, sharing it through a consumer file-sharing link or uploading it to a general video platform can leave copies around far longer than half an hour.
Risks that remain, and where cloud processing is the wrong call
Short retention narrows the exposure; it does not eliminate it.
- While the job runs, the audio and its transcript exist on someone else's systems.
- Voices are personal data in many contexts and can identify speakers even without names.
- Whatever is said aloud, including names, health details or financial figures, becomes text.
- Retention policies describe intent and design; you are trusting the provider to implement them.
- Your organization's rules, client contracts or research ethics approvals may forbid sending certain recordings to any external processor, however briefly.
If any of those apply, consider a fully offline speech tool run on hardware you control, or a human transcriber under a confidentiality agreement. For a recording that may simply contain the odd sensitive passage, trimming that passage out before uploading is often enough.
Reducing what you send in the first place
The least risky audio is audio you never upload. Before starting a job:
- Trim the recording to the section you actually need.
- Cut pre-meeting chatter, breaks and off-topic asides, which often contain the most personal remarks.
- Give the file a neutral name.
- Confirm everyone recorded agreed to their speech being processed by an online service.
These steps work with any provider, and they pair naturally with the data minimization already built into a tool that extracts audio locally. The upload-size side of this is covered in reducing upload size for video translation.
Where to go next
If your recording passes the questions above, the private video translation page is the place to start a job; for speech-only files, audio to text follows the same flow. For the boundary between what stays on your device and what does not, read on-device video privacy.
Frequently asked questions
Does mydubly keep a transcript after I download it?
No. Results, including transcripts, are deleted from the server within 30 minutes of the job finishing. The copy you downloaded is the one that remains, and it is under your control.
What happens to my audio if I close the tab mid-job?
The job cannot finish, because the final video is assembled in your browser. The unfinished job expires on the server after 24 hours and its data is deleted.
Is background conversation in my recording uploaded too?
Yes. The whole soundtrack is mixed down and sent in chunks, so anything audible, including people in the background, is part of the uploaded audio and may appear in the transcript.
Does mydubly hold a security or compliance certification?
mydubly does not claim any security or compliance certification. If your organization requires one before using a processor, treat that as a requirement to resolve before uploading.
Is it safer to transcribe an audio file than a video file?
The speech exposure is the same, because only audio is uploaded in both cases. The difference is that with a video, the picture stays on your device rather than being sent anywhere.