Where the time goes in an AI translation
An AI run is a sequence of stages, and each scales differently with the length of the video.
- Audio extraction
- Done on your device. Depends on file length and device speed, but involves decoding audio only, not the picture.
- Upload
- Depends on your connection and on how much audio there is, not on the size of the video file.
- Recognition and translation
- Work through the audio in short chunks, several of which can be in flight at once.
- Voice generation
- Only in a dubbing run. Usually the longest stage, because every translated sentence becomes new audio.
- Timing fit and reassembly
- Fitting lines to the original timing, then combining the new track with the video in the browser.
- Review
- Human time. Often longer than everything above combined.
File length: the main driver
Because the work is done chunk by chunk, a longer file simply means more chunks. A 10-minute video is about 20 windows of roughly 30 seconds; a 90-minute lecture is about 180. Parallel processing means doubling the length does not necessarily double the wait, but you should expect long files to take noticeably longer than short ones, especially with a voice track. For files at the long end, read translating long videos, which covers splitting and keeping the browser tab alive.
Upload size: why your connection matters less than you think
Uploading a raw video file is slow on most connections. Uploading only compressed audio is not. In mydubly each 30-second chunk is encoded to Opus at 32 kb/s for a transcript, about 120 KB, or at 48 kb/s, about 180 KB, for a dub, because dubbing chunks also carry the music and effects that stay under the new voice. Either way the arithmetic is friendly.
Suppose a 60-minute recording is a 3 GB MP4 and your upload speed is 10 megabits per second. Sending the whole file would mean 24,000 megabits, about 40 minutes of uploading before anything else could start. Sending only the audio for a transcript means about 120 chunks of 120 KB, roughly 14 MB or 115 megabits, which takes around 12 seconds at the same speed (a dub's larger chunks come to about 22 MB, around 17 seconds), and those chunks go up in parallel while the rest of the pipeline gets started.
This is also why weak mobile connections are workable for translation, and the article on reducing upload size goes through the numbers in more detail.
Why voice generation is usually the slowest stage
Recognition and translation produce text, which is compact and quick to generate. A voice model has to produce audio, piece by piece, for every sentence in the video, and the result then has to be normalized, placed against the original timing and joined into a continuous track. mydubly also merges short fragments into whole sentences before synthesis so the delivery sounds natural, which is better for quality but means each synthesis call handles a full sentence. If you only need subtitles and transcripts, transcript mode skips this stage entirely, which is the fastest way to get a usable translation; the subtitle generator is built around that output.
Review time is the real bottleneck
Machine output can be ready before your reviewer has opened their inbox. A fluent reader needs roughly as long as the video to read a translated transcript carefully, more if they are checking against the source line by line or correcting a subtitle file as they go. If your plan involves a colleague in another time zone, their calendar, not the software, decides your release date. Schedule the review when you schedule the translation, and give the reviewer a short list of terms and names to check first so they don't spend an hour on style. Machine translation post-editing explains how to scope that work as light or full.
How human translation timelines compare
A professional workflow adds steps that each take calendar time: quoting, assigning a translator, translation, a second linguist's review, subtitle spotting or script adaptation, and for dubbing, casting, recording sessions, mixing and quality control. Subtitles from a vendor commonly take days rather than hours, and dubbing with voice actors takes longer still, but timelines vary so much between vendors, languages and volumes that the only reliable number is the one in your vendor's written schedule. Rush service usually exists and usually costs more. The trade-offs beyond time are covered in AI vs human video translation.
Pitfalls that add hours or days
- Translating before the edit is locked, then translating and reviewing again after changes.
- Closing the browser tab, or letting a laptop sleep, while the job is still being prepared in the browser.
- Not downloading results promptly; in mydubly they are deleted within 30 minutes of the job finishing.
- Files over the 2-hour limit that have to be split at the last minute.
- Waiting on a reviewer who was never told the job was coming.
- Discovering in review that the source audio is the problem, which only re-recording or careful subtitle correction can fix.
Planning a release around translation
Work backward from your publish date.
- Lock the edit. Every change after translation means paying for and reviewing the changed version again.
- Book your reviewer for a specific day, and send them the glossary of names and terms in advance.
- Run the translation as soon as the edit is final, and download every output straight away.
- Give the reviewer the translated transcript and subtitle files, and the dubbed video if you made one.
- Apply subtitle corrections, and decide whether any voiced errors warrant a re-run or a different approach.
- Keep a buffer of at least one working day for a re-run, a second language, or a reviewer who runs late.
Suppose a creator publishes a 20-minute episode every Thursday and wants Spanish and Portuguese versions on the same day. The edit locks on Monday evening and both dubbed versions are run that night, 1,000 credits each, $2.00 in total. Reviewers read the transcripts on Tuesday, corrections go into the subtitle files on Wednesday morning, and Wednesday afternoon is held free for a re-run if one is needed.
Turnaround in mydubly
mydubly processes files of up to 2 hours entirely as file-based jobs; nothing is live or real-time. Your browser extracts the audio, uploads compressed audio chunks in parallel and assembles the finished video at the end, so keep the tab open until downloads are offered. Transcript mode returns subtitles and transcripts without generating a voice; full output adds the translated voice track and the translated MP4. Unfinished jobs expire after 24 hours, and finished results are removed within 30 minutes, so treat the download step as part of the job. Start with the video translator and time a short clip yourself; that is the most reliable estimate for your device and connection.
Frequently asked questions
Is AI video translation instant?
No. It is fast compared with a human workflow, but each stage takes processing time, and a dubbed voice track takes longer than subtitles. It is also file-based, so it cannot translate a live stream as it happens.
Does a bigger video file take longer to upload to mydubly?
Not much. The video file stays on your device, and only compressed audio is uploaded, so upload time depends on the length of the audio rather than on resolution or file size.
Can I close the tab while my video is being translated?
Keep it open. The browser extracts and uploads the audio and assembles the final video, so closing the tab or letting the device sleep can interrupt the job before you get your files.
How much time should I allow for reviewing a translation?
Plan for at least the running time of the video for a careful read of the translated transcript, and more if the reviewer is correcting subtitle files or checking against the source line by line.
Why do subtitles come back sooner than the dubbed version?
Subtitles only need recognition and translation, which produce text. A dub also generates speech for every sentence and fits it to the original timing, which is the heaviest part of the job.