Two problems that look like one
Live and recorded translation use the same building blocks, speech recognition, machine translation and sometimes speech synthesis, so they are easy to confuse. The difference is when the system must commit to an answer.
A live system has to produce output while the speaker is still talking. It never knows how a sentence will end. A file-based system receives the whole recording first, so it can look ahead, use the full context and take as long as it needs. That single difference explains most of the quality, cost and workflow gaps between them.
How a streaming pipeline works
- Audio is captured in short frames, often fractions of a second long, and sent to a recognizer continuously.
- The streaming recognizer emits partial hypotheses: a best guess at the words so far, which it may revise as more audio arrives. You see this when live captions change a word a moment after showing it.
- An endpointing step decides when a phrase or sentence is complete enough to translate, usually from pauses and the recognizer's confidence.
- The translation step translates that segment, sometimes re-translating as the source text is revised.
- The output is shown as live captions, or passed to speech synthesis and played to listeners.
- Live captions
- Lowest delay of the output options; readers can glance back at recent lines. Text flickers as words are revised.
- Synthetic speech
- Hands-free and works for people not looking at a screen. Adds delay, because a sentence must be translated before it can be spoken.
- Human interpreter
- Handles nuance, humor and context far better. Costs more, needs booking, and also works with a lag of a few seconds.
The latency and accuracy trade-off
Every live system tunes one dial: how long to wait before committing. Wait less and listeners hear the translation sooner, but the system guesses more, revises more and makes more mistakes at sentence boundaries. Wait more and output is steadier and more accurate, but the delay grows until conversation becomes awkward.
There is no setting that removes the trade-off, only choices about where to sit on it. Conversation tools tend to favor low delay because turn-taking breaks down otherwise. Lecture captioning can tolerate a little more lag in exchange for fewer corrections. Recognition itself also suffers from the short window: a misheard word early in a sentence cannot be fixed by context that arrives later if the translation has already been shown or spoken.
Why word order makes simultaneous translation hard
Languages put information in different places. German often places the main verb at the end of a clause, and Japanese and Korean put verbs last in general. Translating into English from those languages means the most important word of a sentence may arrive last. A live system must either wait for it, adding delay, or guess and risk being wrong.
Researchers study policies for this, such as waiting for a fixed number of source words before starting, or learning when to wait and when to commit. Human simultaneous interpreters face the same problem and handle it with anticipation, restructuring and knowledge of the topic, skills that machines approximate only partly. Speech itself adds difficulty, with false starts and fragments that the article on translating spoken language covers in detail.
What changes when the recording is finished
With a complete file, the system can segment speech by looking at the whole recording, use following sentences to resolve ambiguity, and produce a transcript and translation that a person can review before anyone else sees them. Timing can be planned too: a dubbed track can be fitted against the original speech rather than trailing behind it.
File-based processing still has to split long audio into manageable pieces, as described in audio chunking for speech recognition, but it can choose those cut points carefully, such as at quiet moments, rather than being forced by the clock. The cost is that nobody gets the translation during the event.
Choosing live or recorded translation
- A conversation, a negotiation, a customer call or travel: live tools, or a human interpreter if accuracy matters.
- A conference session with a multilingual audience in the room: live captions or interpreters for the moment, then translate the recording for publication.
- A webinar, lecture or talk that will be watched later: file-based translation of the recording, reviewed before posting.
- Legal, medical or safety communication: a qualified human interpreter or translator; machine output can support but not replace them.
- Content you need to search, quote or archive: a translated transcript from the recording.
An industry association runs a panel in English for an audience that includes Spanish and French speakers. During the event, a hired interpreter covers Spanish, and French speakers use a live captioning app on their phones. Afterward, the association translates the recording into both languages, corrects the panelists' names and figures, and publishes subtitled replays.
The live channels gave attendees access in the room; the recorded workflow produced something accurate enough to keep. The conference talk translation page describes the publishing side in more detail.
Translating an event afterward
- Record the main speech feed directly from the sound desk or meeting software, not from a microphone in the room.
- Export the recording with the main speech as the only or first audio track.
- Trim long silences, walk-in music and breaks before translating, or note where they are.
- Generate the transcript and translation, then check names, numbers and terminology with someone who attended or knows the subject.
- Publish subtitles, a translated transcript or a dubbed version, depending on how people will watch.
- Keep the original recording and reviewed text together for future reuse.
Guidance on getting clean audio from events is in transcribing live event recordings, and the specific issues with stream archives are in translating livestream recordings.
Limits of live machine translation
- Errors are shown or spoken before anyone can catch them, so mistakes reach the audience directly.
- Overlapping speakers, accents, jargon and poor room audio degrade live output faster than file-based output, which can at least use more context.
- Delay can make conversation feel stilted, especially when translated speech is synthesized.
- Live systems depend on a stable network connection in most setups; a dropout means a gap in translation.
- Recording, consent and confidentiality rules still apply to live captured audio. Check the policies of the event or organization, and get professional interpretation for high-stakes settings.
mydubly translates recorded files, not live speech
mydubly does not do live or real-time translation. It works on recorded files you already have: video in MP4, MOV, WebM, MKV or M4V, and audio in MP3, WAV, M4A, AAC, OGG or FLAC, up to two hours each. The audio translator and video translator transcribe the recording with Whisper, translate it into one of 21 languages, and give you a translated transcript and subtitles, or a dubbed voice track with one stock voice.
The browser processes the recording locally, splitting the audio into roughly 30-second chunks at quiet points and sending only compressed audio for processing; the original file stays on your device. Keep the page open until the job finishes, since leaving cancels it. For the 50-minute panel, a translated transcript and subtitles cost 50 credits (5¢) per language, and a dubbed version costs 2,500 credits ($2.50). Speakers are not labeled, so a panel transcript needs names added by hand.
Use live for the moment, files for the record
If people need to understand each other right now, use a live tool or an interpreter. If the words will be published, searched or kept, translate the recording afterward and review it. To see what that looks like for an existing recording, start with the audio translator.
Frequently asked questions
How much delay is normal for live speech translation?
It varies by system, language pair and settings. Live captions in the same language can appear within a second or two, while translated speech usually lags more because the system waits for enough of a sentence to translate and then synthesizes it. Human simultaneous interpreters also work a few seconds behind the speaker.
Why do live captions keep changing words after they appear?
Streaming recognizers show their best guess immediately and revise it as more audio arrives and context improves. That is a feature rather than a fault: the alternative is waiting longer before showing anything. Recordings translated afterward don't flicker, because the system has the whole sentence before producing output.
Can I use live translation captions as subtitles later?
You can, but they usually need substantial correction. Live captions reflect early guesses, may miss sentence boundaries and often don't align cleanly with the final recording. Generating subtitles from the finished recording is typically faster than repairing a live caption log, and the timing will match the published video.
Is live machine translation good enough for international meetings?
For informal meetings where participants can ask for clarification, it often helps. For negotiations, legal proceedings, medical consultations or anything where a misunderstanding has consequences, professional interpreters remain the safer choice. Many organizations use live machine captions as a supplement rather than a replacement.