AI voice & dubbing

How AI Voice Translation Turns Speech into Another Language

AI voice translation takes speech in one language and produces speech in another. Most working systems do it as a cascade: speech recognition writes down what was said, machine translation converts the text, and text to speech voices the result. Research models that translate speech directly into speech exist and are improving, but the cascade remains the practical choice when you need to check, correct and reuse the words along the way.

6 min read · Updated

Two ways to get from one spoken language to another

There are two broad architectures. A cascaded system breaks the problem into three well-understood tasks and runs a separate model for each. A direct, or end-to-end, speech-to-speech system trains one model, or a tightly coupled set of models, to map source audio to target audio without producing readable text in the middle.

Both start from the same raw material, a recording of someone talking, and both must cope with accents, background noise, disfluent speech and languages that order ideas differently. Where they differ is in what you can see and fix between the input and the output.

The cascaded approach: recognition, translation, synthesis

In a cascade, each stage hands plain text or timed text to the next.

  1. Automatic speech recognition transcribes the source audio into text with timestamps for each segment.
  2. Machine translation converts each segment into the target language, ideally with enough surrounding context to resolve pronouns and ambiguous words.
  3. Text to speech generates a new voice reading the translated text.
  4. For audio or video that must stay in sync, a timing step places each translated line where the original was spoken.

Each component can be developed, evaluated and replaced independently. Speech recognition has public benchmarks and models such as Whisper; translation engines are tested on text; voice models are tested by listening. The text in between is a natural checkpoint that people and software can inspect. How the voice stage works is covered in how AI voice generation works.

Direct speech-to-speech translation

Direct models try to skip the text. Google's Translatotron (2019) was an early demonstration that a single network could map a spectrogram in one language to a spectrogram in another. Later research, including Translatotron 2 and Meta's work on predicting discrete speech units rather than text, improved quality substantially. Meta's Seamless family of models, released from 2023, brought multilingual speech-to-speech translation to a wide audience as open research, with variants aimed at preserving expressive qualities of the speaker and at streaming.

The motivation is real. A cascade throws away everything that is not words: tone of voice, emphasis, hesitation, laughter. A direct model can in principle carry some of that across, and it can be faster because there is no hard handoff between stages. Many so-called direct systems still predict text internally as an auxiliary task, which shows how useful the written layer remains.

Comparing the two architectures

Inspectability
Cascaded: transcript and translation can be read, searched and corrected. Direct: little or no intermediate text to check.
Error behavior
Cascaded: recognition mistakes flow into translation and voice, but are visible. Direct: errors are harder to locate and attribute.
Expressiveness
Cascaded: original tone is lost at the text step unless added back. Direct: can preserve some prosody and vocal style.
Latency
Cascaded: three handoffs add delay. Direct: potentially lower, which matters for live use.
Flexibility
Cascaded: each stage can be swapped or upgraded, and outputs such as subtitles fall out for free. Direct: one model to train and replace as a whole.
Maturity
Cascaded: each stage is widely deployed. Direct: improving quickly, still largely research and specialist products.

Why cascaded systems remain practical

For recorded content that someone will publish, the cascade's advantages usually outweigh its losses. The written layer turns into useful deliverables: transcripts for search and accessibility, subtitles, a translated script for review. When a viewer reports a mistranslation, you can find the line in seconds rather than scrubbing through audio.

It also allows targeted quality control. If speech recognition is strong but the voice sounds flat, you change the voice; if a domain term keeps coming out wrong, you can see whether recognition or translation is to blame. Our explainer on translating spoken language describes why the text stage is where most meaning is won or lost.

Trade-offs and limits of a cascade

The cascade has three well-known weaknesses, and it is worth being honest about them.

  • Error propagation: a misheard word becomes a confidently translated and confidently spoken wrong word. Nothing downstream knows the input was wrong.
  • Lost paralinguistics: sarcasm, warmth, excitement and pauses for effect do not survive conversion to plain text, so the output voice expresses the generic emotion of the sentence, not the speaker's.
  • Fragmentation: speech arrives in segments that are not always full sentences, so translation can lack context and synthesis can sound choppy unless segments are regrouped.
  • Speaker identity: unless the system adds voice cloning or per-speaker voices, every speaker comes out in the same synthetic voice.
How one misheard number travels

Suppose a podcast host says "we raised fifteen thousand dollars in the first month." The recognizer writes "fifty thousand." The translation engine faithfully renders 50,000 in Spanish, and the voice reads "cincuenta mil dólares" with complete confidence. A listener has no reason to doubt it. In a cascade, though, both transcripts sit side by side, and a quick check of every number in the source transcript catches the slip before anything is published.

How mydubly's voice translation pipeline works

mydubly is a cascaded system and keeps the text at every step. In the browser, the file's audio is decoded and split into windows of about 30 seconds, each cut at a quiet moment so words are not sliced in half. Those speech chunks are transcribed with Whisper, translated by a neural machine translation engine, and voiced by Chatterbox Multilingual. Short fragments are merged into whole sentences before synthesis, and each line is fitted to the original timing with only a small tempo change before the clips are joined into one track.

You get the voice and the words. For audio files, a full run returns the translated audio file and transcripts in both languages; for video it also returns the translated MP4 and SRT and VTT subtitles. The spoken language is detected automatically, and you choose one of 21 target languages and one of 8 voices. A 30-minute podcast episode with full output costs 1,500 credits ($1.50), while transcript mode for the same file costs 30 credits. The audio translator handles MP3, WAV, M4A, AAC, OGG and FLAC; the video translator handles video files.

Checking a voice translation using the transcripts

Because the intermediate text is kept, you can review a translation far faster than by listening alone.

  1. Open the source-language transcript and scan for names, numbers and technical terms, the items recognition most often gets wrong.
  2. Compare those lines with the translated transcript to see whether any error came from recognition or from translation.
  3. Listen to the translated audio only around the timestamps you flagged, plus the first and last minute.
  4. If the source transcript is badly wrong throughout, improve the audio and re-run rather than correcting line by line.
  5. If only translation is off, share the translated transcript with a fluent reviewer; the post-editing guide explains how deep to go.

Where to go next

If your content is audio-first, such as a podcast, interview or voice memo, start with the audio translator and run one episode in the language your audience asks for most. For a deeper look at the voice side, read AI dubbing vs human dubbing.

Frequently asked questions

Is AI voice translation the same as real-time interpretation?

No. Live interpretation translates as someone speaks, which demands very low latency and usually accepts rougher output. File-based voice translation processes a finished recording, so it can use more context and produce a cleaner, timed result. mydubly is file-based only and does not translate live audio.

Does AI voice translation keep the original speaker's voice?

Only if the system includes voice cloning or a direct model designed to preserve vocal style. Most cascaded tools, including mydubly, voice the translation in a chosen stock voice. mydubly offers 8 preset voices and does not clone anyone.

Why not just translate speech directly without text?

Direct models can preserve tone and reduce latency, but they leave little readable text to check and correct, and they are harder to debug. For content that will be published, being able to read and fix the transcript and translation is usually worth more than the extra expressiveness.

What kinds of audio translate worst?

Overlapping speakers, heavy music beds, strong reverberation and very fast or mumbled speech hurt recognition, and those errors carry through every later stage. Clear single-speaker recordings translate most reliably. If several people talk, see translating videos with multiple speakers.

Can I get just the translated text without a voice?

Yes. Transcript mode returns a timestamped transcript and subtitles in the spoken language, optionally translated, without generating a voice. It costs 1 credit per minute, with a minimum of 5 credits per file.