AI translation

Why spoken language is harder to translate than written text

Translating spoken language is harder than translating written text because speech was never designed to be read. People restart sentences, leave thoughts unfinished, chain clauses together with 'and so', and rely on tone and shared context instead of punctuation. A machine translator trained mostly on edited writing receives all of that through a speech recognizer, which adds its own mistakes, and then has to choose a level of politeness the speaker never had to state. Knowing where these problems come from tells you what to fix in the recording, what to check in the output, and what no tool will solve for you.

7 min read · Updated

Speech was never written to be read

Written text is planned. A writer revises, punctuates and puts each idea into a complete sentence. Speech is produced in real time, so it carries the traces of thinking aloud: hesitations, corrections, half-finished clauses and constant reliance on the listener to fill gaps. Listeners handle this without noticing. Machine translation systems, by contrast, learned from parliament proceedings, news, documentation and subtitles, which are mostly tidy, sentence-shaped text.

In a typical speech translation pipeline the audio is transcribed first, then the transcript is translated, then the result is shown as subtitles or spoken by a synthetic voice. Every characteristic of speech described below hits the middle step, translation, in a form it was not trained for.

Disfluencies: fillers, restarts and self-corrections

Fillers like "um", "uh", "you know" and "like" carry little meaning, and most modern recognizers drop many of them. OpenAI's Whisper, for example, tends toward clean output rather than strict verbatim. Restarts and self-corrections are harder, because they contain real words.

Before and after: a self-correction

Spoken: "So we, um, we shipped it on Tuesday, no, sorry, Wednesday, and the client, the client was happy." A transcript that keeps everything gives the translator two shipping days and a repeated subject. A faithful French rendering of the intended meaning is "Nous l'avons livré mercredi, et le client était content." A literal translation of the raw transcript keeps both days and the stammered "le client, le client", and a viewer reading subtitles at speed may simply remember "Tuesday".

Interpreters handle this by translating the speaker's intention rather than their words. Machine translation translates the words. When a recording is full of repairs, the cleanest fix is upstream: edit the transcript before relying on the translation, or re-record a scripted version for anything important.

Fragments and run-on sentences

Spoken sentences rarely match written ones. Some are fragments: "Bigger budget, smaller team. Great." Others never end: "And then we moved the launch and so marketing had to redo everything and honestly that was fine because the first version wasn't working anyway."

The recognizer has to invent punctuation for both. Where it places full stops decides what the translator treats as a unit. A fragment with no verb may be expanded into a full sentence with a guessed verb, or translated word by word into something that reads like a list. A run-on sentence pushes the translator into long-distance dependencies that raise the chance of dropping a clause. Languages with very different word order, such as Japanese or Korean from English, are hit hardest, because the translator must reorganize an entire clause chain to produce natural output.

How recognition errors travel into the translation

A speech translation pipeline is a cascade: each stage trusts the one before it. If the recognizer mishears a word, the translator receives clean-looking text with no sign that anything went wrong, and translates it fluently.

  • Negations and contractions: "You can't upload files over two gigabytes" heard as "You can upload files over two gigabytes" becomes a confident, opposite statement in every target language.
  • Similar-sounding numbers: "fifteen" and "fifty" are easy to confuse in fast or accented speech, and the translator will dutifully write "cincuenta".
  • Names and jargon: an unfamiliar surname or product name is often transcribed as ordinary words, which the translator then translates as ordinary words.
  • Homophones: "accept" and "except", or "affect" and "effect", change meaning entirely.

This is why reviewing the source-language transcript matters as much as reviewing the translation. The guide to improving transcription accuracy covers recording habits that reduce these errors at the source.

Register, politeness and who is speaking

English hides a lot of social information. "Can you send me the file?" works between friends and with a new client. Many target languages force a choice: tu or vous in French, du or Sie in German, tú or usted in Spanish, aap, tum or tu in Hindi, and several distinct speech levels in Japanese and Korean. The translator has to infer the relationship from the words alone, usually one passage at a time, and casual spoken English often comes out either stiffly formal or inappropriately familiar.

Some languages also mark the speaker's gender in the grammar. In Hindi, a man says "main thak gaya hoon" and a woman "main thak gayi hoon" for "I'm tired"; Hebrew, Arabic, Russian and Polish have similar agreement in various forms. A text translator does not hear the voice, so it falls back on defaults, and in an interview with several speakers it cannot tell who is saying which line. How engines handle these choices across longer stretches is covered in context in machine translation.

Numbers, times and spoken shorthand

Spoken numbers are ambiguous in ways written ones are not. "Fifteen fifty" might be a time (15:50), a price ($15.50), a year or a quantity. British "half eight" means 8:30, which a literal translation can turn into "half of eight". Speakers say "a couple of k", "Q3", "the 101" or "twenty-four seven", and the recognizer has to decide whether to write digits, words or abbreviations before the translator ever sees them. Each choice changes what the translator produces, so numbers deserve a dedicated pass during review.

Where speech translation works well, and its limits

It works well when the recording resembles written language:

  • Scripted narration, explainers and tutorials with one clear speaker.
  • Lectures and talks, where speakers tend to use complete sentences and define terms.
  • Recordings with a good microphone and little background noise or music.

It struggles when speech is most conversational:

  • Overlapping speakers and rapid back-and-forth banter.
  • Heavy slang, in-jokes and references that depend on shared history.
  • Switching between languages mid-sentence, which confuses both recognition and translation.
  • Music or crowd noise under the voice, which raises recognition errors that then cascade.

Preparing a recording for translation

If you control the recording, a few habits improve results more than any setting:

  1. Speak in complete sentences and pause briefly between them, so the recognizer places full stops where you intended.
  2. Say names, product terms and acronyms clearly the first time, and repeat the noun instead of "it" or "they" after a long digression.
  3. Avoid crosstalk; in interviews, let each person finish before the other starts.
  4. Keep music and effects well below the voice; loud beds raise recognition errors and make the voice harder to separate from the background.
  5. State numbers unambiguously: "three thirty in the afternoon" rather than "three thirty".
  6. Read the source transcript before the translation, and correct misheard words in your own notes so you know which translated lines to check.

How mydubly handles spoken input

mydubly is built around speech rather than documents. The browser decodes your file's audio and splits it into windows of about 30 seconds, placing each cut at the quietest moment in the last few seconds before the mark so words and sentences are not chopped in half. Whisper large-v3-turbo transcribes each window into timestamped segments, and a neural machine translation engine translates them. The spoken language is detected automatically; recordings that switch languages often are worth checking closely.

You get transcripts in both languages, so you can separate recognition errors from translation errors, along with SRT and VTT subtitles and, if you choose a voice, the translated audio. Transcripts do not label speakers, which matters for interviews: keep your own note of who speaks when. The audio translator accepts MP3, WAV, M4A, AAC, OGG and FLAC files up to two hours long, and audio to text gives you the transcript alone if you want to clean it before deciding on translation. For specific formats, see podcast translation and interview transcription.

Where to go next

Try a real recording rather than a test phrase. Suppose you have a 45-minute podcast episode: a transcript costs 45 credits (under five cents), and a full translated voice track costs 2,250 credits, or $2.25. Start with the audio translator, read the source transcript first, then the translation, and you will see exactly which of the problems above your recording has.

Frequently asked questions

Should I remove filler words before translating a recording?

Fillers like 'um' and 'uh' are usually dropped by the recognizer already. Restarts and self-corrections are the ones worth handling, because they contain real words that get translated. For important content, edit those out of the transcript or re-record the passage.

Why does the translation of a casual conversation sound so formal?

Many languages require choosing a politeness level that English leaves unstated, and translation engines often default to the neutral or formal option when the text doesn't make the relationship clear. Short, chatty lines give the engine little to go on, so the default wins.

Can speech translation handle two people talking over each other?

Poorly. Overlapping speech degrades recognition first, often merging or dropping words from both speakers, and the translation inherits those gaps. Recording each speaker on a separate microphone, or asking guests not to interrupt, helps more than anything done afterwards.

What happens when a speaker switches languages mid-sentence?

Recognizers typically expect one language per stretch of audio, so a switch may be transcribed phonetically in the wrong language or skipped. The translator then works from that flawed text. Expect to review code-switched passages by hand.

Is it better to translate the audio directly or to clean the transcript first?

For casual listening, translating directly is fine. For anything published or relied upon, cleaning the source transcript first gives the translator well-formed sentences and correct names, which improves the translation more than fixing its output afterwards.