Why a taxonomy helps reviewers
Professional evaluation frameworks such as MQM (Multidimensional Quality Metrics) classify translation errors into categories like accuracy, fluency, terminology, style and locale conventions, each with a severity. You don't need the full framework to benefit from the idea. Reading with a short list of error types in mind turns a vague "this feels off" into specific checks: are all the numbers right, did any negation disappear, is the product name intact.
The categories below are ordered roughly by how much damage they do. The scoring side of this, including how to turn error counts into a quality judgment, is covered in machine translation quality evaluation.
Accuracy errors: mistranslation, omission and addition
Mistranslation is the core accuracy error: a word or phrase translated with the wrong meaning, usually the wrong sense of an ambiguous word.
Source: "The patient was discharged on Monday." A German translation using "entladen" (to discharge electrically, to unload) instead of "entlassen" (to release from hospital) is grammatical, fluent and nonsensical. A reviewer who skims for flow may never notice.
Omission is when part of the source disappears. Short function words are the dangerous ones. "Do not restart the device during the update" rendered in French as "Redémarrez l'appareil pendant la mise à jour" has lost "ne... pas" and now instructs the opposite. Omissions are more common in long, clause-heavy sentences.
Addition is the reverse: content that was never in the source. It ranges from a harmless intensifier to a whole invented clause, and it is more frequent when the input is garbled or when a general-purpose language model "helpfully" explains something. Additions are hard to catch without comparing against the source line by line.
Terminology errors and inconsistency
A terminology error uses a correct general-language word where the field has an established term, or translates a term that should have stayed in the source language. In a software tutorial, "dashboard" might appear in German as "Dashboard" (the term used in the actual interface), "Armaturenbrett" (a car's dashboard) and "Übersicht" (overview), all in the same video. Each choice is defensible in isolation; together they confuse viewers trying to follow along on screen.
Inconsistency is the more common symptom, because many engines translate short passages independently. Building a glossary and checking terms against it is the cure; translating names and technical terms walks through it.
Literal idioms and false friends
Idioms are phrases whose meaning is not the sum of their words. Machine translation handles frequent ones well and rarer ones literally. "We're on the same page" becomes French "Nous sommes sur la même page", which is understandable but unnatural; a French speaker would say "Nous sommes sur la même longueur d'onde". German "Ich verstehe nur Bahnhof" literally means "I only understand train station"; the meaning is "It's all Greek to me".
False friends are words that look alike across languages but differ in meaning. Spanish "embarazada" means pregnant, not embarrassed. Spanish "actualmente" and French "actuellement" mean currently, not actually. German "eventuell" means possibly, not eventually. French "librairie" is a bookshop, not a library. Modern engines rarely fall for the famous cases in clear contexts, but short or ambiguous lines still trip them up, especially in the less common direction of a language pair.
Numbers, units, dates and currencies
Numbers seem safe because they look the same in every language. They are not.
- Separators: English "1,500.75" is "1.500,75" in German and "1 500,75" in French. Copying the English format can make a figure read as a thousand times smaller.
- Scale words: English "billion" is "mil millones" in Spanish; Spanish "billón" means a million million, so a literal swap inflates the figure enormously.
- Dates: "03/04/2026" is March 4 in the United States and 3 April in most of Europe.
- Units and currencies: whether to convert miles to kilometers or dollars to euros is a localization decision, and engines are inconsistent about it.
Numbers in speech add another layer, since the recognizer must first decide whether "fifteen" or "fifty" was said.
Named entities: people, brands and places
Names should usually pass through unchanged or be transliterated, never translated as words. Engines sometimes translate them anyway: "Mr. Baker" becoming "Herr Bäcker", a company called Notion rendered as the word for "idea", "Apple" treated as fruit when the context is thin. Place names are the opposite trap: some have established exonyms (Munich is München, Lisbon is Lisboa), and leaving those untranslated, or translating ones that have none, both read as errors. For non-Latin scripts the question becomes how to transliterate, which is where inconsistency creeps in.
Register, tone and gender bias
Register errors put the right meaning in the wrong voice. A relaxed "Hey everyone, quick one today" can come out stiff and formal, and a formal policy statement can come out chatty. In languages with formal and informal "you", switching between them mid-text is a register error even when each sentence is correct.
Gender bias is a related problem. When the source does not mark gender, engines pick the form most frequent in training data, so "my friend is a surgeon" tends to become masculine in Spanish ("mi amigo es cirujano") and "the nurse" feminine, regardless of who is meant. For speaker-gender agreement, see context in machine translation.
The limits of a quick review: errors that slip through
Not every category is equally easy to spot, and it helps to know what a fast read will miss:
- Easy to catch without the source
- Broken grammar, untranslated words, garbled names, obvious register clashes.
- Need the source side by side
- Omissions, additions, wrong sense of an ambiguous word, swapped numbers.
- Need domain knowledge
- Terminology choices, established exonyms, units and conventions.
- Need the audio or video
- Speaker gender, who is addressed, what "this" refers to on screen.
A monolingual reader in the target language catches the first row and little else. For anything published, someone must compare against the source.
Spotting these errors in mydubly output
In a speech pipeline, some apparent translation errors started earlier, in recognition. mydubly gives you transcripts in both languages, which makes the split visible: if the source-language transcript already says "fifty" where the speaker said "fifteen", the translation faithfully reproduced a recognition error. If the source transcript is right and the translation is wrong, the translation step is responsible. Whisper models have also been widely reported to produce stray phrases like "Thank you for watching" over long silences or music, which then get translated like any other line.
Because the SRT and VTT files are timestamped, you can jump from a suspect subtitle to the exact moment in the video and listen. Keep in mind that on-screen text in the picture is not translated at all; that is a gap to plan for, not a translation error. The video translator and subtitle generator both produce the paired transcripts used for this kind of check.
Next step
Take a translated video, open both transcripts, and do three passes: numbers and names first, then negations and missing clauses, then terminology and register. Fix what you find using the post-editing workflow. If you have not translated anything yet, start with a short clip in the video translator; a two-minute dub costs 100 credits, or 10 cents.
Frequently asked questions
Which machine translation errors are the most serious?
Accuracy errors that change meaning: dropped negations, wrong numbers, mistranslated instructions and invented content. These can mislead or endanger viewers. Style and register problems are usually less serious, though they still undermine trust in published material.
Why does machine translation get idioms wrong?
Engines learn idioms from examples. Common idioms appear often enough in training data to be translated as units, but rarer ones look like ordinary word sequences, so the engine translates each word. The result is grammatical but literal.
How can I tell whether an error came from speech recognition or from translation?
Compare the source-language transcript with the audio. If the source transcript is already wrong at that point, the error came from recognition and the translator simply carried it over. If the source transcript is right but the translation is wrong, the translation step made the mistake.
Do machine translation systems still confuse false friends?
Much less often than older systems, because context usually makes the intended meaning clear. Errors still appear in short lines with little context and in less common language pairs, so it is worth checking the well-known false friends for your pair.
Should units like miles and pounds be converted in a translation?
It depends on the content. A recipe or a safety instruction for a European audience usually should be converted; a historical quote or a product specification often should not. Decide on a policy before review rather than accepting whatever the engine did in each line.