Why there is no single accuracy number
Accuracy in video translation is really three accuracies stacked on top of each other: did the recognizer write down the right words, did the translator carry the right meaning, and did the voice say it clearly at the right moment. Each can be measured, but each varies with the material. A benchmark score published for a model was measured on someone else's test audio, often read speech or curated clips, and tells you little about your conference recording with a humming projector and three accents.
That is why claims such as near-human accuracy should be read as marketing rather than measurement. The question to ask is not how accurate the tool is, but how accurate it is on videos like yours, and how many errors you can tolerate for the way the video will be used.
How recognition errors compound into translation
A translation engine has no access to the audio. It sees only the transcript, and it translates whatever is there with full confidence. A single misheard word therefore becomes a mistranslated word, and in a dubbed video, a confidently spoken wrong word.
Source audio: 'We pushed the launch to the third quarter.' Recognized in a noisy room as 'We pushed the lunch to the third quarter.' Spanish output: 'Trasladamos el almuerzo al tercer trimestre.' The Spanish is a perfect translation of the transcript and a wrong translation of the speaker. A Spanish-speaking viewer hears a sentence about moving lunch, with no hint that anything went wrong.
This is why the first place to look when a translation seems off is the source-language transcript. If the transcript is wrong, no amount of translation review will fix it; the remedy is better audio, or correcting the text by hand.
The factors that move accuracy up or down
- Audio quality
- Close microphones, little echo and no music under speech help most. Distant room mics, reverberant halls and loud backing tracks hurt most.
- Speaking style
- Scripted narration and paced lectures do well. Fast, mumbled, overlapping or heavily disfluent speech does worse.
- Vocabulary
- Everyday words are reliable. Product names, acronyms, people's names and niche jargon are the commonest errors.
- Language pair
- Widely spoken languages with plenty of training data in both directions generally do better than less-resourced ones.
- Accent and dialect
- Accents well represented in training data are recognized more reliably than less common ones.
- Content type
- Literal instructions survive translation; jokes, idioms, sarcasm and wordplay often don't.
The improving transcription accuracy guide covers what to change at recording time, and speech recognition and background noise explains why noise matters so much.
Why the language pair matters more than people expect
Some features of the target language force the translator to guess information the source never stated. Translating English 'I'm ready' into French, Hebrew, Polish or Russian requires choosing a grammatical gender for the speaker. Japanese and Korean require choosing a politeness level. Spanish, German and many others require choosing between formal and informal 'you'. A machine translating a 30-second stretch of speech has limited clues for these choices and may pick differently from one section to the next.
Writing systems also change how accuracy is measured. For Chinese and Japanese, which don't separate words with spaces, recognition accuracy is usually counted per character rather than per word.
How to measure accuracy on your own video
A practical evaluation takes three short excerpts and a fluent reader. It costs very little because transcript mode is billed at 1 credit per minute.
- Pick three excerpts of about two minutes each: your cleanest audio, your hardest (noise, fast speech, jargon), and something typical.
- Run the full video in transcript mode with translation switched on, so you get timestamped transcripts in both languages. A 40-minute video costs 40 credits, or 4 cents.
- For each excerpt, listen carefully and correct the source transcript into a reference. Count substitutions, deletions and insertions to get a word error rate.
- Ask a bilingual reviewer to grade each translated segment as correct, minor (awkward but meaning intact), major (meaning changed) or critical (misleading or potentially harmful).
- For every major or critical error, note whether the source transcript was already wrong at that point. That tells you whether to fix audio or review translation.
- If you plan to publish a voice track, dub one excerpt and listen for mispronounced names and rushed lines.
- Word error rate
- (substitutions + deletions + insertions) divided by the number of words in the reference.
- Example
- A 200-word reference with 6 substitutions, 2 deletions and 2 insertions has a WER of 10 / 200 = 0.05.
The word error rate article covers normalization details, such as ignoring punctuation and casing before counting.
Limits of quick accuracy checks
Short tests have blind spots. Three excerpts can miss a problem that appears only in the Q&A section or when a guest speaks. Correcting a machine transcript into a reference tends to bias you toward the machine's wording, so a truly independent transcript would score it slightly worse. Back-translation, translating the output back into your language with another tool, is a tempting check for people who don't read the target language, but it can hide errors that happen to reverse cleanly and invent errors that weren't there. Automatic metrics such as BLEU and COMET need professional reference translations and are designed for comparing systems over large test sets, not for judging one video; machine translation quality evaluation explains how they work.
Realistic expectations by kind of video
- Scripted narration recorded on a good microphone: typically needs only a light check of names and terms.
- Lectures and tutorials: generally solid on explanation, with technical terms and formulas read aloud needing attention.
- Interviews and podcasts: good when speakers take turns, weaker during interruptions.
- Casual vlogs full of slang, in-jokes and cultural references: expect literal renderings that miss the point.
- Panels with crosstalk, live events with crowd noise, and anything sung: expect substantial errors and plan for heavy review or subtitles only.
What affects accuracy in mydubly specifically
mydubly's recognition runs on Whisper large-v3-turbo, and the audio is cut at the quietest moment near each 30-second boundary so that words are not split between chunks. Each chunk's segments are translated by a neural machine translation engine, so the translator works with roughly half a minute of context at a time, enough for most sentences but not a guarantee that a term is rendered the same way throughout. The voice reads exactly what the translation says, so it can't correct anything upstream.
The practical upside is visibility. Every run returns transcripts in both languages, which makes the measurement method above straightforward, and testing in transcript mode first costs a fiftieth of a dubbed run. Common machine translation errors gives a checklist of what to look for when reading the output.
Test before you commit
Before translating a whole library, measure one representative video. Run it through the video translator in transcript mode, give the excerpts to a fluent reader, and decide from real errors on your own material whether you need light review, careful review, or a professional.
Frequently asked questions
Is AI video translation accurate enough for subtitles on a public video?
For clear speech with a quick review of names and terms, usually yes. For noisy, fast or jargon-heavy content, budget for a fluent reviewer to correct the subtitle file before publishing.
Why did my translation get a name completely wrong?
Most likely the speech recognizer misheard it and the translator then translated or respelled the wrong word. Check the source-language transcript at that timestamp; if it is wrong there, the error started with recognition.
Can I judge translation accuracy if I don't speak the target language?
Only roughly. You can check the source transcript yourself, which catches errors that start with recognition. For the translation itself, a fluent reviewer, even a colleague reading for twenty minutes, is far more reliable than back-translation.
Is the dubbed voice less accurate than the subtitles?
The words are the same translation, so meaning errors are shared. The voice can add its own problems, such as mispronounced names or rushed delivery when a line must be sped up to fit.
Does a better microphone really change translation quality?
Yes, often more than anything else you control. Cleaner audio means fewer recognition errors, and fewer recognition errors means fewer translation errors downstream.