Two questions evaluation tries to answer
Every evaluation method is really answering one of two questions. The first is comparative: is engine A better than engine B, or this month's model better than last month's, on a given kind of text? Automatic metrics were designed for this, because researchers need to score thousands of sentences many times a day. The second is absolute: is this particular translation good enough for its purpose? That requires someone to look at the output, because "good enough" depends on the audience and the stakes.
Two classic dimensions run through both. Adequacy (or accuracy) asks whether the meaning of the source is preserved. Fluency asks whether the output reads naturally in the target language. Modern engines are usually fluent, so most of the useful information is in adequacy.
BLEU: counting overlapping words
BLEU, introduced in 2002, measures how many word sequences (n-grams) in the machine output also appear in a human reference translation. It computes the precision of matching single words, pairs, triples and four-word sequences, combines them with a geometric mean, and multiplies by a brevity penalty so that very short outputs cannot score well by being cautious. Scores are usually reported on a 0 to 100 scale and are meaningful at the level of a whole test set, not a single sentence.
Reference: "the cat sat on the mat". Machine output: "the cat is on the mat". Unigram precision: five of the six output words appear in the reference ("is" does not), so 5/6. Bigram precision: of the five output pairs (the cat, cat is, is on, on the, the mat), three appear in the reference, so 3/5. Trigrams and four-grams score lower still, and the combined BLEU drops noticeably for a sentence any reader would call nearly correct.
That example shows BLEU's main weakness: it rewards surface overlap, so a correct synonym or a valid reordering is penalized, especially with only one reference. Scores also depend on tokenization, which is why the SacreBLEU tool was created to standardize how BLEU is computed and reported. BLEU remains common because it is cheap, transparent and reproducible.
chrF: matching characters instead of words
chrF computes an F-score over character n-grams (sequences of up to six characters) rather than words, weighting recall more heavily than precision. Matching at the character level gives partial credit for words that are almost right. If the reference has the German "des Hauses" and the output has "das Haus", word-level BLEU sees two misses, while chrF credits the shared characters.
That makes chrF more informative for morphologically rich languages such as Finnish, Czech, Turkish or Polish, where a single word can take many inflected forms. It shares BLEU's dependence on references and its blindness to meaning: a fluent sentence with a dropped "not" can still score well.
COMET and other learned metrics
COMET belongs to a newer family of metrics that are themselves neural networks. It encodes the source sentence, the machine output and the reference with a multilingual language model, and predicts a quality score. It was trained on large sets of human quality judgments, so it learns to reward meaning-preserving paraphrases that BLEU would punish. In the annual WMT metrics shared tasks, learned metrics like COMET have generally agreed with human judgments more closely than BLEU has. BLEURT is a similar approach.
Some variants work without a reference at all. These quality estimation models, CometKiwi for example, look only at the source and the output, which makes them usable on live content where no human translation exists. The trade-offs are opacity and calibration. A COMET score has no intuitive unit, scores from different model versions are not comparable, and learned metrics can be fooled by outputs that are fluent but subtly wrong.
Human evaluation: MQM and direct assessment
MQM (Multidimensional Quality Metrics) is an annotation framework. Trained reviewers read the source and the translation, mark each error span, assign a category (accuracy, fluency, terminology, style, locale conventions and their subtypes) and a severity. A common scheme weights a minor error as 1 point and a major error as 5, with critical errors treated as disqualifying or weighted more heavily; the penalty total is normalized by text length. MQM is slow and needs skilled bilingual annotators, but it tells you not just how good a translation is but exactly what is wrong with it.
Direct assessment asks raters to score each translation on a 0 to 100 scale for how well it conveys the source meaning, often with quality-control items mixed in to detect careless raters. It is faster than MQM and was used for many years in WMT evaluations, but scores vary between raters and say nothing about error types. Pairwise ranking, where raters pick the better of two outputs, is a simpler alternative for comparing systems.
What scores can't tell you
Automatic metrics have hard limits worth keeping in mind:
- They need reference translations, which you almost never have for your own video or podcast.
- Scores are not comparable across languages, test sets or tokenizations. A BLEU score for English-German says nothing about English-Japanese.
- Small differences between systems are often within noise; serious comparisons use significance testing.
- None of them measure whether a subtitle can be read in time, whether a dubbed line sounds rushed, or whether a term matches what is shown on screen.
- A high average can hide a single catastrophic error, such as a reversed safety instruction, which matters more than everything else combined.
A practical spot check for translated video
When you need to decide whether a specific translated video is fit to publish, a lightweight version of MQM works well:
- Pick three samples of one to two minutes each: the opening, the densest technical stretch, and the fastest or most conversational part.
- Find a reviewer fluent in both languages. Give them the source transcript, the translated transcript and the video.
- Have them mark every error with a category (accuracy, terminology, fluency, style, numbers and names) and a severity: minor (awkward but clear), major (meaning changed or confusing) or critical (could mislead or cause harm).
- Separately check that subtitles are readable at their timing and, for dubbed audio, that lines don't sound rushed or cut off.
- Decide the thresholds before looking at results. For example: any critical error means fix and recheck; more than one major error per sampled minute means full post-editing; otherwise a light pass is enough.
The reviewer samples five minutes in total. They find one major error (a figure written as "1,500" for a Spain-based audience, where the comma marks decimals and the figure reads as one and a half), three minor terminology inconsistencies ("panel" and "tablero" both used for "dashboard") and no critical errors. Under the thresholds above, that is a light post-edit: fix the number format everywhere, standardize the term, and publish. The full dubbed translation for 30 minutes costs 1,500 credits ($1.50), so the review time is the main cost.
The errors you find map onto the categories in common machine translation errors, and the fixes onto the post-editing guide.
Evaluating mydubly output
mydubly's pipeline gives you what this kind of review needs. Every job returns timestamped transcripts in the spoken language and the target language, plus SRT and VTT subtitles, alongside the translated video and audio. A reviewer can line up segments by timestamp and check meaning directly, and can tell recognition errors (already wrong in the source transcript) from translation errors (source right, translation wrong).
Translation is done by a neural machine translation engine that is configurable in the deployment, so the sensible benchmark is your own sample rather than a published score for some named model. If you do have professional reference translations for a few minutes of material, you can compute chrF or COMET yourself against the downloaded translated transcript using the open-source tools. The subtitle generator produces the same paired transcripts without a voice track, at transcript pricing, if you want to evaluate translation quality before paying for dubbing.
Next step
Choose one real video, translate it with the video translator, and run the five-step spot check above with a bilingual colleague. Write down the error counts and your thresholds; after a few videos you will have a quality baseline for your content and language pairs that is far more useful than any public leaderboard.
Frequently asked questions
What is a good BLEU score?
There is no universal threshold. BLEU depends on the language pair, the test set, the number of references and the tokenization, so a score only means something next to other systems measured the same way. Use it to compare engines on your own test set, not to judge a translation in isolation.
Is COMET more reliable than BLEU?
For judging whether meaning is preserved, generally yes: learned metrics like COMET have tracked human judgments more closely than BLEU in recent WMT metrics evaluations. BLEU is still useful because it is simple, transparent and stable across versions, and many teams report both.
Can I evaluate a translation without a reference translation?
Yes, in two ways. Quality estimation models score a translation from the source and output alone, and human review frameworks like MQM need only a bilingual reviewer. For a single video, a structured human spot check is usually more informative than any automatic score.
How many minutes of a translated video should I review?
Enough to cover the different kinds of speech in it. Three samples of one to two minutes (the opening, the most technical part and the most conversational part) catch most systematic problems. Raise the sample size for content where errors carry legal or safety consequences.
What do MQM severity levels mean in practice?
Minor errors make the text slightly awkward without changing meaning. Major errors change meaning or would confuse a reader. Critical errors could mislead someone into a harmful action or cause legal or reputational damage. Severity drives the decision, so agree on examples of each before reviewing.