Speech recognition & transcription

Word error rate explained, with a worked example

Word error rate (WER) is the standard measure of speech recognition accuracy: the number of word substitutions, deletions and insertions needed to turn a machine transcript into a correct reference transcript, divided by the number of words in the reference. A WER of 0.10, or 10 percent, means roughly one error for every ten reference words. It is simple and comparable, but it depends heavily on how the text is normalized, and it treats every word as equally important.

7 min read · Updated

The WER formula

WER = (S + D + I) / N, where each term is counted after aligning the machine transcript, called the hypothesis, with a human-checked reference transcript.

S, substitutions
Reference words replaced by a different word in the hypothesis
D, deletions
Reference words missing from the hypothesis
I, insertions
Extra words in the hypothesis that are not in the reference
N, reference words
The number of words in the reference transcript, which is the denominator

Lower is better, and zero means a perfect match. Because insertions are counted while the denominator stays fixed, WER can exceed 1.0, or 100 percent, when a system adds many words that were never spoken, as happens when a model hallucinates through a long silence. Accuracy is sometimes quoted as 1 minus WER, but that figure can go negative for the same reason, so WER itself is the cleaner number to report.

A worked example

Reference, 10 words

please send the quarterly report to Anna before Friday morning

Hypothesis

send a quarterly report to Ana before Friday in the morning

Counting the errors

One deletion ("please" is missing), two substitutions ("the" became "a" and "Anna" became "Ana") and two insertions ("in" and "the"). WER = (2 + 1 + 2) / 10 = 0.50, or 50 percent.

That score looks alarming for a transcript whose meaning is almost intact. Most of the damage comes from a dropped politeness word and a rephrased time expression, while the one error that could actually matter, the misspelled name, counts exactly the same as "the" becoming "a". This is also a very short sentence, so each error carries a lot of weight; across an hour of speech the same kinds of errors are spread over thousands of words.

  1. Normalize both texts the same way, as described below.
  2. Split each into a list of words.
  3. Find the alignment that needs the fewest edits.
  4. Count substitutions, deletions and insertions.
  5. Divide the total by the number of reference words.

How the alignment is found

Counting errors by eye is unreliable, because there are often several ways to line up two sentences. WER uses the alignment with the smallest total number of edits, computed with the same dynamic-programming algorithm as Levenshtein edit distance, applied to words instead of characters. Tools then trace back through the table to label each edit as a substitution, deletion or insertion.

The minimum-edit rule has a subtle consequence: a substitution counts as one error, while the equivalent deletion plus insertion counts as two, so the algorithm prefers substitutions. When two alignments tie on total edits, different tools may split the count differently among S, D and I, but the WER comes out the same.

Text normalization changes the score

Before comparing, both transcripts must be normalized the same way, or formatting differences get counted as recognition errors. Is "Dr." the same as "doctor"? Is "10" the same as "ten", or "don't" the same as "do not"? A modern model that writes "$25" against a reference that says "twenty five dollars" would otherwise be charged one substitution and two deletions for a perfectly correct recognition.

  • Lowercase everything.
  • Remove punctuation, or decide explicitly which marks to keep.
  • Expand or standardize numbers, currencies, dates and common abbreviations.
  • Standardize spelling variants such as "color" and "colour".
  • Decide how to treat fillers like "um" and "uh", which some systems drop on purpose.

The Whisper paper made this point strongly: OpenAI released a text normalizer alongside the model because Whisper's formatted output was being penalized against references written in other conventions. The lesson for anyone comparing systems is that WER figures computed with different normalizers are not comparable, and a published number should always say which normalization was used.

Character error rate for Chinese and Japanese

WER assumes words are separated by spaces. Chinese and Japanese are written without spaces between words, and splitting them into words requires a segmentation tool whose choices would themselves change the score. The usual answer is character error rate (CER): the same formula computed over characters instead of words.

CER numbers are not directly comparable with WER numbers. One wrong word in English is one error; in Chinese, a wrong two-character word may count as one or two character errors depending on how much of it is wrong. Thai, which also lacks spaces, is often scored with CER too. Korean puts spaces between word groups, so both measures appear in published work, and it's worth checking which one a figure uses. CER can also be useful in any language when you want a measure that is gentler on small spelling slips, since "Ana" against "Anna" is one character error rather than a whole wrong word.

Limits of WER: what the number hides

  • Every error counts the same: a misheard "a" and a misheard drug dosage each add one.
  • Punctuation and casing are usually stripped before scoring, though they affect readability and meaning.
  • It says nothing about timestamps, so subtitle timing needs separate checks.
  • It does not measure who said what; speaker diarization has its own metrics.
  • It depends on the reference, and a sloppy reference makes every system look worse.
  • An average hides variation, so one noisy file can be far worse than the overall figure suggests.

In translation workflows WER is only half the story, because a recognition error that changes meaning will carry into the translation while a harmless one won't. How accurate is AI video translation follows that chain from recognition to the final output.

When WER is the right tool

Despite those limits, WER is valuable whenever you need a repeatable number. It lets you compare two systems on the same files, check whether a change to your recording setup or audio processing helped, and set a quality bar for a project. It is cheap to compute once a reference exists, it is understood across the industry, and it tracks editing effort reasonably well: a transcript with a lower WER usually takes less time to correct.

How to measure WER on your own files

  1. Choose 5 to 10 minutes of audio that represents your real recordings, difficult parts included, rather than your cleanest clip.
  2. Transcribe it with the system you want to evaluate and download the plain transcript.
  3. Build the reference by correcting a copy of that transcript carefully against the audio, word by word. Starting from the machine output is faster than typing from scratch, but be strict, or the reference will inherit the machine's mistakes.
  4. Normalize both texts with the same rules: lowercase, punctuation removed, numbers written consistently.
  5. Compute WER with an established tool. The open-source Python library jiwer is a common choice, and NIST's sclite has long been a standard in research.
  6. Read the list of errors, not just the score, and note which ones matter for your use, such as names, numbers and terms.

Measuring a mydubly transcript

mydubly's transcripts come from Whisper large-v3-turbo by default, and you can download them as a timestamped transcript or as SRT and VTT subtitles in the spoken language. For WER, use plain text only: strip timestamps and cue numbers from an SRT file, or copy the transcript text, before normalizing. Subtitle line breaks stop mattering once the text is split into words.

Cost of a WER test

Suppose you evaluate three 8-minute samples from a lecture series. Transcribing them in mydubly costs 8 credits each, 24 credits in total, which is 2.4 cents and fits inside the 100 free credits new Google sign-ups receive. The larger cost is your own time spent building careful references.

Because the spoken language is detected automatically, glance at the first lines of each transcript to confirm they are in the expected language before scoring; a misdetected clip would make the number meaningless.

Next step: score a sample of your own

Pick one representative recording, transcribe it in video to text, and build a reference for a few minutes of it. The error list will tell you more than any published benchmark about what to proofread. Once you know your typical errors, how to proofread an AI transcript turns that into a review routine, and improving transcription accuracy covers fixes at the recording stage.

Frequently asked questions

What is a good word error rate?

It depends on the audio and the purpose. Clear single-speaker recordings transcribed by a modern model need far fewer corrections than noisy multi-speaker calls, so a figure that is excellent for one would be poor for the other. Decide how much editing you can accept, then measure your own files against that bar rather than a generic threshold.

Can word error rate be higher than 100 percent?

Yes. Insertions add to the error count, but the denominator is the number of reference words, so a transcript with many invented words can have more errors than there are reference words. This typically happens when a model hallucinates text through long stretches of silence or music.

Is WER the same as transcription accuracy?

Accuracy is often reported as 1 minus WER, so a WER of 0.08 becomes 92 percent accuracy. The conversion is convenient but can mislead, because WER can exceed 1 and the single figure hides which errors occurred. Report WER together with the normalization you used.

Why do two tools report different WER for the same transcript?

Usually because they normalize text differently, for example one removes punctuation and expands numbers while the other doesn't, or they treat filler words differently. Ties in the alignment can also split errors differently among substitutions, deletions and insertions, although the total stays the same.

Should I use WER or CER for Japanese?

CER. Japanese is written without spaces between words, so WER would depend on whichever word segmentation tool you happened to use. Character error rate applies the same formula to characters and avoids that problem; the same reasoning applies to Chinese.