Speech recognition & transcription

When Whisper Writes Words Nobody Said: Hallucinations Explained

A Whisper hallucination is text in a transcript that was never spoken: a stock phrase over silence, a sentence invented during music, or a line repeated over and over. It happens because Whisper's decoder is a language model trained on web transcripts, and when the audio gives it little to go on, it writes what usually comes next rather than nothing. Cutting audio at pauses, removing long silences and checking a few predictable spots in the transcript prevent or catch most cases.

6 min read · Updated

What Whisper hallucinations look like

Hallucinations fall into a few recognizable patterns. Knowing them makes them easy to spot.

  • Stock phrases over silence or music, such as "Thank you for watching", "Please subscribe" or a subtitle credit line, often at the end of a video or during an intro.
  • Invented sentences that sound plausible for the topic but do not match the audio, typically during long pauses, applause or background noise.
  • Repetition loops, where one phrase or sentence repeats many times, sometimes for the rest of a 30-second window.
  • Language drift, where the output switches into another language or into an English translation for a stretch of audio.

Skipping is the mirror-image failure: the model sometimes jumps over a sentence of real speech. It is not strictly a hallucination, but the causes and checks overlap.

Why a speech model invents words

Whisper is an encoder-decoder transformer. The encoder turns 30 seconds of audio into a representation; the decoder then writes text token by token, each time predicting the most likely next token given the audio and the text so far. That decoder is effectively a language model. When the audio is clear, the acoustic evidence dominates. When the audio is silent, musical or noisy, the language-model side fills the gap with whatever text is statistically likely.

The training data explains the specific phrases. Whisper was trained on a very large amount of audio paired with transcripts and subtitles found on the web. Many of those subtitle files contain lines that were never spoken, such as subscribe reminders, translator credits and sign-offs, and those lines often sit over the closing seconds of a video. The model learned that silence or music near the end of a clip is frequently accompanied by exactly that kind of text.

Whisper does have a special no-speech prediction for windows without speech, but it is a learned probability, not a guarantee. Background music, breathing or room noise can keep it from triggering.

Repetition loops and conditioning on previous text

Long-form transcription with Whisper usually feeds the text from the previous window to the decoder as a prompt, so names and style stay consistent across windows. The open-source reference implementation does this by default. The side effect is a feedback loop: if one window ends in a repeated or invented phrase, the next window is primed with it and may continue the pattern.

Within a window, a decoder that falls into a high-probability phrase can keep choosing it, because each repetition makes the next one more likely. Greedy decoding, where the model always picks the single most likely token, is especially prone to this.

Safeguards built into the reference implementation

The reference implementation includes heuristics that catch many failures. Exact option names and defaults can change between versions, so check the official repository before relying on them.

No-speech probability
If the model thinks a window contains no speech and its confidence in the text is low, the window's text is dropped.
Average log probability
Low average confidence across the output marks the decode as suspect.
Compression ratio
Highly repetitive text compresses unusually well, so a high gzip compression ratio flags a likely loop.
Temperature fallback
When a decode fails these checks, it is retried with more randomness, which often breaks a loop.
Previous-text conditioning
Can be switched off, trading some cross-window consistency for a lower risk of loops spreading.

These checks reduce the problem without eliminating it. A fluent invented sentence over applause can pass every one of them.

Mitigations that work before the model sees the audio

The most effective defenses keep non-speech audio away from the decoder in the first place. If you build your own Whisper pipeline, these are the usual steps.

  1. Run voice activity detection and send only speech regions to the model, so long silences and music beds are never transcribed.
  2. Cut long audio into windows at natural pauses rather than at fixed 30-second marks, so windows do not start or end mid-word.
  3. Avoid tiny final windows containing a second or two of trailing silence or room tone.
  4. Trim long musical intros and outros before transcription when you do not need them transcribed.
  5. Consider turning off previous-text conditioning for recordings with many silent gaps.

Noise reduction is a double-edged tool here: removing steady hum can help, but aggressive processing that leaves warbling artifacts can make things worse. The article on speech recognition and background noise covers that balance.

How mydubly's audio chunking reduces the risk

mydubly's transcription runs on Whisper, with large-v3-turbo as the default model, so the same failure modes apply. Its audio preparation is designed to give the model clean windows. In the browser, the audio is decoded to 16 kHz mono and split into windows of about 30 seconds, and each cut is placed at the quietest 50-millisecond frame within the last 6 seconds before the 30-second mark. Words and sentences are therefore not cut in half at a window edge, which avoids a common trigger for invented or repeated fragments. A short tail at the end of the file is folded into the last window rather than sent as a separate, nearly empty chunk. The details are in audio chunking for speech recognition.

That lowers the risk; it does not remove it. A window that is mostly music, applause or silence can still produce invented text, and because a full mydubly job translates and voices the transcript, an invented "thanks for watching" over the end credits would be translated and spoken in the dub. That makes a quick hallucination check worthwhile before you produce a voice track.

Example: a tutorial with an outro

Suppose a 12-minute tutorial ends with 40 seconds of music and a logo. The transcript's final segment reads "Thanks for watching, see you in the next video!" although the presenter's last words were "That's it for the export settings." Searching the transcript for "watching" finds it in seconds; deleting the line, or trimming the outro before processing, keeps it out of the subtitles and the dub.

Where hallucination risk is highest, and where it is low

Risk is low for continuous speech: lectures, narrated tutorials, interviews and podcasts with little dead air. The decoder has strong acoustic evidence throughout, and hallucinations there are rare.

Risk rises with:

  • Long silences, such as pauses in a slow interview or an isolated microphone track where the other person is talking.
  • Music-only passages, intros, outros and transitions.
  • Applause, laughter and crowd noise in talks and events.
  • Speech with long hesitations, which leaves silent gaps mid-sentence.
  • Very quiet speech that the model can barely hear.

What to check in a transcript

  1. Read the first and last few segments; intros and outros are the most common place for stock phrases.
  2. Search for "watching", "subscribe", "subtitles" and "thank you", and confirm each hit against the audio.
  3. Scan for any sentence repeated two or more times in a row.
  4. Look for segments whose timestamps cover a music bed, a pause or applause, and check whether that text was spoken.
  5. Check for passages in the wrong language, which can indicate drift or a language identification problem.

The broader review workflow is in how to proofread an AI transcript.

Next step

Transcribe a file with video to text and run the five checks above on the output; on most recordings they take a few minutes. If your videos have long musical intros or outros, trim them before uploading so neither the transcript nor a dub picks up invented lines.

Frequently asked questions

Why does Whisper write 'Thank you for watching' when nobody said it?

Whisper learned from web subtitles, many of which include sign-offs over the closing seconds of a video. When it hears silence or music, especially near the end of a recording, its decoder can produce that familiar phrase instead of nothing.

Does a larger Whisper model hallucinate less?

Not reliably. Larger models are generally more accurate on real speech, but hallucinations come mainly from non-speech audio reaching the decoder, so audio preparation matters more than model size.

How do I stop Whisper from repeating the same line?

Make sure windows are cut at pauses and contain speech, use the compression-ratio check and temperature fallback the reference implementation provides, and consider disabling conditioning on previous text so a loop in one window does not prime the next.

Are hallucinations dangerous or just annoying?

Mostly annoying, but they can matter. An invented sentence in a quoted interview, a medical note or a legal record can misrepresent what someone said, so any transcript used as evidence should be checked against the audio.

Can a hallucinated line end up in a translated video?

Yes. In a pipeline that translates and voices the transcript, any invented text is translated and spoken like the rest. Checking the transcript, or trimming music-only intros and outros before processing, prevents that.