Speech recognition & transcription

How background noise, echo and music affect speech recognition

Background noise hurts speech recognition by masking the parts of the sound spectrum that distinguish one speech sound from another, and reverberation and music do the same in subtler ways. The key quantity is the signal-to-noise ratio: how much louder the voice is than everything else. Modern models like Whisper tolerate moderate noise well, but heavy-handed noise reduction can strip out the very detail they rely on, so a better recording usually helps more than aggressive cleanup afterwards.

8 min read · Updated

How noise confuses a recognizer

A recognizer reads speech as a spectrogram, a map of energy across frequencies over time. Vowels are loud and sit in strong bands called formants; consonants such as "s", "f", "th", "p" and "t" are much quieter and spread into higher frequencies. Those quiet consonants carry a large share of what separates one word from another, and they are the first cues to disappear when noise fills the same parts of the map.

The result is a recognizable pattern of errors: substitutions between similar words, such as "fifty" and "sixty", dropped short words like "a", "of" and "to", and endings lost from plurals and past tenses. In a model like Whisper there is a further effect. Its decoder predicts the most plausible text given what it can hear, so when the audio is ambiguous it fills gaps with fluent guesses, which can be right, slightly wrong or invented.

Not all noise is equal. Steady noise such as a fan, air conditioning or road hum has a consistent spectrum that models handle relatively well. Irregular noise such as clattering dishes or keyboard clicks is harder. Babble, meaning other people talking, is hardest of all because it looks like speech, and the model may transcribe fragments of the wrong voice.

Signal-to-noise ratio, in practice

Signal-to-noise ratio (SNR) compares the power of the speech with the power of the noise and is expressed in decibels: SNR = 10 × log10(speech power ÷ noise power). Every 10 dB is a tenfold difference in power, so a 30 dB SNR means the voice carries a thousand times more power than the noise. The descriptions below are rough guides to how recordings sound, not accuracy figures.

30 dB or more
Noise is faint and barely noticeable, like a quiet room
Around 20 dB
Noise is clearly audible in pauses, but speech is easy to follow
Around 10 dB
Noise competes with the voice and a listener has to concentrate
0 dB
Noise is as loud as the speech, which is hard even for people

Recognition accuracy declines gradually as SNR falls rather than collapsing at a single threshold, and it declines faster for babble and music than for steady hum at the same level.

Measuring SNR in an editor

Select a pause with only background sound and read its average (RMS) level, say minus 52 dBFS. Then select a stretch of speech and read its level, say minus 22 dBFS. The difference, 30 dB, is your approximate SNR. Now suppose the microphone had been 25 cm from the speaker instead of 1 m: halving the distance twice raises the direct voice by roughly 6 dB each time, about 12 dB in total, while the room noise stays the same.

Reverberation: the noise the room adds

Echo works differently from background noise. In a room with hard surfaces, each word reaches the microphone directly and then again as many delayed reflections. On a spectrogram, the tail of each syllable smears over the start of the next, blurring exactly the transitions a recognizer uses to tell sounds apart.

Because reverberation is a copy of the speech itself, typical noise reduction can't separate it, and dereverberation tools tend to leave a hollow, processed sound. The factor you control is the ratio of direct sound to reflected sound, which depends mostly on how close the microphone is to the mouth. A ceiling microphone in a meeting room or a laptop across a classroom captures mostly reflections; a microphone near the speaker captures mostly voice.

Music under speech

Music occupies the same frequency range as the voice and changes constantly, so it masks speech more unpredictably than steady noise. Songs with vocals are the worst case: the model may transcribe the lyrics, or switch between the lyrics and the speaker. Passages of music with no speech are a classic trigger for invented text in Whisper-style models, such as stock phrases or repeated lines, which Whisper hallucinations examines in detail.

If you produced the content yourself, the cleanest fix is upstream: transcribe the dialogue track or voice stem rather than the final mix. For finished videos where only the mixed track exists, voice separation tools can help, but they introduce their own artifacts, so test their output before relying on it.

Risks of aggressive noise reduction

Noise reduction tools estimate the noise and subtract it, using spectral subtraction, noise gates or neural denoisers. Used gently, they make audio more pleasant to listen to. Used aggressively, they remove the quiet consonants and breaths along with the noise, chop the starts and ends off words when a gate closes too early, and leave warbling artifacts that engineers call "musical noise".

That can leave a recording that sounds cleaner to a person and transcribes worse. Whisper was trained on a vast amount of ordinary, imperfect audio, so it has learned to look past moderate natural noise; it has seen far less of the particular distortions a denoiser introduces. Heavy compression and loudness processing can cause a related problem by pulling the noise floor up during pauses. The only reliable way to know whether cleanup helps a given file is to transcribe a short excerpt both ways and compare.

When light cleanup does help

  • A narrow filter can remove steady electrical hum at 50 or 60 Hz without touching the voice.
  • A gentle high-pass filter, commonly set somewhere around 80 Hz, removes rumble from traffic, handling and wind below the speaking range.
  • One mild pass of noise reduction on steady noise such as a fan is usually safe.
  • If a recorder captured separate channels and one holds the speaker clearly, using that channel avoids mixing in the noisier one.
  • Cutting long music-only intros, outros and breaks removes the passages most likely to produce invented text.

Turning up the volume, by contrast, does nothing for recognition: the noise rises along with the voice, and the SNR stays the same.

What to fix before uploading

These checks apply to a file you already have. For advice on capturing better audio in the first place, see how to improve transcription accuracy.

  1. Listen with headphones to half a minute from the beginning, middle and end, and note what kind of noise you hear: steady, babble, music or echo.
  2. Estimate the SNR as in the example above, so you know whether the recording is merely imperfect or genuinely difficult.
  3. If the file is stereo and the speaker is mostly on one channel, export that channel as mono.
  4. Trim music-only sections and long silences at the start and end.
  5. If strong steady noise remains, apply one gentle pass of noise reduction and compare both versions on a short excerpt.
  6. After editing, export a high-quality copy, such as WAV or FLAC, rather than re-encoding repeatedly at a low bitrate.

Noise and the mydubly pipeline

When you open a file in mydubly, the browser decodes its audio to 16 kHz mono. All channels are combined in that step, which is why exporting the cleaner channel beforehand matters if your recorder captured a noisy second track. The audio is then cut into windows of about 30 seconds at the quietest 50 ms frame near each boundary. In a quiet recording that frame is a genuine pause; under constant loud noise the quietest moment is less distinct, so a cut is somewhat more likely to touch speech.

Each window is compressed to Opus at 32 kb/s, a setting suited to speech, and passed to Whisper large-v3-turbo. Compression is not a noise filter, so whatever noise is in the file travels with the voice, and any cleanup belongs before upload. If you go on to create a translated voice track, the dubbed MP4 replaces the original voice but keeps the background, so steady noise and music stay under the new voice, and recognition errors caused by the noise carry into the translation. Loud noise or music under speech also makes it harder to separate the original voice cleanly, which keeping background music when translating discusses.

Testing cleanup cheaply

Suppose you export a 3-minute excerpt twice, once untouched and once with gentle noise reduction. Each file is billed at the 5-credit minimum, so the comparison costs 10 credits, one cent, and tells you which version to use for the full recording.

Next step: test your noisy recording

Take the hardest few minutes of your recording, transcribe them in audio to text, and read the result while listening. If errors cluster in music or crosstalk, trim or rethink those sections; if they are spread evenly, the SNR is the problem and the next recording is where to fix it. Podcasters can find format-specific advice in podcast transcription.

Frequently asked questions

Does Whisper handle background noise well?

Better than many older systems, because it was trained on a large amount of real-world web audio that included noise. It still degrades as noise gets louder, struggles most with competing voices and music, and may invent text in passages without speech. Clean audio remains the most reliable route to an accurate transcript.

Should I run noise reduction before transcribing?

Only lightly, and test it. A gentle pass that removes steady hum can help, but strong noise reduction tends to strip quiet consonants and add artifacts that confuse recognition models. Transcribe a short excerpt both ways and keep whichever version produces fewer errors.

Is echo worse than background noise for transcription?

It can be, because reverberation is a smeared copy of the speech itself, so it overlaps every syllable and noise reduction can't separate it cleanly. A recording made far from the speaker in a hard-walled room often transcribes worse than a close recording with a steady fan in the background.

Why does my transcript contain words during music?

Whisper-style models generate text with a language model, and music-only passages give the decoder ambiguous input that it sometimes fills with plausible phrases or song lyrics. Trimming long music sections before uploading, and checking those passages in the transcript, removes most of these invented lines.

Does turning up the volume improve recognition?

Not by itself. Raising the volume makes the noise louder along with the speech, so the signal-to-noise ratio stays the same. What helps is a recording where the voice is louder relative to the noise, which comes from microphone placement and the environment rather than gain.