Loud enough versus clear enough
Two different things get called "quiet". In the first, the whole recording is low in level, but the speech is clean: turn it up and it sounds fine. In the second, the speech is faint compared with everything else in the recording, the hiss of the recorder, the hum of the room, other people nearby. Turning that up makes everything louder together.
The difference is the gap between the speech and the noise floor, the constant background level you hear in pauses. Gain, normalization and the volume knob all raise speech and noise floor by the same amount, so the gap stays the same. A low-level recording with a large gap is easy to rescue. A recording where speech barely clears the noise floor is not, however loud you make it.
Speech recognizers generally care more about that gap than about absolute loudness. That is why a quiet but clean file often transcribes better than expected, and why a loud but distant one disappoints. The wider subject of noise and signal-to-noise ratio is covered in speech recognition and background noise; this article deals with level and distance.
Why boosting can't put back what was never recorded
The parts of speech that carry the most information for telling words apart are often the quietest. Consonants such as s, f, th, t and k are brief, high-frequency and much softer than vowels. When a speaker is far away or the recorder's gain was set too low, those consonants are the first to sink into the hiss. Raise the level afterward and you get louder vowels with the same smudged consonants, which is why "fifteen" and "fifty", or "can" and "can't", become guesses.
Distance does a second kind of damage. The farther a talker is from the microphone, the more of what reaches it is reflected sound from the walls rather than the direct voice. That reverberant blur is baked in, and level changes don't touch it. A phone on a conference table records the nearest person crisply and the far end of the table as a hollow, echoing voice, and no amount of gain gives the far end the clarity of the near end.
Very low recording levels in a digital file are less of a problem than people fear. A 16-bit or 24-bit recording has far more resolution than speech needs, so a well-made recording peaking a long way below full scale can be raised cleanly. The limit almost always comes from acoustic noise and distance, not from the file.
Whispered and murmured speech
Whispering removes the voicing, the buzz of the vocal folds that gives normal speech its pitch and much of its energy. What remains is breath shaped by the mouth, quiet and noise-like. People understand whispers through context and lip-reading; recognizers trained mostly on normal voiced speech have a much harder time, and whispered passages often come back with missing or wrong words.
Murmured speech, such as a speaker trailing off at the end of sentences or someone talking to themselves, sits between the two. Sentence endings that fade into the noise floor are a common source of truncated transcripts. If a project depends on quiet asides, plan to transcribe or check those passages by ear.
Long quiet stretches and invented text
A separate risk appears when a recording contains long stretches where nobody is speaking clearly, only room tone or distant mumbling. Whisper-style models are trained to produce text, and given near-silence they occasionally produce plausible phrases that nobody said, or repeat a previous line. Raising the level of the whole file makes faint background sound more prominent, which can encourage this. Whisper hallucinations explains the behavior and how to spot it.
Two practical consequences follow. Trim long silent sections at the start and end of a recording before uploading, and treat any transcript text that appears during a passage you hear as silent with suspicion.
A worked example: questions from the back of the room
A hypothetical 45-minute guest lecture was recorded on a phone at the lectern. The lecturer is clear, but students' questions from the back of the room are faint and echoey. The first transcript has the lecturer's answers and garbled or missing questions. Normalizing the whole file makes no difference to the questions. Applying gentle compression brings the questions closer to the lecturer's level, and that version recovers some of them, but not all. For the rest, the reviewer listens and writes a short summary such as [audience question about funding, partly inaudible]. The transcript costs 45 credits (4.5¢).
The fix for next time is cheap: ask the lecturer to repeat each question into their microphone before answering. That one habit is worth more than any processing of the faint originals.
Preparing a quiet recording before you upload
- Listen on headphones at a comfortable volume. If you can follow most words with effort, processing can help; if you can't, skip to the limits section below.
- Look at the waveform. If the loudest peaks are a long way below the top of the display throughout, the recording is low-level, and step 3 is safe.
- Normalize or apply gain so the loudest peaks sit just below full scale, around -1 dBFS. This makes review easier and does no harm.
- If one person is much quieter than another, apply gentle compression or use a loudness-leveling tool so the quiet passages come up relative to the loud ones. Go gradually and listen for pumping breaths and swelling background noise. Loudness normalization explained covers the difference between peak and loudness approaches.
- Try light noise reduction on a one-minute copy only. If the speech sounds watery or metallic afterward, back it off or skip it.
- A gentle low-cut filter around 80 to 100 Hz removes rumble that otherwise rises with the gain.
- Trim long silences at the start and end.
- Export losslessly, as WAV or FLAC, rather than re-encoding the boosted audio to a low-bitrate MP3.
- Transcribe a one-minute test of the hardest passage before running the whole file.
When re-recording is the only fix
- If the words aren't intelligible to a careful listener on headphones after level and compression, they won't be recovered by recognition either.
- If the important speaker was across the room from the only microphone, the reverberation is permanent.
- If whispered or murmured speech carries the content that matters, a transcript will need a human ear regardless.
- If the recording is replaceable, re-recording is faster than any rescue. Ask the speaker to repeat the key part, re-read a statement, or redo the interview with the microphone close.
- If it's irreplaceable, such as a recorded voicemail, a family recording or a one-off event, use the AI draft for the clear parts and transcribe the faint parts by hand, marking uncertain words rather than guessing. For a statement that matters legally, have a qualified professional transcriber produce the record.
Muffled audio is a related but different problem, where the speech is loud enough but dull because of fabric over the microphone or a covered phone; that is covered in muffled audio in recordings.
Quiet files in mydubly
mydubly transcribes the audio as it is in the file. It doesn't offer gain, compression or noise-reduction controls, so level work happens in an editor before upload. The browser decodes the audio to 16 kHz mono, splits it into chunks of about 30 seconds at the quietest moment near each boundary, and sends only compressed audio chunks over HTTPS for transcription with Whisper. Quiet passages are transcribed as well as their clarity allows, and faint ones may come back incomplete.
You receive a plain transcript, a timestamped transcript and SRT and VTT subtitles, and optionally a translated transcript. The timestamps are particularly useful for quiet recordings, because they let you jump straight to the passages that need listening. The audio to text tool accepts MP3, WAV, M4A, AAC, OGG and FLAC; for video, use video to text. If a dubbed version is the goal, any word lost in a faint passage is missing from the translated voice as well, so correct the transcript problems you can identify before deciding to dub.
Next step: test before you process the whole file
Pick the faintest minute you care about, prepare it with the steps above, and compare the transcript with one from the untouched original. If the prepared version is better, process the whole file the same way. For habits that prevent the problem next time, especially microphone distance, see how to improve transcription accuracy.
Frequently asked questions
Should I normalize audio before uploading it for transcription?
It is safe and often helpful for your own review, but on its own it rarely changes the transcript much, because it raises speech and background equally. What tends to help more is compression or loudness leveling when one speaker is much quieter than another. Normalize first so you can hear the file properly, then decide whether leveling is needed.
Is it better to record quietly or risk recording too loud?
Recording a little low is safer. A clean recording with peaks well below full scale can be raised later without harm, while a recording that hits full scale clips, and clipping distorts speech in a way that can't be fully undone. A common aim for speech is peaks somewhere around -12 to -6 dBFS, leaving headroom for laughter and emphasis; see clipped audio for what happens when levels run too hot.
Can an AI voice-enhancement tool rescue a faint recording?
Sometimes, partly. Speech enhancement tools can make faint speech easier for people to listen to, but they work by estimating what clean speech should sound like, and on very poor input they can smear or invent sounds. Compare a transcript of the enhanced version with one of the original on the same passage, and keep whichever reads more accurately against your own listening.
Why does my transcript stop partway through a quiet speaker's sentence?
Sentence endings are usually the quietest part of speech, because people drop their voice as they finish a thought. When the last words fall close to the noise floor, the recognizer may treat them as background and end the segment early. Gentle compression helps a little. Otherwise, review sentence endings in quiet passages and complete them by ear.
Does a higher bitrate help with quiet recordings?
Not much once the bitrate is reasonable for speech. Bitrate controls how faithfully the file stores what was recorded; it can't add clarity that the microphone didn't capture. Avoid very low bitrates when re-exporting a boosted file, since compression artifacts become more audible at higher levels. Audio bitrate for speech recognition covers sensible settings.