What the five minutes are for
A transcript, a translation or a dubbed voice inherits every problem in the recording. Most of those problems are visible in minutes if you know where to look, and some, like a channel wired backwards, make a file transcribe as near-silence. The point of a pre-flight check is not to fix anything. It is to decide one of three things: the file is fine as it is, it needs a specific fix first, or the speech isn't recoverable and you need a different source.
The checks below assume an audio editor such as Audacity (free on Windows, macOS and Linux), headphones, and optionally ffmpeg. Each step links to a dedicated article for the fix, rather than repeating it here.
Minute one: read the file's vital signs
Open the file in MediaInfo, or run ffprobe -hide_banner recording.wav, and note:
- Duration. Does it match what you expect? A 90-minute interview that shows 41 minutes is incomplete.
- Format and codec. A low-bitrate compressed file, such as speech at well under 64 kb/s, can be smeared before you even listen.
- Sample rate. 44.1 or 48 kHz is normal; 8 kHz means telephone-band audio, which transcribes worse and can't be improved by resampling.
- Channels. Mono or stereo, and whether a "stereo" file might be two different microphones.
- Number of audio tracks in a video file. If there are several, make sure you know which one has the speech; the multiple audio tracks article explains how tracks are chosen.
Minute two: listen on headphones at chosen spots
Laptop speakers and earbuds in a noisy office hide exactly the problems speech recognition stumbles on. Use closed headphones, at a comfortable volume, and listen to four short spots rather than the whole file:
- The first 30 seconds, where level and setup problems show up.
- A stretch in the middle, to catch changes such as a microphone that slipped or a door that opened.
- The quietest speaker, wherever they talk.
- The last minute, to confirm the recording didn't cut off early or degrade as a battery ran down.
Listen for distance and echo (does the voice sound like it's in a hallway?), steady noise (fans, hum, hiss), bursts (wind, handling, clicks), distortion on loud words, and dropouts. Write each problem down with a timestamp.
Minute three: read the waveform and the meter
In Audacity, the waveform tells you a lot at a glance. Speech looks like irregular blocks separated by thinner gaps. Learn to read four patterns:
- Blocks that fill only a sliver of the track height mean the recording is very quiet. That can be fine if the noise is low, and it is covered in transcribing quiet audio.
- Blocks with flat, squared-off tops touching the edges mean clipping. Audacity can mark clipped samples in red when you turn on Show Clipping in the View menu.
- Gaps between phrases that are almost as thick as the speech mean a high noise floor.
- One channel showing speech and the other nearly flat means a one-sided recording.
If you prefer numbers, ffmpeg -i recording.wav -af volumedetect -f null - prints the mean and maximum volume. A maximum of 0.0 dB on a file that wasn't deliberately normalized is a strong hint of clipping, and a mean far below the rest of your recordings means it's quiet.
Minute four: noise floor, clipping and channels
Noise floor: select two or three seconds of a pause, where nobody talks, and play it with the meter visible. Then compare it with the level of normal speech. There's no universal threshold, but as a rule of thumb, if the pause sits only a little below the speech, recognition will struggle and words at the ends of sentences will go missing. Note what kind of noise it is, because steady hum, broadband hiss and chatter need very different fixes.
Clipping: zoom in on the loudest moments, such as laughter or emphasis. A few clipped peaks are harmless; clipped syllables throughout are not, and clipping can't be fully undone. The clipped audio article explains how far repair can go.
Channels: create a mono mix and listen to it, for example with ffmpeg -i recording.wav -ac 1 mono-test.wav, or with Audacity's Mix Stereo Down to Mono on a copy. If the voice gets thin, hollow or nearly disappears in the mono version, one channel is polarity-inverted relative to the other, and the two cancel when combined. This matters because many transcription tools mix stereo to mono before recognition.
Minute five: run a short test file
The final check is the one that answers the real question. Cut one or two minutes from the hardest part of the recording, such as the quiet speaker or the noisiest stretch, and export it as its own file. Transcribe it and read the result against the audio.
Count the problems that matter for your use: missed or wrong names, whole phrases that vanished, invented text in pauses, the wrong language. If the sample is clean, the full file will be too. If it's poor, you've spent a minute instead of an hour finding out, and you know which fix to try first.
Example: a venue recording that would have transcribed as silence
A nonprofit receives a 75-minute stereo WAV of a keynote from the venue's sound desk and wants a transcript and Spanish subtitles. On headphones, the keynote sounds normal, if a little wide.
The waveform shows healthy speech on both channels and no clipping. The noise floor in pauses is low. But the mono test mix makes the speaker almost vanish, leaving a faint, phasey whisper. One cable or channel at the desk had been wired with reversed polarity, so the two channels are near mirror images and cancel out when combined.
The fix takes a minute: in Audacity, split the stereo track into two mono tracks, delete one, and export the other as a mono WAV. A two-minute test clip then transcribes cleanly for 5 credits (0.5¢), and the full keynote follows for 75 credits (7.5¢). Without the check, the first sign of trouble would have been an almost empty transcript.
What a quick check can't tell you
A five-minute check samples the file; it doesn't audit it. A dropout at minute 63 or a guest who switched to another language halfway through can slip past four listening spots. For long or high-stakes recordings, skim the whole waveform for sudden changes and spot-check more places.
The check also can't judge content: whether a speaker's technical terms will be recognized, whether accents or fast speech will cause errors, or whether a translation will be accurate. Those depend on the speech itself, and the only real test is reading a sample transcript. And a clean-sounding file isn't proof of good recognition; heavily processed audio can sound pleasant to people and still confuse a recognizer.
When the check finds a problem
- Muffled, dull voice
- see muffled audio in recordings
- Steady hum or buzz
- see removing hum from audio
- Traffic, chatter or room noise
- see speech recognition and background noise
- Gusts and rumble outdoors
- see wind noise in video
- Flat-topped, distorted peaks
- see clipped audio
- Voice disappears in mono or plays on one side
- keep the good channel, or see mono vs stereo for speech
Fix problems on a copy, never on the original, and re-run the test clip afterwards. Heavy cleanup can make speech harder to recognize, so check that the fix actually improved the transcript, not just the sound.
Checking audio before a mydubly job
mydubly decodes audio in the browser to 16 kHz mono before sending compressed chunks for recognition with Whisper, so the mono test in minute four is a direct preview of what the recognizer will receive. If a video file has several audio tracks, only one is used, which is why the vital-signs check includes counting tracks.
The test-file step is cheap: transcripts cost 1 credit per minute with a 5-credit minimum per file, so a two-minute sample costs 5 credits (0.5¢). If you plan to dub the file, test with a transcript first, because the voice is generated from the transcript and any recognition errors carry into the translation and the new voice. The transcription accuracy guide has advice for the next recording, when the current one can't be saved.
If a file passes the check but the transcript still contains odd phrases over silence or music, Whisper hallucinations explains the likely cause.
Make the check a habit
Run the five-minute check whenever a recording comes from someone else, a new device or a new location. Once a file passes, upload it to audio to text, or to video to text if you're working from the original video.
Frequently asked questions
What level should speech peak at in a finished recording?
For a raw recording, normal speech peaking somewhere around minus 12 to minus 6 dBFS, with nothing touching 0 dBFS, is a healthy range. A file that peaks lower can still transcribe well if the background is quiet. A published podcast or video is usually louder because it has been mixed and normalized, so don't compare a raw interview with a finished episode.
Can I check audio quality without installing anything?
Partly. Careful headphone listening at a few points catches most problems, and many media players show duration, format and channels in an info panel. A waveform view and a mono mix test need an editor such as Audacity, which is free. The cheapest objective test of all is transcribing a short sample, which needs nothing beyond a browser.
How do I know if audio is good enough for dubbing as well as transcription?
For dubbing, what matters is the transcript, because the new voice is generated from it, not from the original sound. If a test transcript is accurate, the recording is good enough to dub. The original speech is removed and the music and background sounds are kept under the new voice, so loud music or noise can leave faint traces of the original voice in the mix, and they can cause recognition errors that carry into the dub.
Should I check every file in a large batch?
For a batch from one source and one setup, check the first file thoroughly, then do a quick vital-signs read and a single listening spot on the rest. Look out for outliers: a different duration, sample rate or channel count often signals a different device or a setup change. Any file from a new device, location or person deserves the full five minutes.
Does normalizing a file before transcription help?
Normalizing raises or lowers the whole file to a target level. It can make a very quiet file easier for you to review, but it doesn't change the ratio between speech and noise, so it rarely changes recognition much. It never fixes clipping. If you do normalize, use a copy and keep peaks below 0 dBFS.