A baseline you can set once
Most recorders, recording apps and audio interfaces offer more options than speech needs. This baseline works for interviews, lectures, voice-overs and meetings, and you can leave it alone for years.
- Sample rate
- 48 kHz, or 44.1 kHz for audio-only projects
- Bit depth
- 24-bit, or 32-bit float if the recorder supports it
- Channels
- mono, one channel per microphone
- File type
- WAV (or BWF) while recording
- Normal speech level
- peaks around minus 12 dBFS
- Loudest moments
- below minus 3 dBFS
- Limiter
- on, as a safety net only
- Low-cut filter
- on, around 80 to 100 Hz
- Automatic gain
- off when someone can watch the meter
The rest of this article explains each line, so you know when to deviate. It deliberately skips microphone choice and room treatment; those matter more than any setting and are covered in choosing a microphone for speech and the transcription accuracy guide.
Sample rate: why 48 kHz is plenty
The sample rate sets the highest frequency a recording can contain: roughly half the sample rate. At 44.1 or 48 kHz that limit sits above the range of human hearing, which more than covers speech. 48 kHz is the convention for anything that will end up in video, so choosing it avoids a conversion later.
Speech recognizers usually work at a much lower rate internally. Whisper, for example, processes audio at 16 kHz, so a 96 kHz recording is downsampled before recognition and the extra information is discarded. Higher rates have uses in music production and sound design; for speech they only make files bigger. Recording directly at 16 kHz is also a poor idea, because it throws away detail you might want for editing and listening. The sample rate explainer goes into the theory.
Bit depth and headroom
Bit depth sets how much dynamic range the recording can hold, which in practice means how far below the maximum you can record before the noise floor becomes a problem.
At 16-bit, a well-set level sounds fine, but recording very quietly to stay safe costs you some quality when you raise it later. At 24-bit there is so much range that you can leave generous headroom, peaks around minus 12 dBFS, and still turn the level up afterwards without adding audible noise. That headroom is what protects you when an interviewee suddenly laughs.
32-bit float recorders store values above the digital maximum, so a file recorded too hot can usually be turned down afterwards without clipping, provided the recorder's input electronics weren't overloaded. They are a real safety benefit for unpredictable situations. They don't make a well-recorded voice sound any better.
Mono, stereo and multiple channels
One microphone produces one channel of information. Recording it as stereo just stores the same signal twice, or worse, puts it on one side only, which then plays in one ear. Choose mono for a single microphone.
Stereo makes sense in two situations: a true stereo microphone capturing a space, which speech rarely needs, or two separate microphones recorded to the left and right channels. The second is common for interviews and is really two mono recordings in one file. It lets you balance the speakers later. The trade-offs, including what happens when a stereo file is mixed down, are covered in mono vs stereo for speech.
File format on the recorder
Record uncompressed WAV whenever the device allows it. It costs storage, not quality, and it means every later edit or export starts from the original. Some recorders write BWF, a WAV variant with timestamps and metadata, which any audio editor opens.
The storage cost is predictable. A mono file at 48 kHz and 24-bit takes about 518 MB per hour; the same in stereo takes about 1 GB. A 96 kHz stereo recording uses around 2 GB per hour for no benefit to speech. WAV files also have a 4 GB size limit, which is why many recorders split long sessions into several consecutive files.
If a recorder or phone app only offers compressed formats, choose AAC or MP3 at a high setting, such as 192 kb/s or more for stereo. That is clean enough for transcription and editing, though not ideal for heavy processing.
Gain, limiter and low-cut filter
Gain is the setting that matters most, and the one most often wrong. Set it with the real speaker, at their real distance, talking as loudly as they will during the session, not with a quiet "testing, one, two". Aim for normal speech peaking around minus 12 dBFS and the loudest laugh or emphasis staying under minus 3. A recording that is too quiet buries words in hiss, and one that is too loud clips; clipped audio explains why clipping damages transcription permanently.
A limiter is a safety net that stops sudden peaks from reaching the maximum. Leave it on, but don't use it as a substitute for correct gain: a voice that constantly hits the limiter sounds squashed.
A low-cut (high-pass) filter removes rumble from traffic, air conditioning and handling noise, which lives below the voice's useful range. Around 80 to 100 Hz is a safe setting for speech.
Automatic gain control adjusts the level continuously. It rescues unattended recordings, but it also raises background noise in every pause and can make levels pump. Use manual gain when someone can watch the meter, and automatic gain only when nobody can.
Example: setting up a handheld recorder for interviews
A researcher records hour-long interviews on a handheld recorder with two lavalier microphones. The first session was recorded at 96 kHz stereo with the built-in microphones and automatic gain, and the interviewee's voice faded in and out with the room noise.
For the next session, she plugs the two lavaliers into the recorder's inputs, sets 48 kHz, 24-bit WAV, with the interviewer on the left channel and the interviewee on the right. She turns automatic gain off, switches on the limiter and an 80 Hz low-cut, and sets each input while the person speaks at full volume. The file is about 1 GB per hour instead of 2 GB, and both voices sit at a steady level.
A 60-minute session transcribed with a speech-to-text tool costs 60 credits (6¢). Because each person has their own channel, she can also rebalance a quiet answer in an editor before transcribing.
Configuring a recorder step by step
- Set sample rate to 48 kHz and bit depth to 24-bit or 32-bit float.
- Choose WAV and mono, or stereo only when two microphones go to separate channels.
- Turn off automatic gain and any built-in noise reduction or effects.
- Switch on the limiter and a low-cut filter around 80 to 100 Hz.
- Place the microphone, then set gain while the speaker talks at their loudest expected level.
- Record 30 seconds, listen back on headphones, and check the meter peaks and the noise in pauses.
- Confirm the date, time and file naming, and that the card has space for the whole session.
Settings that are overkill or backfire
- 96 or 192 kHz sample rates: larger files, no benefit for speech recognition or listening.
- Stereo for a single microphone: wasted space and a risk of one-sided playback.
- Heavy noise reduction or "voice enhancement" in the recorder: it can't be undone, and aggressive processing creates artifacts that confuse speech recognition. Clean up afterwards on a copy if needed.
- Recording very quietly "to be safe": at 16-bit especially, raising the level later raises the hiss with it.
- Low-bitrate compressed formats to save space: storage is cheaper than a re-interview.
- Trusting the default settings of a new device: check them, because some default to automatic gain or compressed formats.
What mydubly does with your recording
mydubly accepts WAV, FLAC, MP3, M4A, AAC and OGG audio, as well as common video formats, up to 2 hours per file. In the browser, the audio is decoded to 16 kHz mono and compressed to Opus at about 32 kb/s before only those chunks are sent over HTTPS for recognition with Whisper. A 96 kHz, 24-bit WAV and a 48 kHz one therefore look the same to the recognizer; what changes the transcript is the level, the noise and how close the microphone was. The upload size breakdown explains why 16 kHz mono is enough for speech.
Two consequences follow. A stereo interview recorded with one person per channel is mixed to mono, so both voices are transcribed together; mydubly doesn't label speakers, so separate channels don't produce speaker names. And if a recorder split a long session into several WAV files, each is a separate job, billed at 1 credit per minute with a 5-credit minimum per file, so joining very short files first can save a little.
No setting fixes the room itself; speech recognition and background noise covers that side of recording quality.
Next step
Set the baseline above on your recorder once, run a 30-second test, and keep the WAV as your master. When a session is done, upload the file to audio to text for a transcript and subtitles, or see the interview transcription use case for the rest of an interview workflow.
Frequently asked questions
Is 44.1 kHz or 48 kHz better for voice recording?
For speech, neither is audibly better. Both capture everything a voice produces and far more than speech recognition uses. Choose 48 kHz if the audio will be combined with video, because video projects use it by default, and 44.1 kHz if your whole workflow is audio-only and your editor is set to it. The important thing is to keep one rate through a project and avoid needless conversions.
Should I record speech in 16-bit or 24-bit?
Use 24-bit when the recorder offers it. It gives you room to leave plenty of headroom without the recording becoming noisy when you raise the level later. A 16-bit recording made at a good level is perfectly usable for transcription and listening, so there's no need to redo old recordings. The difference shows mainly when levels were set too low.
Does a higher bitrate improve transcription accuracy?
Only up to the point where compression stops audibly damaging the voice. Very low bitrates smear consonants and can increase errors, but a clean recording at a moderate compressed bitrate transcribes as well as uncompressed audio, because recognizers resample and compress internally anyway. Gain, noise and microphone distance have far more effect than bitrate once you're above that threshold.
Should I use the noise reduction setting on my recorder?
Usually not. Built-in noise reduction is applied permanently while you record, and if it is too aggressive the voice becomes watery or clipped at the start of words, which hurts both listening and speech recognition. Record clean audio with the microphone close, then apply gentle noise reduction afterwards on a copy if the background is distracting. That way you can always go back to the original.
What level should I record a voice-over at?
Aim for normal speech peaking around minus 12 dBFS with the loudest words below minus 3, the same as for interviews. Loudness targets for publishing are applied later, during mixing and loudness normalization, not while recording. Recording too hot to sound loud is a common mistake; you can always raise a clean recording, but you can't repair one that clipped.