Quantization: turning a smooth wave into numbers
A digital recorder measures the incoming voltage thousands of times per second. How often it measures is the sample rate, covered in audio sample rate explained. How precisely it writes down each measurement is the bit depth.
Each sample has to be stored as one of a fixed set of values. A 16-bit sample can take 65,536 different values; a 24-bit sample can take about 16.8 million. The real voltage almost never lands exactly on one of those steps, so the converter rounds it to the nearest one. That rounding is called quantization, and the small difference between the true value and the stored value is quantization error.
When the signal is reasonably loud and complex, like speech or music, that error behaves like a very quiet, steady hiss spread across all frequencies. This hiss is the digital noise floor. Bit depth decides how far below full scale it sits.
How much dynamic range each bit adds
Every additional bit doubles the number of available steps, which halves the size of each step and lowers the rounding error by about 6 dB. That gives a simple rule of thumb: dynamic range in decibels is roughly six times the bit depth.
- 8-bit
- About 48 dB. Audibly hissy; found in old games and some telephony formats
- 16-bit
- About 96 dB. The CD standard and the most common delivery format
- 24-bit
- About 144 dB in theory. The usual recording and editing format
- 32-bit float
- A 24-bit-precision value with an exponent, giving range far beyond any real microphone or room
The theoretical figures for 24-bit are never reached in practice. Converters, preamps and microphones all generate their own electronic noise, and real recording chains land well short of 144 dB. Even a quiet room has air handling, traffic and the faint noise of the talker's own breathing. In almost every voice recording the acoustic and electronic noise is far louder than the quantization noise of a 24-bit file, and usually louder than that of a well-leveled 16-bit file too.
Bit depth has nothing to do with frequency response or how bright a recording sounds. A 16-bit file and a 24-bit file of the same speech, both at a good level, sound identical to listeners.
16-bit versus 24-bit: where the difference actually shows
The extra range of 24-bit matters most when levels are set low. Suppose you record with peaks around -30 dBFS to stay safe from a loud laugh. In a 16-bit file, those peaks now sit only about 66 dB above the quantization floor; when you raise the level later, the floor comes up with it. In a 24-bit file the same recording still has well over 100 dB of headroom below the peaks, so raising it adds no audible digital noise. The meter scale and what dBFS means are explained in what dBFS means.
That is why recorders default to 24-bit. It turns level-setting from a precise task into a forgiving one. You can leave generous headroom for surprises and fix the level afterwards.
Where 16-bit is enough:
- Final delivery of finished audio that has already been mixed and leveled.
- Speech recorded at a sensible level, peaking somewhere around -12 to -6 dBFS.
- Files headed for speech recognition, which cares about the voice, not the floor beneath it.
Where 24-bit earns its keep:
- Recording unpredictable sources, such as interviews, panels and live events.
- Projects that will go through heavy editing, compression or level changes.
- Archival masters of material you cannot record again.
32-bit float: headroom, not quality
Integer formats, 16-bit and 24-bit, have a hard ceiling. Any sample louder than full scale is cut off, which is digital clipping. A 32-bit floating-point file stores each sample as a number with a separate exponent, so values above full scale are kept intact. A file recorded too hot can be turned down afterwards and the peaks come back unclipped.
Some field recorders now capture 32-bit float straight from their converters, often using more than one converter per input to cover a very wide range. For run-and-gun interviews, that is a genuine safety feature. Two caveats apply. First, the analog input before the converter can still overload with an extremely loud source, so check the recorder's documentation. Second, float does nothing for a voice recorded at a good level; it sounds the same as 24-bit. How clipping damages speech, and what it does to a transcript, is covered in clipped audio and transcription.
Editing software works in 32-bit float internally for a similar reason: processing steps can push levels above full scale temporarily without damage, as long as the final output is brought back down before export.
Dither: why reducing bit depth needs a little noise
When a 24-bit or float project is exported to 16-bit, the extra bits have to go. Simply chopping them off, called truncation, makes the rounding error follow the signal. On loud material that hardly matters. On very quiet passages, such as the fade at the end of a podcast or a reverb tail, the error turns into a gritty, buzzing distortion that tracks the sound.
Dither fixes this by adding a tiny amount of random noise before reducing the bit depth. The noise breaks the link between the signal and the rounding error, so the error becomes a constant, smooth hiss instead of distortion. Quiet detail below the last bit remains audible inside that hiss rather than disappearing. Noise-shaped dither goes a step further, pushing the hiss toward frequencies where hearing is less sensitive.
Practical rules for dither:
- Apply it once, at the very end, when converting to the final integer bit depth.
- Do not dither when staying at 24-bit or float, or when exporting to a lossy codec such as AAC or MP3, which does not store a bit depth in the same way.
- Leave it on in your editor's export dialog if you are unsure; at 16-bit it is inaudible on normal speech levels.
- Avoid dithering repeatedly by bouncing a file to 16-bit, editing it, and bouncing again.
Bit depth and speech recognition
Speech recognition models listen for the patterns of the voice: formants, consonant bursts and the rhythm of syllables. All of that sits tens of decibels above the quantization floor of a 16-bit file. Once a recording is at a sensible level, extra bits add nothing a model can use.
Most recognition pipelines also convert the audio before the model ever sees it. They resample to a low rate such as 16 kHz, mix to mono and normalize the samples to a fixed numeric range. Whatever bit depth the source used, the model receives the same kind of input. What does change a transcript is the level relative to the background, the distance to the microphone and any clipping, which are the factors a recorder's settings should protect.
A journalist records a 40-minute interview on a recorder that writes three files at once: 16-bit, 24-bit and 32-bit float WAV, all mono at 48 kHz, peaking around -10 dBFS. The files take roughly 230 MB, 346 MB and 461 MB. Played back, they sound the same. Transcribed, all three give the same text, because the recognizer works on a 16 kHz mono version of each. The 32-bit file only proves its worth on the next interview, when a guest shouts and the peaks go over full scale without clipping.
Common bit depth mistakes and their limits
- Converting a 16-bit file to 24-bit to improve it. The new file is bigger, but the extra bits are empty; no detail is recovered.
- Recording extremely quietly at 16-bit because a meter looked scary. Leave sensible headroom instead, or switch to 24-bit.
- Treating 32-bit float as permission to ignore levels entirely. It protects against clipping in the file, not against a noisy room or a distant microphone.
- Exporting to 16-bit without dither from a project with long, quiet fades.
- Confusing bit depth with bitrate. Bitrate describes the size of a compressed stream; a lossy MP3 or AAC file does not carry a meaningful bit depth at all.
- Expecting more bits to remove noise. Bit depth only sets the digital floor; the room and the electronics usually set the real one.
Baseline recorder settings, including gain, limiters and low-cut filters, are covered in audio recording settings for speech.
How mydubly handles your file's bit depth
mydubly accepts WAV and FLAC files at common bit depths, along with MP3, M4A, AAC and OGG audio and MP4, MOV, WebM, MKV and M4V video, up to 2 hours per file. In the browser, ffmpeg.wasm decodes the audio and downmixes it to mono 16 kHz. It is then cut into chunks of about 30 seconds and compressed to Opus at 32 kb/s before being sent over HTTPS for recognition. Opus is a lossy codec, so the original bit depth does not survive that step, and it does not need to: a 24-bit or 32-bit float master and a well-leveled 16-bit copy reach the recognizer as essentially the same signal.
That means there is no benefit in exporting a higher bit depth for transcription, and no need to convert a 16-bit file before uploading. Keep the highest-quality master for your own editing. If you are uploading a WAV, the WAV to text page covers that format specifically.
Next step: check your recorder and export settings
Look at the settings on your recorder or phone app and note the bit depth, then check the default in your editor's export dialog and whether dither is on. Record at 24-bit or 32-bit float if available, export finished audio at 16-bit with dither or straight to your delivery codec, and send whichever file you have to mydubly's audio to text tool. A 40-minute interview costs 40 credits (4¢) to transcribe.
Frequently asked questions
Is 24-bit audio better than 16-bit?
For recording, 24-bit is more forgiving because it leaves far more room below the peaks, so you can record with generous headroom and raise the level later without adding digital noise. For listening to finished speech or music at a sensible level, most people cannot tell the two apart. 16-bit remains a perfectly good delivery format.
What is 32-bit float audio used for?
It stores samples as floating-point numbers, so values louder than full scale are preserved instead of being clipped. Recorders that capture 32-bit float let you turn down an over-hot recording afterwards and recover the peaks. Editing software also uses it internally so processing never clips between steps.
Do I need dither when exporting audio?
Use dither when reducing a 24-bit or float project to a 16-bit integer file, and apply it once at the final export. It is unnecessary when exporting to 24-bit, to float or to a lossy format such as AAC or MP3. On speech at normal levels the added hiss is inaudible.
Does higher bit depth improve speech recognition?
No, once the recording is at a sensible level. Recognition models listen for features of the voice that sit far above a 16-bit noise floor, and most pipelines convert audio to a standard internal format first. Level, background noise, microphone distance and clipping matter far more to a transcript.
Can I fix a quiet 16-bit recording?
Usually, yes. Raising the level brings up the noise floor along with the voice, but in most real recordings that floor is dominated by room and microphone noise rather than by quantization. If the speech is clearly audible when turned up, it will normally transcribe well.
What bit depth does an MP3 or AAC file have?
Strictly, none in the way a WAV file does. Lossy codecs store compressed frequency information rather than individual integer samples, and the decoder produces samples at whatever precision the player requests. The quality of a lossy file is described by its bitrate and encoder, not its bit depth.
Why does my WAV file say 32-bit when I recorded at 24-bit?
Many editors and converters export 32-bit float by default, or the file was processed in float and saved without conversion. The audio itself is not degraded. If another app refuses the file, export it again as 16-bit or 24-bit PCM, which is the most widely supported WAV variant.