What a decibel actually measures
A decibel is not a fixed amount of sound. It is a way of expressing a ratio between two values on a logarithmic scale, which suits audio because hearing itself works roughly logarithmically. Every decibel scale has a reference point, and the letters after "dB" tell you what that reference is.
A few relationships are worth remembering. A change of about 6 dB doubles or halves the amplitude of a signal, and a change of 20 dB multiplies or divides it by ten. A change of 1 dB is around the smallest difference many people notice in a direct comparison, and a 10 dB increase is commonly described as sounding roughly twice as loud. These are rules of thumb, but they make meters much easier to read.
dBFS, dBu, dBV and dB SPL compared
- dBFS
- Decibels relative to digital full scale. 0 dBFS is the maximum sample value; levels below it are negative. Used on every digital meter
- dBu
- Decibels relative to 0.775 volts. Used for analog equipment; professional line level is commonly quoted as +4 dBu
- dBV
- Decibels relative to 1 volt. Consumer line level is commonly quoted as -10 dBV
- dB SPL
- Sound pressure level in the air, relative to roughly the quietest sound a young listener can hear. Used for microphones, rooms and hearing safety
- dBTP
- True peak in dBFS, estimating peaks that occur between samples when audio is converted back to analog
The important point is that these scales do not convert neatly into one another without knowing the equipment. How many dBFS a given dBu level produces depends on how a particular interface or recorder is calibrated. Broadcast and studio alignment conventions often place a reference tone somewhere around -18 or -20 dBFS, but consumer gear varies. Similarly, how loud a voice in dB SPL appears on your meter depends on microphone sensitivity, distance and preamp gain.
Why digital levels are negative
In a fixed-point digital recording, each sample is a number with a maximum possible value. Full scale is that maximum, so it is defined as 0 dBFS, and every quieter level is measured downward from it: -6 dBFS is half the full-scale amplitude, -12 dBFS a quarter, and so on.
Nothing can exceed 0 dBFS in a fixed-point file. If the input is too strong, the converter writes the maximum value repeatedly and the waveform's tops are flattened. That is clipping, and it sounds like crackle or harsh distortion. What can and cannot be recovered from a clipped recording is covered in clipped audio and transcription.
32-bit floating-point audio, used inside most editors and by some recorders, can represent values above 0 dBFS without clipping internally. That gives a safety net during editing, but the analog-to-digital converter in front of it still has a maximum, and the final export to a fixed-point format will clip anything left above 0. The relationship between bit depth, noise floor and dynamic range is explained in audio bit depth.
Peak, RMS and loudness meters
Different meters answer different questions, and many apps show more than one at once.
- Sample peak meters show the highest sample value in each moment. They react instantly and tell you how close you are to clipping. Most recording apps default to these.
- True peak meters estimate the peaks between samples that appear when audio is converted or encoded. They can read slightly higher than sample peaks.
- RMS meters show an average level over a short window, closer to how loud a signal feels. Speech typically reads well below its peaks on an RMS meter, because speech is spiky.
- VU-style meters are a traditional averaging meter with slow, smoothed movement. Their scale is referenced to an alignment level, not to 0 dBFS.
- Loudness meters measure in LUFS or LKFS using a frequency weighting that approximates hearing, with momentary, short-term and integrated readings. They are used to match loudness between programs; targets for that are covered in loudness normalization explained.
The gap between peak and average level is called the crest factor. Raw speech has a large crest factor, which is why a voice can look quiet on an RMS or loudness meter while its peaks are already close to 0 dBFS. Compression and limiting reduce that gap.
Headroom and the noise floor
Headroom is the space between your typical peaks and 0 dBFS. It exists because speech is unpredictable: a laugh, a cough or an emphatic word can be several decibels louder than the rest. Without headroom, those moments clip.
At the other end is the noise floor: the hiss of the preamp, the hum of the room, the converter's own noise. Recording too quietly brings speech closer to that floor, and raising it later raises the noise by the same amount. Good gain staging means setting the input so the voice sits comfortably above the noise floor while leaving enough headroom for surprises. With 24-bit recording, the digital noise floor is so low that the analog noise of the microphone and preamp is almost always the limiting factor, which is why modest levels are safe.
Recording level targets for speech
There is no single correct number, but these ranges are commonly suggested for spoken word on a sample peak meter:
- Normal speech peaks
- Roughly -12 to -6 dBFS
- Loudest moments, such as laughs
- Staying below about -3 dBFS
- Average or RMS level
- Often somewhere around -24 to -18 dBFS, depending on the speaker and meter
- Room tone with nobody speaking
- As low as the room and equipment allow; well below the speech level
To set levels in practice:
- Ask the speaker to talk as they will during the recording, including a laugh or an emphatic sentence, not a quiet "testing, testing".
- Adjust the input gain on the interface, recorder or microphone, not the playback volume, until normal speech peaks land in the target range.
- Check the loudest moment stays clear of the top of the meter and the clip indicator stays off.
- For a 24-bit recorder, if in doubt, err slightly lower rather than higher; you can add gain later without harm.
- Watch the meters during the session. Speakers get louder as they relax and quieter as they tire.
Reading meters in common apps
Most recording and editing software, from free editors such as Audacity to video editors and streaming apps, shows levels in dBFS with 0 at the top. Some details worth knowing:
- Colors usually move from green to yellow to red as levels approach 0 dBFS. Red does not always mean clipping; it often means "close".
- A peak-hold line or number shows the highest recent peak, which is useful because peaks pass too quickly to read.
- A clip indicator stays lit after an overload until you reset it. If it lights during a take, check the waveform at that moment.
- Some apps draw an average level as a lighter bar inside the peak bar. Check the app's documentation if you are unsure which you are looking at.
- Meter scales are not always linear. A meter that spends a lot of its height on the top 20 dB can make quiet recordings look emptier than they are.
- Phone voice-memo apps often show only a simple waveform or no meter at all; keep the phone at a steady distance and do a short test recording.
Two people record a 45-minute interview on a USB interface. During the sound check, the host's normal speech peaks at about -9 dBFS and the guest's at about -20 dBFS. The producer raises the guest's input gain until her peaks also sit around -9 dBFS, then asks both to laugh; the loudest laugh peaks at -4 dBFS, so the levels stay. Midway through, the guest leans in and starts peaking near -2 dBFS, so the producer turns her gain down a few decibels at a natural pause. Nothing clips, and both voices need little level adjustment in the edit.
Common level mistakes
- Setting gain from a quiet test phrase, then clipping when the real conversation starts.
- Recording very low to be safe on 16-bit or lossy recorders, where the noise floor is higher.
- Turning up the headphone or speaker volume and thinking the recording is now louder.
- Using digital gain or normalization afterwards to rescue a recording that was far too quiet, which raises the noise along with the voice. If that is your situation, transcribing quiet audio explains what still works.
- Treating dBFS numbers from different devices as equal in loudness. The same reading can come from very different acoustic levels depending on microphone and gain.
Levels and mydubly
For transcription, a recording does not need to be loud; it needs speech clearly above the noise and no clipping. mydubly decodes your file in the browser, mixes it to mono at 16 kHz and sends compressed chunks for recognition, with a voice activity filter that skips silent stretches while keeping timestamps aligned to the original audio. mydubly has no level controls, so a recording with healthy levels in your editor is the right starting point.
For dubbing, the synthesized voice clips mydubly generates are loudness-normalized before the voice track is assembled, so you do not need to match your original recording's levels for the dubbed version. The original speech is removed from the dubbed track, while any music and effects are kept under the new voice and lowered automatically while it speaks. More on that output is on the AI dubbing page.
Next step
Open a recent recording in your editor and note the peak and average readings during normal speech. If peaks sit roughly between -12 and -6 dBFS with nothing clipped, your levels are in good shape. Transcribe it with mydubly's audio to text tool, or, for a phone recording, see the voice memo to text page.
Frequently asked questions
Is 0 dBFS too loud?
0 dBFS is the absolute ceiling of a digital recording, so reaching it during recording usually means clipping. Most people record speech with peaks well below it, commonly around -12 to -6 dBFS, and use a limiter later if a finished file needs to be louder.
What is the difference between dBFS and LUFS?
dBFS measures a sample's level relative to digital full scale, typically shown on a peak meter. LUFS measures perceived loudness over time with a frequency weighting that approximates hearing. A recording can peak near 0 dBFS yet have a low loudness in LUFS if most of it is quiet.
Why does my meter show negative numbers?
Because digital levels are measured downward from full scale. 0 dBFS is the maximum, so everything quieter is negative: -6 dBFS is about half the maximum amplitude, -12 dBFS about a quarter. Negative numbers are normal and expected.
How much headroom should I leave when recording voice?
A common approach is to keep normal speech peaks around 6 to 12 dB below full scale, so sudden laughs or emphasis still fit. With 24-bit or 32-bit float recording you can afford a little more headroom, since the extra noise from recording slightly lower is negligible.
Can I convert dBFS to dB SPL?
Not without knowing the microphone's sensitivity, the preamp gain and the distance to the source. The same dBFS reading can come from a quiet voice with high gain or a loud voice with low gain. Calibrated measurement setups make the conversion, but everyday recording setups do not.
Does recording level affect transcription accuracy?
Extreme levels do. Clipped audio distorts words, and very quiet recordings bring speech close to the noise floor, which can cause missed words. Within a sensible range, the exact level matters much less than background noise, echo and microphone distance.