Audio engineering for speech

How to EQ a Voice Recording: Frequencies, Cuts and Sensible Boosts

Good EQ for voice usually starts with a high-pass filter somewhere around 70–120 Hz, a small cut in the muddy low mids if the room or microphone added boom, and only then a gentle lift in the presence region around 2–5 kHz if the voice still sounds distant. Cut narrow problems before you boost broad regions, make every move while listening at a normal volume, and compare with the EQ bypassed. For transcription, a light touch is enough; heavy EQ is for listeners, not for speech recognition.

9 min read · Updated

Where a voice sits in the frequency range

Equalization changes the level of particular frequency bands. To use it well on speech, you need a rough map of what each region of the spectrum contributes. The ranges below are common rules of thumb rather than precise boundaries; every voice, microphone and room shifts them.

Below about 80 Hz
Rumble: traffic, air conditioning, desk bumps, handling noise. Very little useful speech energy
About 80–250 Hz
The fundamental pitch of most adult voices and the sense of body or warmth. Too much sounds boomy
About 200–500 Hz
The muddy zone. Room build-up and proximity effect pile up here and blur words together
About 500 Hz–1 kHz
Boxiness and nasal or honky tones, especially in small untreated rooms
About 2–5 kHz
Presence. Consonants that carry intelligibility live here; the ear is also very sensitive to harshness in this band
About 5–10 kHz
Sibilance and crispness: s, sh, t and f sounds
Above about 10 kHz
Air and breath. Pleasant in small amounts, mostly irrelevant to understanding words

Two practical points follow from this map. First, the parts of speech that make words understandable are mostly consonant energy in the low kilohertz range, not the low-end warmth that makes a voice sound rich. Second, problems often show up far from where you hear them: a voice that sounds dull may be suffering from too much low-mid energy rather than too little treble.

Start with a high-pass filter

A high-pass filter removes everything below a chosen frequency and leaves the rest alone. It is the single most useful EQ move for spoken word because almost every recording carries low-frequency energy that adds nothing: building rumble, footsteps, a laptop fan, the thump of a hand on the table.

Where to set it depends on the voice. A common starting point is around 80 Hz for deeper voices and somewhere around 100–120 Hz for higher voices, then adjusting by ear. Raise the cutoff until the voice starts to sound thin, then back it off a little. A gentle slope such as 12 dB per octave sounds more natural than a steep one, although a steeper slope can help when rumble is severe.

Many microphones and recorders have a high-pass switch, sometimes labeled with a bent line symbol. Using it at the source keeps rumble out of the recording in the first place, which is better than removing it later. Electrical hum at 50 or 60 Hz and its harmonics is a different problem that a high-pass filter only partly addresses; narrow notch filters for hum are covered in removing hum and buzz from audio.

Mud and boxiness in the low mids

Once rumble is gone, the next common problem is too much energy between roughly 200 and 500 Hz. Two things usually cause it. Directional microphones used close to the mouth boost low frequencies, an effect called proximity effect. Small rooms with hard walls also reinforce certain low and low-mid frequencies, making the voice sound as if it was recorded in a cupboard.

A modest cut of 2 to 4 dB with a medium-width bell filter in this region often makes a voice clearer without changing its character. Boxiness a little higher up, around 500 Hz to 1 kHz, responds to the same treatment. Keep these cuts small. Taking out too much low-mid energy leaves a voice thin and brittle, and it does not remove the room itself; it only changes its tone. Treating the room is the real fix, and EQ is a cosmetic one.

If the whole recording sounds dull and blanketed rather than boomy, the cause may be a covered microphone, a distant placement or heavy noise reduction, which is a separate diagnosis covered in fixing muffled audio.

Presence, harshness and air

The presence region, roughly 2 to 5 kHz, is where EQ can make a voice sound closer and easier to follow. A broad boost of 1 to 3 dB is usually plenty. Larger boosts quickly become tiring, because the ear is especially sensitive in this range and harsh tones there make listeners turn the volume down.

Harshness is the opposite problem: a piercing, edgy quality, often from a bright condenser microphone, a voice that is naturally sharp, or an aggressive presence boost added earlier. A narrow cut somewhere between about 2.5 and 4 kHz often softens it. Find the exact spot by ear, as described in the next section.

Above that, sibilance around 5 to 10 kHz is better handled with a de-esser than with static EQ, because s sounds are short and a static cut dulls every other sound too. How de-essers work and how to set them is covered in our guide to taming sibilance. A small high-shelf boost above about 10 kHz adds air to a clean recording, but it also lifts hiss, so skip it if the noise floor is audible.

Cut before you boost: finding a problem frequency

Cutting removes something that is already wrong; boosting adds energy and also raises any noise in that band. Starting with cuts keeps the result natural and avoids pushing levels toward clipping. A reliable way to find a problem frequency:

  1. Pick a bell filter, set it fairly narrow and boost it by 8 to 10 dB. This is temporary.
  2. Play a section where the problem is obvious and slowly sweep the frequency up and down.
  3. Stop where the problem jumps out most clearly: the boom, the honk or the harsh edge.
  4. Turn the boost into a cut, typically 2 to 6 dB, and widen the filter until it sounds natural.
  5. Bypass the EQ and compare at the same perceived loudness. If the change is not clearly better, reduce it.

Level matching matters in that last step. A slightly louder version almost always sounds better at first, which tricks people into keeping boosts they do not need; use the plug-in's output gain to even them out.

Order within a processing chain also matters. A common order for speech is high-pass, corrective cuts, compression, then any gentle tonal boosts and a de-esser. Compression raises quiet parts, so cleaning up mud and rumble first stops the compressor from reacting to them. The settings for that stage are covered in compression for voice.

A worked example

Hypothetical: a home podcast with a dynamic microphone

A host records in a spare bedroom with a dynamic microphone about 5 cm from the mouth. The raw track sounds boomy and slightly boxy. The editor sets a high-pass filter at 90 Hz with a 12 dB per octave slope, which removes desk thumps without thinning the voice. Sweeping a narrow boost finds a build-up near 250 Hz, so she cuts 3 dB there with a medium width, and a further 2 dB near 700 Hz to reduce the cupboard sound. After light compression the voice is clear but a little distant, so she adds a broad 2 dB lift around 3.5 kHz. She skips the air shelf because the room has a faint fan noise. Bypassing the whole chain, the processed version sounds closer and less muddy, and still sounds like the same person.

Notice how small the moves are. If an EQ setting needs boosts or cuts of 10 dB or more to sound acceptable, the problem usually lies upstream: microphone distance, room, or the microphone itself.

EQ for listeners versus EQ before transcription

Speech recognition models are trained on a very wide range of recordings, from studio audio to phone calls, so they do not need a polished, broadcast-style tone. What helps a recognizer is a voice that stands clearly above noise and reverberation. In practice, that means EQ before transcription should be modest:

  • A high-pass filter is worth applying, because it removes rumble that adds nothing to speech.
  • A small low-mid cut can help when proximity effect or room boom blurs consonants.
  • Large presence boosts do not make a recognizer more accurate and can exaggerate hiss and harshness.
  • Air and sparkle above about 8 kHz matter little for recognition, and many systems resample audio to 16 kHz, which cannot represent anything above 8 kHz at all.
  • Noise and echo are rarely fixed by EQ. The broader picture is in how background noise affects speech recognition.

EQ for a podcast or video audience is a different job: it is about comfort, consistency between speakers and a pleasant tone across the whole episode. You can make that version for listeners and still send an unprocessed or lightly processed copy for transcription.

Common EQ mistakes and their limits

  • Boosting the low end for warmth on a recording that already has proximity effect, making it boomier.
  • Applying a large presence boost to compensate for distance. The voice gets harsher, but the room sound and noise rise with it.
  • Using EQ to fight broadband noise. A noise floor covers many frequencies, so cutting one band only changes its color.
  • Copying a preset or a chart of settings without listening. Frequencies in any guide, including this one, are starting points.
  • Making decisions on laptop speakers or only on headphones. Check on at least two playback systems.
  • Expecting EQ to remove reverberation. It can reduce the boom of a room, but not the echo.

Where EQ fits with mydubly

mydubly has no EQ or audio repair controls, so any tonal work happens in your editor before you choose the file. When you start a job, the page decodes the soundtrack in your browser, mixes it down to mono at 16 kHz and sends it in roughly 30-second chunks of compressed audio for recognition. Because of that 16 kHz sample rate, boosts above 8 kHz never reach the recognizer, while rumble filtering and clearing low-mid mud carry through.

If you are dubbing, keep in mind that the translated video replaces the original voice with a synthesized one, keeping only the music and effects underneath, so the EQ you applied to the original voice does not shape the dubbed voice. Your EQ still matters for recognition, which drives the translated script, the subtitles and both transcripts. A 20-minute recording costs 20 credits (2¢) to transcribe on mydubly's audio to text tool.

Next step

Take one recent recording and apply only a high-pass filter and, if needed, a single small low-mid cut. Compare it with the original at matched volume. If the voice is clear, stop there and transcribe it; the audio transcription guide walks through the rest. If it still sounds dull, boxy or harsh after small moves, fix the microphone position or the room before reaching for bigger EQ settings.

Frequently asked questions

What is a good starting EQ setting for a speaking voice?

A high-pass filter around 80–100 Hz, a small cut of a few decibels somewhere between 200 and 500 Hz if the voice is boomy, and an optional broad lift of 1 to 3 dB around 2–5 kHz if it sounds distant. Treat these as starting points and adjust by ear. Many clean recordings need only the high-pass filter.

Should I EQ before or after compression?

Corrective EQ such as the high-pass filter and mud cuts usually goes before compression, so the compressor does not react to rumble and boom. Gentle tonal boosts often go after it, because compression changes the balance you hear. Either order can work; the key is not to compress energy you plan to remove anyway.

Why does my voice sound thin after EQ?

Usually the high-pass filter is set too high or the low-mid cut is too deep or too wide. Lower the high-pass cutoff until the voice regains body, and reduce the cut to a few decibels. Compare with the EQ bypassed at matched loudness to check you are actually improving the sound.

Does EQ improve transcription accuracy?

Modest EQ can help a little, mainly by removing rumble and reducing boom that blurs consonants. It does not fix noise, echo or distant microphones, which affect recognition much more. Large boosts in the presence or treble region usually make no difference to a speech recognizer and can make hiss worse.

Can EQ remove background noise?

Only noise that sits in a narrow band, such as low rumble or a specific whine, can be reduced with EQ without harming the voice. Broadband noise like hiss, fans or traffic overlaps with speech frequencies, so cutting it also cuts the voice. Dedicated noise reduction or a better recording is the answer for those.

Do I need different EQ for male and female voices?

Not different rules, but different numbers. Higher voices generally tolerate a higher high-pass cutoff, and their low-mid build-up may sit a little higher. Listen to each speaker separately, especially on a multi-microphone recording, and EQ each track on its own before mixing them.