Audio engineering for speech

Compressing Voice Audio: Settings, Limiters and When to Leave It Alone

Audio compression for voice means dynamic range compression: automatically turning down the loudest parts of speech so loud and soft words sit closer together, then raising the whole thing back up. Gentle settings, roughly 2:1 to 4:1 with a few decibels of gain reduction, make a voice easier to follow over music or in a noisy room. Pushed too hard, compression raises background noise in every pause and makes levels pump. It is not file compression, and it rarely improves a transcript on its own.

9 min read · Updated

Dynamic range compression is not file compression

The word compression means two unrelated things in audio. File compression, done by codecs such as MP3, AAC and Opus, makes the data smaller; it is covered in what an audio codec is. Dynamic range compression changes the sound itself by reducing the difference between loud and quiet moments. This article is about the second kind.

Speech is naturally uneven. A talker leans in and out, emphasizes some words, trails off at the end of sentences and laughs. On a good monitor in a quiet room, that variation sounds natural. In a car, on a phone speaker or under background music, the quiet words get lost while the loud ones jump out. A compressor narrows that range so the listener doesn't have to keep adjusting the volume.

The controls, one at a time

Threshold
The level above which the compressor starts working. Anything quieter passes through untouched
Ratio
How strongly it pushes back. At 4:1, a sound 8 dB over the threshold comes out only 2 dB over
Attack
How quickly gain reduction starts once the signal crosses the threshold, usually in milliseconds
Release
How quickly the gain recovers once the signal drops back below the threshold
Knee
Whether compression begins abruptly at the threshold (hard knee) or eases in around it (soft knee)
Makeup gain
A fixed boost after compression that restores the overall level lost to gain reduction
Gain reduction meter
Shows how many decibels the compressor is removing at each moment; the most useful readout on the plugin

The threshold and ratio together decide how much work the compressor does. A low threshold with a modest ratio touches most of the speech gently; a high threshold with a steep ratio leaves most words alone and clamps only the peaks.

Attack and release decide how the compression sounds. A very fast attack catches the start of each syllable and can make consonants sound blunted. A slower attack lets the first few milliseconds through, which keeps speech crisp but lets short peaks past. A release that is too fast makes the level flutter between words; one that is too slow keeps the gain turned down through the next quiet word, so it sounds squashed.

Limiters: the ceiling, not the shaping

A limiter is a compressor with a very high ratio and a very fast attack, often with look-ahead so it can react before a peak arrives. Its job is to stop anything going above a set ceiling, such as a peak level just under full scale, not to shape the voice.

For speech, a limiter belongs at the very end of the chain as a safety net. Many recorders also have an input limiter that prevents clipping when someone suddenly shouts; leaving that on during interviews is cheap insurance. What a limiter should not do is carry the whole job of evening out a voice. Driving speech hard into a limiter produces a flat, breathless sound and audible distortion on loud words.

Sensible starting points for speech

There is no universal preset, because voices, microphones and rooms differ. These are commonly suggested starting points, to adjust by ear while watching the gain reduction meter:

  1. Clean up first. Apply a high-pass filter and any noise reduction before the compressor, so it isn't reacting to rumble.
  2. Set the ratio between 2:1 and 4:1.
  3. Lower the threshold until the meter shows a few decibels of gain reduction on normal words and more on the loudest ones.
  4. Start with a medium attack, somewhere around 10 to 30 milliseconds, and shorten it only if peaks are escaping.
  5. Set the release so the meter returns toward zero between phrases without visibly fluttering, often somewhere in the 50 to 200 millisecond range.
  6. Add makeup gain until the compressed voice is about as loud as the original at its peaks, then compare with the compressor bypassed at matched loudness.
  7. Finish with a limiter only to catch occasional peaks.

The comparison in step six matters. A louder signal almost always sounds better at first listen, so judge compression with the levels matched, not by how much more impressive it seems.

Two compressors in series, each doing a little, often sound more natural than one compressor doing a lot. Overall loudness targets for podcasts and video platforms are a separate step after compression, explained in loudness normalization explained.

When compression helps a transcript

For speech recognition, compression is mostly neutral. A recognizer cares about how clearly the voice stands out from the background, and a compressor doesn't change that ratio within a word: it turns speech and the noise underneath it down or up together.

There are a few cases where it genuinely helps:

  • One quiet participant. In a recording where one speaker is much quieter than another, light compression followed by makeup gain brings that voice up relative to the loud one. That helps the human reviewing the transcript, and it can help a recognizer whose input level sits close to its noise floor.
  • Big jumps in level. A voice that swings from a whisper to a shout is easier to proofread when the swings are narrowed.
  • Preparing a listening copy. If you will check a transcript against the audio, a gently compressed version is easier to hear on laptop speakers.

Quiet recordings in general, including when simple gain is the better fix, are covered in transcribing quiet audio.

When compression makes things worse

Compression has a predictable cost: everything below the voice comes up too. With a low threshold, heavy ratio and lots of makeup gain, the hum, room tone and breaths in every pause rise toward the level of the speech. The result sounds noisy even though no new noise was added.

Common symptoms of overdoing it:

  • Pumping. Background noise audibly swells in the gaps after each phrase as the gain recovers, then ducks when the voice returns.
  • Breathing. Breaths between sentences become nearly as loud as the words.
  • Flattened delivery. Emphasis disappears, and the talker sounds tired or shouted at the same time.
  • Distorted consonants. A very fast attack and release can add a gritty edge to s and t sounds.

For transcription, raised noise in pauses has a specific downside. Recognizers occasionally produce words from noise or silence, and a recording whose pauses have been pushed up toward speech level offers more material to misread. If you compress, consider a gentle downward expander or gate before the compressor to keep pauses quiet, and set it carefully so it does not chop off quiet word endings.

Compression in a dubbed mix

Dubbing adds a different reason to compress. When a translated voice is mixed back with music and sound effects, the voice has to stay intelligible over a bed that keeps changing. Light compression on the voice, plus ducking that lowers the music while someone speaks, is the standard approach in most editors. Getting that balance right is covered in editing dubbed audio in a video editor.

Hypothetical: a remote interview with uneven voices

A podcaster records a 30-minute interview. The host sits close to a good microphone; the guest joins over a call and arrives about 8 dB quieter, with some fan noise. The podcaster high-passes both tracks, then sets the guest's compressor at 3:1, a 20 ms attack and a 120 ms release, with the threshold low enough to take off around 4 dB on normal words. After makeup gain the two voices sit close together. A first attempt at 8:1 with 12 dB of makeup gain made the guest's fan roar in every pause, so the podcaster backed off. The transcript of the gently compressed mix reads the same as before; the listening experience is what improved.

Limits of compression and common mistakes

  • Compression cannot recover clipped audio. Peaks that were cut off at the recorder stay cut off.
  • It cannot separate a voice from noise, reverberation or another speaker; it only changes how levels move over time.
  • It is not a loudness target. A compressor makes speech more even; meeting a platform's loudness specification is a separate measurement.
  • Compressing before noise reduction makes noise reduction harder, because the noise level now changes from moment to moment.
  • Presets copied from music production often use ratios and attack times too aggressive for dialogue.
  • Applying compression to a file before archiving it removes options. Keep an unprocessed master.

Where compression fits with mydubly

mydubly accepts audio and video files up to 2 hours long and works on the speech in them. In the browser, the audio is decoded, mixed down to mono 16 kHz, cut into chunks of about 30 seconds and compressed to Opus at 32 kb/s for upload over HTTPS. That is file compression for transport, not dynamic range compression of your voice. If you want a compressed sound, apply it in your editor first; for a transcript, an unprocessed recording at a good level is usually just as good.

For dubbing, the generated voice clips are loudness-normalized, and the server assembles them into a dubbed AAC track. The original speech is removed by AI vocal separation, and the remaining music and effects are mixed under the translated voice and ducked automatically while it speaks. If you want full control over that balance, download the dubbed audio, mix it with your clean music and effects track in an editor, and apply any compression and ducking there. AI dubbing describes what the dubbed output includes, and the dubbing guide walks through the steps.

Next step: try a gentle compressor on one file

Pick a recording where levels jump around, add a compressor at 3:1 with a medium attack, and aim for a few decibels of gain reduction. Compare it with the original at matched loudness and listen especially to the pauses. If the goal is a transcript, upload either version to the audio to text tool and spend your effort on the recording rather than the processing. A 30-minute recording costs 30 credits (3¢) to transcribe.

Frequently asked questions

What compressor ratio should I use for voice?

Ratios between 2:1 and 4:1 are common starting points for speech. Lower ratios sound more natural; higher ratios hold a voice more firmly but make compression more audible. Watch the gain reduction meter, and aim for a few decibels on normal words rather than a fixed number.

Should I compress audio before transcribing it?

Usually it isn't necessary. A recognizer depends on how clearly the voice stands out from the background, and a compressor doesn't change that within a word. Light compression can help when one speaker is much quieter than another, but heavy compression raises noise in pauses, which can make results worse.

What is the difference between a compressor and a limiter?

A limiter is an extreme compressor with a very high ratio and a very fast attack, designed to stop peaks from exceeding a ceiling. A compressor shapes the overall dynamics of a voice more gently. In a speech chain, compression does the evening out and a limiter at the end catches the occasional stray peak.

Why does my compressed audio sound like it is breathing?

That is pumping: the compressor turns the gain down when the voice is loud and lets it recover in pauses, so background noise swells after each phrase. Raise the threshold, lower the ratio or reduce makeup gain. A slower release or a gentle expander before the compressor also helps.

Does compression make a voice louder?

Compression itself reduces the level of loud passages. The overall loudness only rises when you add makeup gain afterwards, which is possible because the peaks are now lower. Because louder audio tends to sound better at first, compare processed and unprocessed versions at the same loudness.

Is compression the same as loudness normalization?

No. Compression changes the dynamics within a recording, making loud and soft moments closer together. Loudness normalization adjusts the whole file up or down to reach a measured average loudness without changing its internal dynamics. Many workflows use both, compression first and normalization last.