Audio engineering for speech

Harsh S Sounds in Voice Recordings: Causes, De-Essers and Technique

To reduce sibilance, fix the cause first: aim the microphone slightly off the mouth, back away a little and remove any presence or treble boost that is exaggerating s sounds. Then use a de-esser, which turns down only the sibilant frequencies at the moments they get too loud, typically by a few decibels. Keep the reduction modest, because over-de-essing makes speakers sound as if they have a lisp. Do this before exporting to a lossy format, since harsh s sounds are among the first things low-bitrate encoders smear.

8 min read · Updated

What sibilance is

Sibilants are the hissing and hushing consonants: s, z, sh, ch, and the sound in the middle of "measure". They are made by forcing air through a narrow gap against the teeth, which produces noise-like energy high in the frequency range. Roughly speaking, sh sounds sit lower and s sounds higher, with most of the energy somewhere between about 4 and 10 kHz depending on the speaker.

Some sibilance is necessary. Without it, "sixty" and "fifty", or "cats" and "cat", become hard to tell apart. The problem is excess: s sounds that jump out several decibels louder than the vowels around them, sound piercing on earbuds, and make a listener wince. That is what people mean by a sibilant recording.

What makes s sounds harsh

Sibilance usually builds up from several small contributions rather than one big fault. The common causes:

  • The microphone. Many condenser microphones have a presence peak or bright top end that flatters voices but exaggerates s sounds. Some speakers and some microphones are simply a bad match.
  • Placement. Pointing the microphone straight at the mouth puts the narrow, directional stream of sibilant energy right on the capsule.
  • The speaker. Teeth, dental work and speaking style all change how much sibilance a voice produces.
  • EQ. A presence or air boost added to make a voice clearer also lifts the s region. Tonal choices for speech are covered in EQ for voice recordings.
  • Compression. A compressor turns loud vowels down and then makeup gain raises everything, so short sibilants end up louder relative to the rest. Fast limiters can add their own edge.
  • Lossy encoding. Low-bitrate MP3, AAC or Opus files can turn sibilants into swishy, watery or metallic artifacts, making s sounds feel harsher than in the original.
  • Playback. Bright earbuds and small speakers can exaggerate sibilance that sounds fine on studio monitors.

Prevent it at the microphone

Technique is the cheapest fix and it does not dull the voice:

  1. Aim the microphone at the chin or the corner of the mouth rather than straight at the lips. Sibilant energy is directional, so a small angle reduces it noticeably.
  2. Back off a little. Moving from very close to a moderate distance softens s sounds and plosives at the same time.
  3. Try a different microphone if you have one. A darker dynamic microphone often suits a naturally sibilant voice better than a bright condenser.
  4. Check the microphone's own settings. Some have presence-boost switches that are worth turning off for sibilant speakers.
  5. Record a test line full of s sounds, for example "She sells sixty silver scissors on Sunday", and judge it on earbuds as well as headphones.

How a de-esser works

A de-esser is a compressor that listens only to the sibilant frequency range. When energy in that range rises above a threshold, it turns the signal down, then releases as soon as the s is over. Vowels and other consonants pass untouched because they do not trigger it.

There are two main designs. A wideband de-esser turns down the whole signal briefly when an s is detected; it can sound natural on mild sibilance but dulls everything during the reduction. A split-band de-esser reduces only the high-frequency band, leaving the body of the voice untouched; it is the more common choice for speech. Dynamic EQ, an equalizer band that only cuts when its level exceeds a threshold, does essentially the same job and is a good alternative if your editor has it.

Setting a de-esser step by step

  1. Find the frequency. Most de-essers have a listen or sidechain monitor button that lets you hear only what is triggering it. Sweep the frequency until you hear mostly the s sounds. For many voices this lands somewhere between about 5 and 8 kHz; sh-heavy voices may need a lower setting.
  2. Set the threshold. Lower it until the de-esser acts on the harshest s sounds but stays idle on normal speech. Watch the gain-reduction meter: it should flicker on sibilants, not sit permanently active.
  3. Limit the range. If the plug-in has a range or maximum reduction control, start with a few decibels. Gentle reduction on many sibilants usually sounds better than heavy reduction on a few.
  4. Keep the timing fast. Attack and release should be short so that only the sibilant is affected; many de-essers handle this automatically.
  5. Compare with bypass at the same volume, in context, on earbuds. The voice should still sound like itself, with s sounds that no longer jump out.

Placement in the chain matters. Because compression raises sibilance, many editors put the de-esser after the compressor, or use one gentle de-esser before and another after. Try both and listen. How compression changes the balance of a voice is explained in compression for voice.

For a short recording with a handful of piercing s sounds, manual clip gain is also fine: select each sibilant and lower it by a few decibels. It is slow but precise.

Over-de-essing and how to spot it

The most common de-essing mistake is doing too much. Signs that you have crossed the line:

  • S sounds turn into something closer to th, giving a lisping effect.
  • The voice sounds dull or muffled whenever the speaker says certain words.
  • Words like "sixty" and "fifty" become harder to distinguish.
  • Breaths and other soft high-frequency sounds disappear or pump.

If you hear any of these, raise the threshold, reduce the range or narrow the frequency band. It is better to leave a little sibilance than to remove the consonants listeners need.

Hypothetical: a bright voice-over with a compressor

A creator records explainer voice-overs with a bright condenser microphone pointed straight at the mouth, adds a 4 dB presence boost and compresses heavily. Viewers mention harsh s sounds on phone speakers. The fix comes in three steps. First, she angles the microphone toward her chin and drops the presence boost to 1 dB, which removes most of the problem at no cost. Second, she adds a split-band de-esser after the compressor around 6.5 kHz with a few decibels of reduction. Third, she exports the final audio as AAC at a moderate bitrate rather than a very low one. The s sounds now sit with the rest of the voice, and nothing sounds lisped.

Sibilance, lossy encoding and transcription

Lossy codecs spend their limited bits where hearing is most sensitive, and noise-like high-frequency sound is hard to represent efficiently. At low bitrates, encoders may also cut off the highest frequencies entirely. The result is that harsh sibilants turn into swirling artifacts or lose their crispness. De-essing before you encode, and avoiding very low bitrates for final versions, keeps s sounds clean. How bitrate affects speech files is explained in audio bitrate for speech recognition, and the speech-oriented codec many browsers use is covered in the Opus audio codec.

For transcription, harsh sibilance is rarely a real problem. A recognizer does not get fatigued, and s sounds that are too loud are still clearly s sounds. Many speech recognition systems, including Whisper-family models, work on audio at 16 kHz, which keeps content up to 8 kHz and discards anything above. Over-de-essing is the bigger risk: if sibilants are pushed down so far that they blur into other sounds, distinctions such as plurals can suffer.

Limits of de-essing

  • A de-esser cannot fix sibilance that was distorted by clipping or a harsh limiter; it only reduces level.
  • Very sibilant voices recorded on a bright microphone straight on may need re-recording for professional results.
  • Static EQ cuts in the s region dull every word and are a poor substitute for a de-esser.
  • Sibilance baked into a low-bitrate file with encoding artifacts cannot be cleaned up by de-essing the decoded audio afterwards.
  • Every voice is different; frequency ranges in any guide are only a starting point.

Sibilance and mydubly

mydubly has no de-esser or audio repair controls, so do this work in your editor before you choose the file. When a job starts, the page decodes the audio in your browser, mixes it to mono at 16 kHz and encodes roughly 30-second chunks to Opus at about 32 kb/s for upload. That encoding is designed for the recognizer, not for listening, and it is the reason extreme high-frequency detail does not matter for your transcript.

If you are dubbing a video, the translated version replaces the original voice with a synthesized one, so sibilance in your recording does not appear in the dubbed track. It can still matter for the original-language version you publish, and if over-de-essed, for recognition accuracy of the words the translation and subtitles are built from. Subtitle files from the subtitle generator follow the recognized speech, so checking a few sibilant-heavy lines is quick.

Next step

Record your test line with the microphone angled slightly off-axis and listen on earbuds. If s sounds still jump out, add a split-band de-esser with a few decibels of reduction and compare with bypass. Once the voice sounds natural, transcribe it with mydubly's audio to text tool; a 15-minute recording costs 15 credits (1.5¢). If your final export is an MP3, the MP3 to text page covers that format specifically.

Frequently asked questions

What frequency should I set a de-esser to?

It depends on the voice, but many speakers' s sounds fall somewhere between about 5 and 8 kHz, with sh sounds lower. Use the de-esser's listen mode and sweep until you hear mostly sibilants. Trust your ears over any fixed number.

Should the de-esser go before or after compression?

Often after, because compression with makeup gain tends to raise sibilance. Some editors use a gentle de-esser before the compressor and another after. Try both orders on the same passage and keep whichever sounds more natural at matched volume.

How much de-essing is too much?

If s sounds start to resemble th, or the voice dulls on certain words, you have gone too far. A few decibels of reduction on the harshest sibilants is a common starting point. Leaving slight sibilance is better than removing consonants listeners need.

Can EQ reduce sibilance instead of a de-esser?

A static EQ cut reduces sibilance but also dulls every other sound in that range for the whole recording. A dynamic EQ band that only cuts when the level crosses a threshold works much like a de-esser and is a good option. Ordinary EQ is better used to avoid creating sibilance with unnecessary boosts.

Why does my audio sound more sibilant after exporting to MP3?

Low-bitrate lossy encoding struggles with noise-like high-frequency sounds and can add swishy or metallic artifacts to s sounds. Export at a moderate bitrate, de-ess before encoding, and avoid repeated lossy re-encoding of the same file.

Does sibilance affect speech recognition?

Usually not much; loud s sounds are still recognized as s sounds. Over-de-essing is more likely to cause problems, because it can blur distinctions such as plurals. Aim for natural, balanced sibilants and check a few passages in the transcript.