Audio & video formats

Loudness Normalization: LUFS, True Peak and Consistent Speech Levels

Loudness normalization adjusts audio so it reaches a target perceived loudness, measured in LUFS, rather than a target peak level. Two clips that peak at the same level can sound very different in loudness; two clips normalized to the same LUFS value sound roughly equally loud. Broadcasters and many streaming platforms work with loudness targets, and for speech the practical goal is simple: every line, clip and episode at a steady, comfortable level without clipping.

8 min read · Updated

Peak level is not loudness

Digital audio has a hard ceiling, 0 dBFS (decibels relative to full scale). Peak meters show how close the highest samples come to it, which matters because anything pushed past the ceiling clips. But peaks say little about how loud something sounds. A single sharp hand clap can peak near 0 dBFS while sounding brief and quiet overall; a heavily compressed voice can peak lower and still sound much louder, because its energy is sustained.

Peak normalization scales a file so its highest peak hits a chosen level. It guarantees headroom but not consistency: normalize ten speech clips to the same peak, and they will still jump up and down in perceived loudness when played in a row. That is the problem loudness normalization solves.

How LUFS measures loudness

LUFS, loudness units relative to full scale, comes from the ITU-R BS.1770 recommendation, which defines a measurement designed to track how loud people perceive programme material. In short:

  • The audio passes through a K-weighting filter, which reduces low frequencies and slightly boosts upper-mid frequencies, roughly mimicking the ear's sensitivity.
  • The energy is averaged over time and expressed on a decibel-like scale. LKFS, used in some North American standards, is the same unit under another name.
  • For a whole programme, gating excludes silence and very quiet passages from the average, so long pauses in speech do not drag the figure down.
  • A change of 1 LU (loudness unit) corresponds to a change of 1 dB.

Several readings come from the same measurement:

Integrated loudness
The gated average over the whole file. The number targets refer to
Short-term loudness
A sliding three-second window, useful for spotting loud and quiet sections
Momentary loudness
A 400-millisecond window that reacts almost instantly
Loudness range (LRA)
How much the loudness varies across the programme, in LU
True peak (dBTP)
An estimate of the highest level the waveform reaches between samples, which can exceed sample peaks after playback conversion or lossy encoding

Loudness targets, described with care

Different organizations publish different reference levels, and they are updated over time. The figures below are commonly cited, but always check the current official documentation for the platform or broadcaster you deliver to.

  • European broadcast generally follows EBU R 128, with an integrated target of -23 LUFS and a maximum true peak of -1 dBTP.
  • United States television generally follows ATSC A/85, with a target of -24 LKFS.
  • Music and video streaming services typically normalize playback, and reference levels around -14 to -16 LUFS are often quoted for them. Not every service publishes its figure, and some apply normalization only to loud material or let listeners switch it off.
  • Podcast guidance from major directories has commonly pointed to around -16 LUFS for stereo, with true peaks kept below about -1 dBTP.

The practical lesson is that, on a platform that normalizes, making your upload louder does not make it louder for viewers. It is turned down to match, and any limiting you used to squeeze in extra loudness simply costs dynamics.

Normalization versus compression and limiting

Loudness normalization in its basic form is a single gain change applied to the whole file: everything goes up or down by the same amount. It does not change the balance between quiet and loud passages. That is both its strength, because it is transparent, and its limit.

Two situations need more than gain. If reaching the target requires raising a file so far that its peaks would exceed the true-peak ceiling, a limiter has to catch those peaks. And if a recording swings between whispering and shouting, normalization makes the average right but leaves the swings; compression is the tool that narrows them. Some normalizers offer a dynamic mode that adjusts gain over time, which is effectively automatic compression and should be listened to carefully.

Worked example: raising a quiet narration

A narration clip measures -22 LUFS integrated with a true peak of -4 dBTP, and the target is -16 LUFS with peaks no higher than -1 dBTP. Reaching the target needs +6 dB of gain, which would push the true peak to +2 dBTP, three decibels over the ceiling. The options are to let a limiter reduce the loudest moments by about 3 dB, to compress the narration gently first so it peaks less, or to accept a slightly lower target. All three are legitimate; the choice depends on how the voice sounds afterwards.

Why speech needs consistent loudness

Listeners adjust their volume once, near the start, and expect it to hold. Speech that keeps changing level forces them to reach for the control, and quiet passages get lost under background noise on phones and laptop speakers. Consistency matters at several scales:

  • Within a recording, where different speakers or microphone distances produce different levels.
  • Between segments, such as a voice-over recorded on different days.
  • Across a series, so each episode or lesson plays at the level of the last.
  • Between voice and music, where the mix must keep words clear over the bed.

Generated speech adds another case. Synthetic voice clips are produced sentence by sentence, and two consecutive clips can come out at noticeably different levels. Without normalization, a dubbed video would audibly jump from line to line, one of the artefacts discussed in AI voice quality.

How mydubly normalizes generated voice clips

In mydubly's dubbing pipeline, every generated voice clip is loudness-normalized to the same target after its leading and trailing silence is trimmed, before it is placed against the original timing. Lines and chunks therefore sit at a consistent level, so the translated voice does not jump between sentences or between the roughly 30-second sections the track is built from. The chunks are then joined into one continuous voice track and mixed over the original music and effects, which are lowered automatically while the voice speaks. The result is delivered as an M4A file and, if you choose, swapped into your video in place of the original soundtrack. How each clip is placed in time is explained in syncing translated audio with video.

The level of that voice is designed to be steady, not to match any particular platform target. If you dub a dialogue-only export and mix the translated voice with a clean music stem in an editor, measure the finished mix and normalize it to your platform's current guidance; adding a music bed always changes the total loudness. How the original music is kept, and when your own mix is worth it, is covered in keeping background music when translating video.

Measuring and normalizing a file step by step

  1. Measure first. Use your editor's loudness meter, a free loudness meter plugin, or ffmpeg: ffmpeg -i input.wav -af loudnorm=print_format=summary -f null - prints integrated loudness, true peak and loudness range.
  2. Listen through the loudest and quietest sections. If the range is very wide, apply gentle compression before normalizing.
  3. Normalize to your target. Many editors, including Audacity, offer a loudness normalization effect that accepts a LUFS value. With ffmpeg, a two-pass loudnorm run, measuring first and feeding the measured values into the second pass, gives a more accurate linear result than a single pass.
  4. Check the true peak afterwards and apply a limiter if it exceeds your ceiling.
  5. Export, then measure the exported file again, since lossy encoding can raise true peaks slightly.
  6. Finally, play the result next to one of your previous uploads on the same device as a sanity check.

What loudness normalization can't fix

  • Noise rises with the voice. Raising a quiet recording raises its hiss and room noise by exactly the same amount; see transcribing quiet audio for what can and cannot be recovered.
  • Clipping stays clipped. Lowering a distorted recording makes the distortion quieter, not cleaner.
  • Balance inside a mix is unchanged. If music drowns the voice, normalizing the mix keeps it drowned; fix the mix first.
  • Loudness is not clarity. A muddy or distant voice at the right LUFS is still muddy and distant.
  • A target is not a quality score. Meeting -16 or -23 LUFS says nothing about whether the audio sounds good.

Next step

Measure the integrated loudness and true peak of one of your published videos or episodes, and compare it with the current guidance for the platform you use. If you are producing translated versions, the AI dubbing voice track gives you a consistently leveled starting point, and the step-by-step dubbing guide covers reviewing the result before you edit or publish. A 4-minute dub costs 200 credits (20¢).

Frequently asked questions

What is the difference between LUFS and dB?

Decibels describe a ratio between levels and appear in many forms, such as dBFS for digital peak level. LUFS is a specific loudness measurement defined by ITU-R BS.1770 that weights frequencies roughly the way hearing does and averages over time with gating. A change of one loudness unit equals one decibel, so a file at -20 LUFS needs 4 dB of gain to reach -16 LUFS.

What LUFS should I use for YouTube videos?

YouTube normalizes playback but does not formally publish a target, and figures around -14 LUFS are widely quoted from measurements. Many creators aim somewhere around -14 to -16 LUFS integrated with true peaks below about -1 dBTP. Because platforms change their processing, treat any figure as a guide and check the current recommendations of the platform you publish on.

Should I normalize audio before transcribing it?

It is rarely necessary. Speech recognizers handle a wide range of levels, and a quiet but clean recording usually transcribes well. Normalizing can make review easier for people listening back, but it does not add detail and raises background noise along with the voice. If a recording is extremely quiet, the bigger gains come from better microphone placement next time.

Why does my video sound quieter on a streaming platform than in my editor?

Usually because the platform normalizes playback and turned a loud upload down, or because your editor's monitoring was set louder than the platform's playback. Many services are reported to reduce loud uploads without raising quiet ones. Measure the integrated loudness of your export; if it is well above the platform's commonly quoted reference, the reduction is expected and costs nothing but the dynamics you limited away.

Is peak normalization ever the right choice?

Yes, when the goal is headroom rather than consistent loudness: for example preparing a recording for further processing, or making sure a single file uses the available range without clipping. For anything played in sequence with other material, such as episodes, lessons or dubbed lines, loudness normalization is the better fit because it matches how loud each item sounds.