Why dialogue gets lost for hard-of-hearing viewers
Most hearing loss is not simply hearing everything more quietly. Age-related and noise-induced hearing loss commonly reduce sensitivity to higher frequencies first, and those frequencies carry many consonants: the s, f, t, k and th sounds that separate fifteen from fifty. Vowels may still come through, so speech sounds present but blurred. Turning the volume up makes the vowels and the music louder too, which does not restore the missing detail and soon becomes uncomfortable.
The second problem is separating one sound from another. A listener with typical hearing can follow a voice in a busy scene by picking it out from music and noise. Many people with hearing loss, including hearing aid and cochlear implant users, find that much harder. A level of background music that feels atmospheric to the editor can be the difference between following a sentence and guessing at it.
That is why clarity is a mixing decision rather than a volume setting. The edit decides how much separation the viewer gets, and no playback device can fully undo a mix in which music sits on top of the words.
Keep speech clearly above everything else
The most useful single rule is to give speech a consistent lead over every other element. Editors often judge balance on good headphones in a quiet room, where everything sounds clear. Viewers watch on laptop speakers, in kitchens, with a fan running, or through hearing aids.
- Set the dialogue level first, then bring music and effects up underneath it, rather than starting with the music and fitting the voice in.
- Judge the balance at a low playback level as well as a normal one. If speech disappears before the music does as you turn down, the music is too loud.
- Listen once with the picture hidden. Without lip movements and on-screen context you hear what a viewer who cannot see well, or who looks away, actually receives.
Some broadcasters publish guidance for hard-of-hearing audiences that recommends keeping background sound well below speech, but there is no single figure everyone uses, and loudness meters measure overall level rather than intelligibility. A listening test with real listeners is more reliable than any number. How overall loudness targets work is a separate subject, covered in loudness normalization explained.
Music beds under dialogue
Music is the most common reason dialogue becomes hard to follow, because it occupies the same frequency range as the voice and keeps changing.
- Prefer instrumental tracks under speech. Sung lyrics compete directly with spoken words, and the listener's attention is pulled between the two.
- Choose arrangements that are sparse in the midrange. Pads, low strings and soft percussion leave more room than busy guitar or piano parts in the voice range.
- Duck the music while someone speaks and let it rise in the pauses, with gentle fades so the movement is not distracting.
- Use EQ to carve a modest dip in the music around the upper midrange, roughly where consonant detail sits, instead of pulling the whole track down further.
- Drop the music entirely for critical information: instructions, safety points, names, numbers and anything a viewer must act on.
Competing sounds and busy scenes
Effects, room tone and location noise can mask speech as effectively as music.
- Keep loud effects such as doors, notifications, whooshes and applause off key words. Shifting them a fraction of a second into a gap is usually enough.
- Keep ambience low and steady. Fluctuating crowd noise or traffic is harder to listen through than a constant, quiet background.
- Trim overlapping speech where meaning matters. Crosstalk is hard for everyone and very hard for listeners with hearing loss.
- Reduce reverberation at the source. A voice recorded far from the microphone in a hard-walled room smears consonants, and processing afterwards only partly helps. The same factors trouble machine transcription, as described in background noise and speech recognition.
Compression and dynamic range
Wide level swings hurt clarity. If a presenter drops their voice at the end of sentences, or a guest is much quieter than the host, viewers either miss the quiet parts or turn up and get blasted by the loud ones.
Moderate compression on the dialogue track evens out these swings so quiet words come up and loud ones come down. Adjusting clip gain on individual phrases before compression often gives a more natural result than a heavy compressor alone. The goal is speech that sits at a fairly stable level from the first minute to the last.
Over-compression has its own costs: breaths and background noise rise in every pause, and the voice sounds flat and tiring. Compressing the full mix can also make music pump around the voice, so process dialogue and music as separate tracks.
How playback devices change what viewers hear
- Small laptop and phone speakers reproduce little bass but plenty of midrange, so music with energy in the voice range competes harder than it does on studio monitors.
- Many TVs and soundbars offer dialogue or speech enhancement modes. They help, but they work best when the mix already gives speech a lead.
- Surround mixes usually place dialogue in the center channel. A badly balanced stereo downmix can bury it, so check the stereo version on its own.- Some viewers listen in mono, through one earbud or because of hearing loss on one side. Hard-panned dialogue, or stereo effects that cancel when folded down, can vanish. Keep speech centered and test a mono fold-down.
Captions complement a clear mix, they do not replace it
Captions are essential for many deaf and hard-of-hearing viewers, and many hearing viewers use them too. They are still not a substitute for a clear mix. Many people with mild or moderate hearing loss prefer to listen and use captions as backup for the words they miss. Reading continuously is tiring, captions carry tone and emphasis only partly, and some viewers cannot read fast enough to keep up.
Plan for both: a mix that lets people who rely on hearing follow the speech, and accurate captions for people who rely on reading. For accessibility use, captions should include relevant non-speech information such as sound effects and music cues, and a person should review them. The difference between same-language captions and translated subtitles is set out in captions vs subtitles.
Remixing a translated voice track
The same principles apply when a video gets a new voice track in another language. A dubbed voice that is technically clean can still be hard to follow if music is pushed up underneath it, or if its level differs from sections that kept the original audio. If you want full control over that balance and still have the original project, mix the dub yourself in an editor:
- Export a dialogue-only version from the project and dub that, so the translated audio comes back as essentially the voice alone. Layering stems over the audio of a normal dub would double the music, because that file already contains the original background.
- Place the translated audio first and set it to a comfortable, consistent level.
- Bring the clean music and effects stems from the project in underneath, ducking under speech as you would for the original language.
- Check lines where the translated speech runs longer than the original. Music swells that used to fall in a pause may now land on words.
Editor-specific steps are in editing dubbed audio in a video editor, and how a dub keeps the original music, and where that has limits, is explained in keeping background music in translated videos.
A charity edits a six-minute appeal with a piano bed running from start to finish. An older supporter who wears hearing aids says she could not follow the two interview sections. The editor removes the music under both interviews, keeps it on the intro and closing montage, raises the quieter interviewee's clip gain and moves a door slam out of a key sentence. Nothing else changes, and at the next screening the same supporter follows every answer.
Mistakes and trade-offs to watch for
- Mixing only on studio headphones at a high level, then publishing without a small-speaker check.
- Expecting platform normalization to fix the balance. Normalization changes the level of the whole file; it does not change the relationship between voice and music.
- Using aggressive noise reduction, which can make speech watery and harder to understand than the original noise.
- Treating captions as permission for a muddy mix.
- Overcorrecting. Removing all music and ambience can make a video feel empty. The aim is a clear lead for speech where speech matters, not silence everywhere else.
- Ignoring the trade-off. A clarity-first mix can sound less cinematic; for informational content clarity usually matters more.
What mydubly does and does not do for dialogue clarity
mydubly is not a mixing tool, but its dubbing output interacts with these decisions. When you dub a video, AI vocal separation removes the original speech, and the remaining music, ambience and effects are mixed under the new voice and automatically lowered while it speaks. The translated MP4 keeps the original picture with that mix, and the audio download, which is AAC audio in an M4A file, is the same mix; the output is mono. Separation is not perfect: in dense or loud mixes, faint traces of the original voice can remain and the background can sound slightly thinner than the original, and sung vocals are removed along with the speech. If you have your edit project, dubbing a dialogue-only export and mixing the translated voice with your clean music and effects stems, as described above, still gives the cleanest, fully controllable result.
Generated voice clips are loudness-normalized so lines sit at a consistent level, but mydubly does not offer manual mixing or ducking controls, EQ or a dialogue enhancement setting. The whole video uses one of 8 stock voices, with no per-speaker voices. A dubbed job also gives target-language SRT and VTT subtitles that follow the dubbed lines. They contain recognized speech only, with no sound tags, music descriptions or speaker labels, so add those cues and review the text by hand before using the files as accessibility captions. The AI dubbing page and the guide to adding subtitles to a video cover the practical steps.
Next step: run a clarity check on one video
Pick a video with music under speech and play it on a phone speaker at low volume, then once in mono, then once with the picture hidden. Note every sentence you have to replay. Those are the places to duck the music, move an effect or raise the voice, and the same check applies to a dub. A two-minute test dub costs 100 credits (10¢).
Frequently asked questions
Why do hard-of-hearing viewers struggle with background music more than other viewers?
Hearing loss often reduces the ability to separate one sound from another, not just the ability to hear quiet sounds. Music that a typical listener tunes out can mask the consonants a hard-of-hearing listener depends on. Hearing aids amplify the music along with the speech, so the balance in the mix still decides what gets through.
Should I remove music from my videos completely?
Not necessarily. Music on intros, transitions and pauses rarely causes problems. The trouble comes from continuous beds under talking, especially tracks with vocals or busy midrange parts. Removing or ducking music under speech, and dropping it entirely under critical information, keeps most of the atmosphere while protecting the words.
Does turning up the volume fix unclear dialogue?
Only partly. Raising the volume raises music and noise by the same amount, so the balance between speech and background stays the same. For listeners who miss high-frequency consonants, louder vowels and music can make the experience more uncomfortable without making words clearer. The fix belongs in the mix.
What level should background music sit at under speech?
There is no single agreed figure. Some broadcaster guidance for hard-of-hearing audiences asks for background sound well below speech, and the right amount depends on the music and the voice. Set it by ear at low volume on a small speaker, then confirm with listeners who have not heard the script.
Are captions enough if the audio is hard to follow?
Captions are essential for many deaf viewers, but many people with mild or moderate hearing loss prefer to listen and use captions only as backup. Reading every line is tiring, and some viewers cannot read quickly. Accessible video usually needs both a clear mix and accurate, reviewed captions.
Will a dubbed video from mydubly keep my original music?
Yes. AI vocal separation removes the original speech, and the remaining music and effects are mixed under the new voice and lowered automatically while it speaks. Separation is not perfect: faint traces of the original voice can remain in dense or loud mixes, and sung vocals are removed along with the speech. For full control, export a dialogue-only version from your project, dub that, and mix the returned translated voice with your clean music stem in an editor, applying the same ducking and balance checks you would use for the original language.