Audio engineering for speech

How Acoustic Echo Cancellation Works, and What It Does to Recorded Calls

Acoustic echo cancellation is the processing in calling software that stops your loudspeaker's sound from being picked up by your microphone and sent back to the other person. It predicts the echo from the audio it just played and subtracts it, then suppresses whatever is left. When both people talk at once, it can clip word beginnings, briefly mute one side or make levels pump, and those artifacts end up in recordings and transcripts. Headphones remove most of the problem at the source.

8 min read · Updated

The problem echo cancellation solves

On a call without headphones, the other person's voice comes out of your speakers, bounces around your room and enters your microphone. Without processing, that sound would travel back to them a fraction of a second later, so they would hear themselves echo. With two open microphones and two sets of speakers, the loop can even build into howling feedback.

Telephone handsets avoided this mostly by design: the earpiece sits against your ear and the microphone is near your mouth. Laptops, phones on speaker, conference-room systems and smart speakers put a loudspeaker and a microphone a few centimetres apart in an open room. Every one of them needs echo cancellation to be usable.

How an echo canceller works

Echo cancellation has an advantage over ordinary noise reduction: it knows exactly what the echo should sound like, because it has the signal it just sent to the speaker. That signal is called the far-end reference.

  1. The canceller keeps a copy of the far-end audio going to the loudspeaker.
  2. An adaptive filter learns the echo path: how the speaker, the room's reflections and the microphone change that sound, and how long it takes to come back.
  3. It applies that learned path to the reference to predict the echo arriving at the microphone.
  4. It subtracts the predicted echo from the microphone signal.
  5. A second stage, often called residual echo suppression or nonlinear processing, turns down whatever echo the filter missed, especially when the near end is silent.
  6. Noise suppression and automatic gain control usually follow in the same chain before the audio is encoded and sent.

The adaptive filter has to keep learning, because the echo path changes whenever someone moves, a door opens or the volume changes. Hardware also matters. Cheap speakers distort at high volume, and distortion is not predictable from the reference, so it passes through the linear filter and has to be suppressed by the second stage.

Double-talk: the hard part

Double-talk is when both ends speak at once. It is the hardest situation for an echo canceller. The microphone now carries your voice and the echo of theirs, and the filter must not mistake your voice for echo and start adapting to it.

Echo cancellers use double-talk detection to spot those moments and freeze or slow the filter's learning. The residual suppressor faces a dilemma at the same time. Suppressing hard removes any trace of echo but also attenuates your voice; suppressing lightly keeps your voice but lets some echo through. Different products strike that balance differently, and the balance shifts with room, hardware and volume.

The result in practice is that a quick interjection like "right" or "sorry, go on" may come through quieter, cut short or not at all while the other person is still speaking.

The artifacts echo cancellation leaves behind

These are the audible traces of the trade-offs above, and they are baked into any recording made from the call's audio:

Clipped word starts
The first syllable after the other person stops is attenuated while the suppressor releases, so words begin abruptly or lose their opening consonant
Half-duplex gaps
One side is heavily suppressed while the other talks, so the call behaves like a walkie-talkie and overlapping speech disappears
Pumping
Background noise and voice level rise and fall as suppression switches on and off
Residual echo
A faint, sometimes metallic copy of the far-end voice slips through, especially after volume changes or with distorting speakers
Warbling or underwater tone
Aggressive suppression combined with noise reduction removes parts of the voice spectrum unevenly

Network problems add their own glitches, such as robotic stretches and dropouts, which are a separate issue from echo cancellation.

Why headphones help so much

Echo cancellation exists because the microphone can hear the loudspeaker. Headphones remove that path. With nothing for the microphone to pick up, there is little echo to cancel, the residual suppressor has little reason to act, and double-talk stops being a problem. Interruptions, laughter and backchannel words come through intact.

Wired earbuds or closed headphones work well. Open-back headphones at high volume can still leak a little. Bluetooth headsets solve the echo problem but often switch to a lower-quality call profile when their built-in microphone is in use; a wired or USB microphone with Bluetooth headphones avoids that on many systems.

Some calling apps offer a setting that disables echo cancellation and other processing, often named something like original sound. It suits people wearing headphones with good microphones, and it backfires badly with laptop speakers. Check your platform's current documentation for what is available.

What echo cancellation does to recorded calls and transcripts

A cloud or host recording captures the call after processing, so every artifact above is in the file. Speech recognition copes well with clean call audio, but the artifacts have specific effects:

  • Clipped onsets make the first word of a turn harder to recognize. Short function words and names at the start of a sentence are the usual casualties.
  • Half-duplex gaps remove overlapping speech entirely. The recognizer cannot transcribe a word that was suppressed before it was recorded, so agreements, objections and questions asked over someone else go missing.
  • Residual echo can make the recognizer hear the far-end voice twice, producing repeated words or a stray phrase.
  • Pumping and suppression raise and lower background noise, which can make quiet stretches more prone to stray text.

None of these are recognition errors in the usual sense; they are gaps and distortions in the audio itself. Missing-word patterns and how to diagnose them are covered in transcription missing words. Narrowband audio from traditional phone lines is a different problem, explained in phone call recording transcription.

Hypothetical: one call, two recordings

A researcher interviews a participant over a video call. She wears earbuds; the participant uses laptop speakers. The platform's cloud recording sounds fine until the two overlap: whenever she says "mm-hm" during an answer, it vanishes, and the participant's sentences after each of her questions start with a clipped syllable. The transcript reflects exactly that. For the next session, she asks the participant to plug in any wired earphones, and she records a local backup of her own microphone. Overlaps now survive in the recording, and the transcript keeps her prompts.

Getting cleaner call recordings

  1. Ask everyone to wear headphones or earbuds. This single change prevents most echo artifacts.
  2. Mute when not speaking in group calls, which reduces echo and noise paths.
  3. Keep speaker volume moderate if headphones are impossible, so the loudspeaker does not distort.
  4. Use a dedicated microphone close to the mouth rather than the laptop's built-in microphone across the desk.
  5. Record locally on each side when quality matters. A local recording of your own microphone captures your voice before the call's processing; if you use speakers, though, that local file will also contain the other person's voice from your room.
  6. Test with a short call and listen back to the recording before an important session.

A fuller guide to separate-track recording for interviews is in remote interview recording setup.

Limits and trade-offs of echo cancellation

  • It cannot be undone afterwards. Once a syllable is suppressed in the call, no editing tool can bring it back.
  • It works best for one talker at a time. Lively group discussions with frequent overlap stress every canceller.
  • Turning it off is only safe with headphones. Without them, the far end hears echo or feedback.
  • Speakerphones in large rooms are hard cases. Long reverberation gives the filter a longer echo path to model, and residual echo is more likely.
  • Processing chains differ between apps and change with updates, so a setup that worked last year may behave differently now.

Call recordings and mydubly

mydubly is not a calling app and has no live mode; it works on recordings you have already made, chosen from your device. You can upload meeting recordings as video (MP4, MOV, WebM, MKV, M4V) or audio (MP3, WAV, M4A, AAC, OGG, FLAC), up to 2 hours per file. Recognition uses Whisper with a voice activity filter that skips silence while keeping timestamps tied to the original recording.

mydubly cannot restore words removed by a call's echo canceller, and it does not separate or label speakers, so overlapping talk appears as plain text in a single transcript. If a recording has more than one audio track, only the default track is used, so if your platform saved per-participant audio separately, combine the tracks or upload them one at a time. For recorded meetings, the meeting transcription page covers the workflow, and the video to text tool accepts the video file directly.

Next step: test your call setup

Make a two-minute test call with a colleague, deliberately talk over each other a few times, and listen to the recording. If interjections vanish or word starts are clipped, switch both sides to headphones and test again. Then transcribe the better recording with the audio to text tool; a 45-minute call costs 45 credits (4.5¢).

Frequently asked questions

Why do the first words get cut off on video calls?

Usually the echo canceller's residual suppression is still active from the other person's speech when you start talking, so your first syllable is attenuated. It is most common when one side uses laptop speakers. Headphones on both ends largely solve it.

Does echo cancellation affect call recordings?

Yes, when the recording is made from the call's audio, as cloud and host recordings are. Clipped word starts, missing overlaps and pumping are all captured in the file. A local recording of each person's own microphone avoids the processing, but picks up speaker audio if that person isn't wearing headphones.

Should I turn off echo cancellation?

Only if everyone involved wears headphones and uses a decent microphone. Without echo cancellation, any open loudspeaker sends the other person's voice back as echo or feedback. Some apps offer a processing-off mode for musicians and podcasters; check the current settings of your platform.

Why does a call sound like a walkie-talkie?

That is half-duplex behavior: the echo canceller suppresses one direction heavily while the other side speaks, so both can't be heard at once. It is a common fallback when the echo is strong, the room is reverberant or the speakers distort. Headphones or lower speaker volume usually reduce it.

Can a transcript recover speech that echo cancellation removed?

No. If a word was suppressed during the call, it is missing or nearly silent in the recording, and no recognizer can transcribe what is not there. A separate local recording from the speaker's side, made with headphones, is the only way to keep it.

Is echo cancellation the same as noise suppression?

No, though they usually run in the same chain. Echo cancellation removes sound from your own loudspeaker using a known reference signal. Noise suppression removes unrelated background sounds such as fans or keyboards, which it has to estimate without a reference.