Audio engineering for speech

Voice Activity Detection: How Software Decides When Someone Is Speaking

Voice activity detection, or VAD, is the step that labels each short slice of audio as speech or not speech. Speech recognition systems use it to skip silence, to split long recordings at pauses and, in live systems, to decide when someone has finished talking. Simple detectors measure loudness; modern ones are small neural networks trained to recognize voices. Both can clip the start of a word, miss very quiet speech or get confused by music.

8 min read · Updated

What a voice activity detector does

A VAD looks at audio in short frames, typically a few tens of milliseconds each, and outputs a decision or a probability for every frame: is a human voice present here? A smoothing step then turns those frame decisions into speech regions with start and end times, ignoring tiny blips and bridging very short gaps.

That sounds trivial, because silence looks obvious on a waveform. Real recordings are rarely silent, though. Room tone, fans, traffic, keyboard clicks, breaths, music and other people all fill the gaps between words. The job of a VAD is to separate a voice from everything else that makes sound, at a speed that is cheap enough to run before the expensive recognition model.

Energy-based detectors

The oldest approach compares the loudness of each frame with a threshold. Frames louder than the threshold count as speech; quieter frames count as silence. Refinements include adaptive thresholds that track the background level, zero-crossing rates that help with hissy consonants, and energy measured only in the frequency bands where speech is strongest.

Energy detectors are fast, simple and transparent. Their weakness is that they measure sound, not voice. A door slam or a music bed counts as speech; a soft-spoken talker in a noisy room may not. They work well in quiet, controlled recordings and fail in exactly the conditions where you most want help.

Statistical and neural detectors

Later detectors model what speech and noise look like statistically. The voice detector in the open-source WebRTC project, used in many calling apps, is a well-known example built on classic statistical models of speech and noise spectra.

Neural VADs go further. A small network is trained on large amounts of labeled audio to output a speech probability for each frame, learning the patterns of voicing, formants and syllable rhythm rather than loudness alone. They cope far better with background noise and music, at the cost of running an extra model. They are still small compared with a recognition model, so the cost is modest. Silero VAD is a widely used open example, and faster-whisper's documentation describes its VAD filter as based on it.

Energy-based
Loudness against a threshold. Very fast, easy to tune, confused by any loud non-speech sound
Statistical
Models of speech and noise spectra. Light enough for real-time calls, moderate robustness
Neural
Trained network gives a speech probability. Most robust to noise and music, needs a model and more computation

How speech recognition uses VAD

VAD shows up in several places in a recognition pipeline, and each use has different priorities:

  • Skipping silence. Removing long non-speech stretches saves processing time. It also reduces a known failure of large recognition models: given silence or noise, they sometimes produce plausible words nobody said. That behavior is covered in Whisper hallucinations.
  • Segmenting audio. Long recordings are split into pieces the model can handle, and pauses are the safest places to split. How chunk boundaries are chosen is the subject of audio chunking for speech recognition.
  • Endpointing. Live dictation and voice assistants use VAD to decide when you have stopped talking so they can respond. Too eager and they cut you off mid-thought; too patient and they feel sluggish.
  • Timestamp mapping. When silence is removed, the system must remember where each kept region sat in the original audio so that transcript timestamps and subtitle cues still line up with the recording.

The last point is easy to overlook. A well-built pipeline translates every timestamp back to the original timeline, so removing a minute of silence does not shift every subsequent subtitle by a minute.

The settings that shape VAD behavior

Most detectors expose a handful of parameters, under various names:

  1. Threshold. The speech probability or loudness above which a frame counts as speech. Lower values keep more quiet speech and more noise.
  2. Minimum speech duration. Very short detections, such as a cough or click, are discarded.
  3. Minimum silence duration. A gap must last at least this long before a speech region is closed, so normal pauses between words don't split a sentence.
  4. Padding. Extra audio kept before and after each region, so soft word starts and trailing consonants are not shaved off.
  5. Maximum region length. Very long speech regions may be split for the recognizer's benefit.

Padding deserves special attention. Speech often starts with a quiet sound, such as an h, f, s or a breathy vowel, and ends with a fading consonant. A detector that switches on only when the voice is clearly present will clip those edges unless padding gives them back.

Where VAD fails

  • Clipped onsets and endings. Soft first syllables and trailing sounds fall below the threshold, so the recognizer hears "elp" instead of "help" or loses a plural s.
  • Missed quiet speech. A distant participant, a mumbled aside or a whisper may score below the threshold and be removed before recognition, so the words never appear in the transcript.
  • Music. Singing is voice, so a VAD may pass it through for transcription; speech over a loud music bed may score lower than expected; and instrumental passages can occasionally pass as speech.
  • Non-speech vocal sounds. Laughter, coughs and filled pauses fall into a grey zone; detectors treat them inconsistently.
  • Very noisy recordings. When noise is close to the level of the voice, even neural detectors become uncertain, and the result swings between keeping noise and dropping speech.
  • Overlapping talkers. VAD answers whether any voice is present, not whose. It does not separate speakers, which is a different task called diarization.
Hypothetical: a seminar recording with a quiet student

A lecturer records a 50-minute seminar on a phone at the front of the room. The transcript covers her talk well, and a five-minute silent exercise in the middle produces no stray text. But a student's question from the back row appears only partly, and two short replies are missing entirely. The student's voice reached the phone well below the lecturer's, close to the room's background level, so the detector judged parts of it as non-speech. Raising the gain on the quiet passages in an editor before transcribing brings most of the question back. A clip-on microphone passed to students would have prevented the problem.

Getting better results when VAD is in the pipeline

  1. Record each talker close to a microphone. VAD trouble is mostly a level problem relative to the background.
  2. Reduce steady background noise at the source rather than relying on the detector to ignore it.
  3. Raise very quiet passages in an editor if you know someone was hard to hear. The approaches are covered in transcribing quiet audio.
  4. Keep music beds low under speech, or transcribe from a voice-only version if you have one.
  5. When a transcript is missing a passage, listen to that part of the recording first. If the speech is faint or buried, the detector is a likely cause.

Trade-offs and limits of VAD

Every VAD setting trades one error for another. A sensitive detector keeps quiet speech but passes more noise, which can produce invented text; a strict one keeps transcripts clean but loses soft voices. There is no setting that is right for every recording.

VAD also only tells a system where speech is. It cannot make faint speech intelligible, separate overlapping voices or decide which speaker matters. It is a filter in front of recognition, not a fix for a weak recording.

Voice activity detection in mydubly

In mydubly's default deployment, recognition runs on a server using Whisper large-v3-turbo through faster-whisper, with faster-whisper's voice activity filter enabled. Silence and non-speech stretches are skipped, and the timestamps in the transcript and subtitles still refer to the original audio. Each piece is recognized without conditioning on previously recognized text, which keeps one bad passage from steering the next.

Before that, the browser decodes your file, downmixes it to mono 16 kHz and cuts it into chunks of roughly 30 seconds. Each cut is placed at the quietest 50-millisecond moment within the last 6 seconds of the window, which is a loudness-based choice rather than a neural VAD, so cuts tend to land in gaps between words. The practical consequence is the same as anywhere VAD is used: very quiet or distant speech can be skipped, so raise it before uploading if it matters. mydubly does not label speakers. The output includes a timestamped transcript and SRT and VTT files from the subtitle generator.

Next step: check a missing passage

If a transcript has a gap, find the timestamp around it and listen to the original at that point. Faint speech points to detection; clear speech points to something else. Fix the level, then transcribe again with the audio to text tool, or follow the audio transcription guide for the full workflow. A 50-minute recording costs 50 credits (5¢).

Frequently asked questions

What is VAD in speech recognition?

VAD stands for voice activity detection. It labels short frames of audio as speech or non-speech so the recognizer can skip silence, split long recordings at pauses and keep timestamps aligned. It runs before the main recognition model and is much cheaper than it.

Why does voice activity detection cut off the first word?

Many words begin with soft sounds such as h, f or a breathy vowel that sit close to the background level. If the detector's threshold is strict and its padding short, those first few milliseconds are labeled non-speech and dropped. More padding or a closer microphone usually fixes it.

Is neural VAD better than energy-based VAD?

In noisy recordings and with music, neural detectors are generally more reliable, because they learn what voices sound like rather than measuring loudness. In quiet, controlled audio, a simple energy detector can work well and costs almost nothing to run. The right choice depends on the material and the computing budget.

Can voice activity detection remove speech by mistake?

Yes. Very quiet, distant or whispered speech can score below the threshold and be removed before recognition, so it never appears in the transcript. Speech buried in loud music or noise is also at risk. Raising the level of those passages, or recording closer, reduces the problem.

Does VAD tell who is speaking?

No. It only says whether a voice is present, not whose voice it is. Working out which speaker said what is speaker diarization, a separate task. mydubly does not label speakers in its transcripts or subtitles.

Does removing silence shift transcript timestamps?

Not in a well-built pipeline. The system records where each kept speech region sat in the original audio and maps recognized timestamps back to it. In mydubly's default setup, timestamps refer to the original recording even though silence is skipped.