Speech recognition & transcription

Speaker Diarization: How Software Works Out Who Spoke When

Speaker diarization is the process of splitting an audio recording into segments according to who is speaking, answering the question 'who spoke when'. It does not identify people by name; it groups stretches of speech into anonymous speakers such as Speaker 1 and Speaker 2. Most systems detect speech, turn short slices of it into voice embeddings, and cluster those embeddings, which works well on clean turn-taking and poorly on overlap and similar-sounding voices.

6 min read · Updated

What diarization produces, and what it does not

A diarization system outputs a list of time ranges, each tagged with a speaker label: Speaker A from 0.0 to 12.4 seconds, Speaker B from 12.6 to 30.1, and so on. Paired with a transcript, those ranges become the familiar "Speaker 1:" prefixes in an interview transcript.

Diarization is not speaker identification. It does not know that Speaker A is the host; it only knows that two stretches of audio sound like the same voice. Attaching real names requires either a human or a separate speaker recognition system enrolled with known voice samples. Diarization is also separate from speech recognition: a speech model like Whisper transcribes words but does not, on its own, say who said them.

The classic diarization pipeline

Most production systems still follow a modular pipeline, even when each module is a neural network.

  1. Voice activity detection marks where anyone is speaking and discards silence, music and noise.
  2. Segmentation splits speech into short pieces, either at detected speaker changes or into uniform windows of a second or two.
  3. Embedding extraction turns each piece into a fixed-length vector that captures what the voice sounds like.
  4. Clustering groups the vectors so that pieces from the same voice end up together, and each cluster becomes a speaker.
  5. Resegmentation refines boundaries, smoothing out implausibly short turns and reassigning edges.

Speaker embeddings and clustering

The embedding step does the heavy lifting. A network trained on thousands of speakers learns to map any short clip to a point in a vector space where clips of the same voice lie close together and different voices lie far apart. The field moved from statistical i-vectors to neural x-vectors and to architectures such as ECAPA-TDNN, but the goal is the same: a compact voice fingerprint that ignores what is being said.

Clustering then has to decide how many speakers there are. Agglomerative hierarchical clustering repeatedly merges the closest groups until the remaining groups are further apart than a threshold; spectral clustering analyzes the similarity matrix as a whole. If you know the number of speakers in advance, many toolkits let you supply it, which removes one of the biggest sources of error. Left to guess, a system may split one speaker into two when their voice changes, or merge two similar voices into one.

Overlapping speech and end-to-end models

The modular pipeline assumes one speaker at a time, so overlapping speech is its weak point. When two people talk at once, the embedding of that slice is a blend of both voices and belongs cleanly to neither cluster.

End-to-end neural diarization, often abbreviated EEND, attacks this directly. Instead of clustering, a network outputs, for every frame, the probability that each speaker is active, so two speakers can be active simultaneously. Pure end-to-end models struggle with long recordings and many speakers, so popular open-source toolkits such as pyannote.audio combine a neural segmentation model that handles overlap within short windows with embedding-based clustering across the whole file. Check the toolkit documentation for current model versions, as these change frequently.

How diarization quality is measured

The standard metric is diarization error rate, or DER: the fraction of reference speech time that is attributed incorrectly. It has three components.

False alarm
Time labeled as speech when nobody was speaking.
Missed speech
Time when someone spoke but the system labeled no speaker.
Speaker confusion
Time assigned to the wrong speaker.
Worked example: computing DER

Suppose a 10-minute recording contains 600 seconds of reference speech. The system misses 30 seconds, adds 15 seconds of false speech and assigns 45 seconds to the wrong speaker. DER = (30 + 15 + 45) / 600 = 0.15, or 15 percent. Evaluations often ignore a short collar around each speaker change, commonly a quarter of a second, so scores from different papers are not always comparable.

Combining diarization with a transcript

To get a labeled transcript, the speaker time ranges must be merged with the words. Each word or segment from the speech recognizer is assigned to whichever speaker was active at its timestamp. This is where small timing differences matter: if the transcript's segments span a speaker change, the whole segment can land on one speaker. Open-source projects such as WhisperX pair Whisper with word-level alignment and pyannote diarization for this reason, and accurate word timestamps make the merge far cleaner.

Where diarization helps

  • Interview and focus group transcripts, where readers need to know who said what.
  • Meeting minutes and action items attributed to the right person.
  • Searching an archive for everything one speaker said.
  • Talk-time analysis, such as how much an interviewer spoke versus the guest.
  • Multi-voice dubbing, where each original speaker needs a distinct synthetic voice.

Common error modes and limitations

  • Overlapping speech and quick back-and-forth exchanges cause confusion and missed speech.
  • Similar voices, such as siblings or two speakers of the same age and accent, may be merged into one.
  • One speaker can be split into several when they move away from the microphone, laugh, shout or change tone.
  • Short interjections like "yes" or "right" are often absorbed into the surrounding speaker's turn.
  • Phone audio, heavy compression and music beds degrade embeddings.
  • Labels are anonymous, so a human still has to map Speaker 1 and Speaker 2 to real people.

mydubly transcripts have no speaker labels: practical workarounds

mydubly's transcription runs on Whisper and produces a timestamped transcript and SRT and VTT subtitles, but it does not perform diarization: transcripts do not label speakers. If you need attribution, these approaches work with what mydubly does provide.

  1. Record each speaker on a separate track when you can, as many podcast and interview setups already do, export each track as its own audio file, transcribe each one separately, and merge the two transcripts by timestamp.
  2. For a single mixed track, add speaker names by hand during your proofreading listen, using the timestamps to jump between turns.
  3. For subtitles, use dashes or name prefixes only where a speaker change would otherwise confuse viewers.

Separate tracks have one caveat: an isolated track contains long silences while the other person talks, and speech models can occasionally invent text over silence, so scan those transcripts for stray phrases, as explained in Whisper hallucinations. For dubbing, note that mydubly uses one chosen voice for the whole video; the trade-offs are covered in single-voice vs multi-voice dubbing.

Example: a two-host podcast

Suppose a 50-minute episode was recorded with each host on their own microphone track. Exporting both tracks and transcribing each costs 50 credits per track, 100 credits or 10 cents in total. Interleaving the two timestamped transcripts gives a speaker-attributed transcript without any diarization model, and with no overlap confusion.

Next step

If your recordings have one main speaker or you can work from separate tracks, upload them to audio to text for timestamped transcripts. The podcast transcription and interview transcription pages show how those workflows fit together.

Frequently asked questions

Is speaker diarization the same as speaker recognition?

No. Diarization groups speech by voice into anonymous speakers within one recording. Speaker recognition matches a voice against known, enrolled voices to say who a person is.

Does Whisper do speaker diarization?

Not by itself. Whisper transcribes speech and produces timestamps, but it does not output speaker labels. Diarization is usually done by a separate model and merged with Whisper's output using timestamps.

Why does diarization sometimes split one person into two speakers?

The system clusters voice embeddings, and a speaker's voice can change enough to form a second cluster when they move away from the microphone, laugh, shout or speak over different background noise. Telling the system the number of speakers, where supported, reduces this.

Can I get a speaker-labeled transcript from mydubly?

Not automatically, because mydubly transcripts do not include speaker labels. If each speaker was recorded on a separate track, transcribe each track and merge them by timestamp; otherwise add names by hand while proofreading.

What is a good diarization error rate?

It depends heavily on the recording. Clean two-person interviews can score far better than noisy meetings with many speakers and frequent overlap, and differences in evaluation settings, such as the collar, make published numbers hard to compare.