Speaker diarization answers the question "who spoke when?" for a recording. It does not identify people by name; it labels stretches of speech as SPEAKER_00, SPEAKER_01 and so on, so that the same voice gets the same label throughout the file. Meeting transcripts, call-centre analytics, podcast editing and training-data pipelines all depend on it, and the quality of a speaker-attributed transcript is usually limited by diarization rather than by speech recognition.
The first generation of diarization systems cut audio into short uniform windows, embedded each one and clustered the embeddings. That design assumes exactly one speaker per window, which breaks on the thing real conversations are full of: overlap and rapid turn-taking. The second generation, which this article calls v2, splits the job in two. A neural segmentation model decides locally, inside a few seconds of audio, how many speakers are active in each frame, including overlap. A clustering stage then decides globally which local speakers are the same person. This article explains both halves from first principles, walks through a worked example with real code, and ends with failure modes, operating advice and a checklist. Speech detection is covered in voice activity detection and transcription in ASR architecture.
Why window-and-cluster pipelines hit a ceiling
A v1 pipeline runs voice activity detection, slices the speech into windows of about 1.5 seconds, extracts a speaker embedding from each window and clusters the embeddings. Each cluster becomes a speaker. It is simple and it works well on clean interviews with long turns.
It fails in three predictable ways. A window containing two voices yields a blended embedding that belongs to neither cluster, so overlapped speech is missed or given to one person. A speaker change mid-window snaps to the window grid. A short interjection such as "no, wait" is absorbed into its neighbour. In meetings these cases are common enough to put a floor under the error rate that better embeddings cannot remove.
The v2 fix is to train a model to output frame-level speaker activity directly. It can represent two people talking at once and place a change point at frame resolution. It cannot keep identities consistent across an hour, because it only sees a few seconds; clustering keeps that job.
Measuring it: diarization error rate
The standard metric is diarization error rate (DER): the total duration of false alarm (speech hypothesised where the reference has none), missed speech, and speaker confusion (speech attributed to the wrong speaker), divided by the total duration of reference speech. Because the hypothesis labels are arbitrary, scoring first finds the best one-to-one mapping between hypothesis and reference speakers, then counts confusion under that mapping. Overlapped regions count once per reference speaker, which is why a system that ignores overlap pays for it in missed speech.
Two scoring settings change the number a lot: whether a forgiveness collar around reference boundaries is applied, and whether overlapped speech is scored. Never compare DER figures produced under different settings, and fix the settings in a script when you evaluate on your own data.
The local stage: powerset segmentation
The segmentation model takes a fixed chunk of audio and outputs, for every frame, which of a small number of local speakers are active. In the pyannote segmentation-3.0 model card, a chunk is 10 seconds of mono audio at 16 kHz, a chunk can contain at most 3 speakers, and at most 2 can be active in any one frame.
There are two ways to frame the output. The multi-label way gives each local speaker an independent sigmoid, and you need a tuned threshold per output to decide who is active. The powerset way, from Plaquet and Bredin's 2023 paper on powerset multi-class cross-entropy, enumerates every allowed combination as one class and uses a single softmax. With 3 speakers and at most 2 at once there are 7 classes: non-speech, each speaker alone, and each pair. Taking the arg-max per frame gives a decision with no threshold to tune, and overlap is a first-class label rather than two thresholds that happen to fire together.
Local speaker labels are arbitrary: in one chunk the woman who opens the meeting may be speaker 1, in the next she may be speaker 3. Training therefore uses permutation-invariant training: compute the loss under every assignment of reference speakers to output slots and keep the smallest. With 3 slots there are only 6 permutations, so this is cheap.
import itertools
import torch
import torch.nn.functional as F
# powerset classes for 3 local speakers, at most 2 active per frame
CLASSES = [(), (0,), (1,), (2,), (0, 1), (0, 2), (1, 2)]
def to_powerset(multilabel): # (frames, 3) of 0/1 -> (frames,) class ids
ids = []
for row in multilabel.tolist():
active = tuple(i for i, v in enumerate(row) if v)
ids.append(CLASSES.index(active)) # chunks with 3 overlapping are dropped upstream
return torch.tensor(ids)
def pit_powerset_loss(logits, reference): # logits (frames, 7), reference (frames, 3)
best = None
for perm in itertools.permutations(range(3)):
target = to_powerset(reference[:, list(perm)])
loss = F.cross_entropy(logits, target)
best = loss if best is None or loss < best else best
return best
Embeddings for each local speaker
After segmentation, every chunk has up to three local speakers, each with a frame mask saying where they talk. The pipeline extracts one speaker embedding per local speaker per chunk, using the audio inside that speaker's mask. The community-1 pipeline uses an embedding model from the WeSpeaker toolkit; any model trained for speaker verification plays the same role: it maps a stretch of speech to a vector where the same voice lands close together.
Two details matter in practice. Frames where the speaker overlaps someone else should be excluded from the embedding when enough clean frames remain, because a blended frame pulls the vector towards the other voice. And very short masks give unreliable embeddings: a local speaker active for 0.3 seconds should usually be skipped for clustering and assigned later by its neighbours, rather than allowed to seed a spurious cluster.
The global stage: constrained clustering
Clustering groups the per-chunk embeddings into global speakers. Agglomerative hierarchical clustering (AHC) is the simplest choice: start with every embedding as its own cluster and repeatedly merge the two closest until the closest pair is farther apart than a threshold, or until a known number of speakers remains. The community-1 pipeline instead uses VBx, a Bayesian hidden Markov model over the sequence of embeddings (Landini and colleagues, 2022), which tends to be more robust to threshold choice because it models how speakers persist over time.
The local stage gives clustering a constraint that v1 never had: two different local speakers from the same chunk are, by construction, different people. That is a cannot-link constraint, and enforcing it prevents the most common merge error, two similar voices in the same conversation collapsing into one.
import numpy as np
def constrained_ahc(emb, chunk_ids, threshold=0.7, num_speakers=None):
# emb: (n, d) L2-normalised embeddings; chunk_ids[i]: chunk the embedding came from
clusters = [{i} for i in range(len(emb))]
def dist(a, b): # average cosine distance between clusters
return 1.0 - float(np.mean(emb[list(a)] @ emb[list(b)].T))
def allowed(a, b): # cannot-link: never merge two speakers of one chunk
return not ({chunk_ids[i] for i in a} & {chunk_ids[j] for j in b})
while len(clusters) > 1:
pairs = [(dist(a, b), x, y) for x, a in enumerate(clusters)
for y, b in enumerate(clusters) if x < y and allowed(a, b)]
if not pairs:
break
d, x, y = min(pairs)
if num_speakers is None and d > threshold:
break
if num_speakers is not None and len(clusters) <= num_speakers:
break
clusters[x] |= clusters[y]
del clusters[y]
return clustersThis sketch is slow and shows only the logic; production code builds a linkage once with a library and sets forbidden pairs to an infinite distance.
Reconstruction: back to one timeline
Chunks overlap, so most frames are covered by several chunks. Reconstruction maps every local speaker to its global cluster and averages each global speaker's activation per frame over the covering chunks. It also averages the per-chunk speaker counts to estimate how many people speak in each frame, and keeps that many top global speakers. Short gaps are then filled and very short turns removed.
The output is a list of turns: start, end and speaker label, often written as RTTM. Turns from different speakers may overlap in time, which is correct for conversation but awkward for transcripts. pyannote's community-1 output therefore also exposes an exclusive diarization, in which at most one speaker is active at a time, specifically to simplify merging with transcription timestamps.
Worked example: a four-person meeting
Take a one-hour recording of four people. With 10 second chunks and a 1 second step, the segmentation model runs about 3,600 times; batched on a GPU this is the cheapest part. Each chunk yields up to 3 local speakers, so clustering sees up to 10,800 embeddings, fewer after short masks are dropped. If you know there are four participants, pass that; if you only know a range, pass bounds; if you know nothing, the clustering threshold decides.
from pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained(
"pyannote/speaker-diarization-community-1", token="HF_TOKEN")
output = pipeline("meeting.wav", min_speakers=3, max_speakers=5)
for turn, speaker in output.speaker_diarization:
print(f"{speaker} {turn.start:8.2f} {turn.end:8.2f}")
exclusive = list(output.exclusive_speaker_diarization) # one speaker at a timeThe model card states the pipeline expects 16 kHz mono and resamples or downmixes other input itself. The token is needed because the model is gated on Hugging Face. Check results by listening: play every boundary in two or three five-minute stretches.
Attaching speakers to words
Most users want a speaker-attributed transcript, which means joining diarization turns with ASR words. The robust rule is to give each word to the speaker whose exclusive turn overlaps it most, and to fall back to the nearest turn when a word falls in a gap. Word timestamps from ASR are themselves approximate, so do not split a word across speakers.
def assign(words, turns):
# words: [(start, end, text)], turns: [(start, end, speaker)] from exclusive output
out = []
for ws, we, w in words:
best, best_ov = None, 0.0
for ts, te, spk in turns:
ov = min(we, te) - max(ws, ts)
if ov > best_ov:
best, best_ov = spk, ov
if best is None: # in a gap: nearest turn by midpoint
mid = (ws + we) / 2
best = min(turns, key=lambda t: min(abs(mid - t[0]), abs(mid - t[1])))[2]
out.append((ws, we, w, best))
return outThen merge consecutive same-speaker words into utterances. If your ASR emits only segments, force-align them to words first.
Streaming diarization is a different design
Everything above is offline: clustering sees the whole file before deciding. Live captioning cannot wait. Online variants keep running speaker centroids and assign each new chunk's local speakers to an existing centroid or a new one. The cost is label stability: when two centroids later prove to be one person, the system must live with the split or relabel text already shown. End-to-end neural diarization (EEND), which outputs global speaker activity directly, is another research line aimed at this setting. For low-latency speech pipelines see streaming ASR architecture.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Two people become one speaker | Similar voices, threshold too loose, or speaker count forced too low | Pass min_speakers, tighten threshold, check cannot-link is applied |
| One person split into two | Channel change (phone vs room mic), long recording drift | Loosen threshold or pass max_speakers; normalise audio levels |
| Phantom speaker with tiny total time | Laughter, music, noise embedded as speech | Drop clusters below a minimum total duration and reassign their turns |
| Overlap missed entirely | Using exclusive output for scoring, or a v1 pipeline | Score the overlap-aware output; keep exclusive only for transcripts |
| Good DER, bad transcript | Word assignment by segment, or ASR timestamps drift | Assign per word by overlap; forced-align segments |
| Slow on long files | Embedding extraction on CPU, unbatched | Run on GPU, batch chunks, split very long files at long silences |
Operating it
Build an evaluation set of twenty to forty minutes of your own audio, because benchmark meetings say little about phone calls with hold music. Record the pipeline version and hyperparameters with every output, and when the speaker count is unknown tune the clustering threshold on this set; it moves results most.
Treat speaker embeddings as biometric data: store them only if needed, scoped to the recording, and delete them with it. Labels are per file; SPEAKER_00 in one call is unrelated to SPEAKER_00 in the next, and cross-file identity is a separate verification problem with its own consent questions.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Speaker count | Pass the known count: removes the hardest decision, wrong if someone joins | Pass bounds or nothing: adapts, relies on threshold quality |
| Clustering | AHC: simple, transparent, threshold-sensitive | VBx: more robust over time, more parameters to understand |
| Output | Overlap-aware: correct for analytics and scoring | Exclusive: clean for transcripts, drops overlap by design |
| Mode | Offline: best accuracy, waits for the whole file | Online: low latency, unstable labels |
What to do next
- Collect 20-40 minutes of your own audio and annotate reference turns, including overlap.
- Write a scoring script with fixed collar and overlap settings and keep it under version control.
- Run the community-1 pipeline with known speaker counts first, then with bounds, and compare DER on your set.
- Listen to every boundary in two short stretches to see which failure mode dominates before tuning anything.
- Merge with ASR per word using the exclusive output, and spot-check the attributed transcript.
- Decide retention for embeddings and outputs, and record pipeline version and parameters with every result.