Speaker diarization answers the question of who spoke when. It does not know names. Its output is a set of time segments, each labelled with an anonymous tag such as SPEAKER_00, so that a transcript can be split into turns, a meeting summarised per participant, or a support call separated into agent and customer. The neural model at the heart of a modern diarizer is covered in the companion article on overlap-aware neural diarization. This article is about the system around that model.
Most production trouble comes not from the model but from eight-hour recordings that do not fit in memory, users who want names instead of SPEAKER_03, the same person appearing in fifty calls under fifty tags, evaluation numbers that cannot be compared, and voiceprints being biometric data. We take each in turn and finish with a checklist.
What a diarizer produces, from first principles
Nearly every current diarization system does four things. Voice activity detection finds the regions that contain speech. A segmentation model decides, frame by frame within a short window, which local speakers are active, including overlapping ones. An embedding model turns each local speaker's audio into a fixed-length vector, typically a few hundred numbers, so that two vectors from the same voice are close under cosine similarity and vectors from different voices are far apart. Clustering then groups those vectors into speakers for the whole file, and a reconstruction step writes the final timeline.
Two properties drive the rest of the article. Labels are local to one run: nothing connects SPEAKER_00 today with SPEAKER_00 tomorrow. And clustering cost grows with the number of embeddings, which grows with audio length. Long audio and identity across recordings are system problems, not model problems.
Results are exchanged as RTTM, plain text with one ten-field line per segment: the type SPEAKER, file id, channel, start and duration in seconds, two unused fields, the speaker label and two more unused fields. Keep references and system output in RTTM and every scoring tool will read them.
SPEAKER board_2026_09 1 12.480 3.920 <NA> <NA> SPEAKER_01 <NA> <NA>
SPEAKER board_2026_09 1 16.120 1.050 <NA> <NA> SPEAKER_00 <NA> <NA>
The service architecture
Treat diarization as an asynchronous job, not a request: a one-hour file takes minutes of compute, clients disconnect, and workers die. Each of the six stages below writes its output to object storage keyed by job id and stage, so a failed stage is retried alone, and records the model versions and parameters it used.
- Ingest: decode and resample once, usually to 16 kHz mono. If each participant has their own channel, as in many call-centre recordings, channel identity is speaker identity and per-channel VAD may replace diarization entirely.
- Voice activity detection over the whole file, which also yields the speech ratio, the first health metric to log.
- Block diarization of fixed-length blocks on GPU workers.
- Global linking of block-local speakers into speakers for the whole file.
- Optional identity resolution against enrolled people or earlier recordings.
- Transcript join and publish.
Long audio: why you split, and how to stitch
A pipeline that handles a 20-minute meeting can fail on an eight-hour deposition, and the reason is arithmetic. With a 10-second window advanced one second at a time and up to three local speakers per window, eight hours is 28,800 windows and up to 86,400 embeddings. A full pairwise similarity matrix over them in 32-bit floats needs 86,400 squared times 4 bytes, about 30 GB. Skipping silent windows helps, but growth stays quadratic, and a crash at hour seven wastes everything before it.
The fix is to diarize fixed blocks independently, for example 30 minutes with 60 seconds of overlap, then link speakers between blocks. Each block speaker is summarised by a centroid, the unit-normalised mean of its embeddings. Linking is an assignment problem: compare each block's centroids with the file speakers so far, accept the best one-to-one matches above a threshold using the Hungarian algorithm, and create new file speakers for the rest.
import numpy as np
from scipy.optimize import linear_sum_assignment
def unit(v):
return v / np.linalg.norm(v)
def link_block(global_spk, block_spk, threshold=0.6):
# global_spk: id -> (unit centroid, seconds). block_spk: local id -> (unit centroid, seconds).
# Returns local -> global mapping and updates global_spk in place.
g_ids, b_ids = list(global_spk), list(block_spk)
mapping = {}
if g_ids and b_ids:
G = np.stack([global_spk[g][0] for g in g_ids])
B = np.stack([block_spk[b][0] for b in b_ids])
sim = B @ G.T # cosine similarity: vectors are unit length
rows, cols = linear_sum_assignment(-sim) # best one-to-one assignment
for r, k in zip(rows, cols):
if sim[r, k] >= threshold:
mapping[b_ids[r]] = g_ids[k]
for b in b_ids:
cen, secs = block_spk[b]
if b in mapping: # duration-weighted centroid update
g = mapping[b]
gcen, gsecs = global_spk[g]
global_spk[g] = (unit(gcen * gsecs + cen * secs), gsecs + secs)
else: # nobody close enough: a new file speaker
g = f"SPK_{len(global_spk):02d}"
global_spk[g] = (cen, secs)
mapping[b] = g
return mappingTwo refinements matter. Use the overlap as evidence: the 60 seconds shared by two blocks was diarized twice, so the block speakers sharing the most overlapping speech are almost certainly one person, and that vote should override a borderline score. And distrust tiny speakers: under ten seconds of speech gives a noisy centroid, so link it with a stricter threshold or assign it last. The 0.6 threshold is a placeholder to tune on your audio.
Longer blocks give clustering more context, so a quiet participant is more likely to stay one person, but cost more memory and costlier retries. Shorter blocks parallelise and bound memory but create more link decisions, each a chance to merge two people or split one.
Worked example: a three-hour board meeting
A board meeting on one room microphone runs 3 hours 10 minutes with nine participants. The job splits it into seven 30-minute blocks with 60-second overlaps on four GPU workers; each block reports three to six local speakers.
Block 1 creates five file speakers. In block 2, five of six match with similarities of 0.71 to 0.84, and the sixth, a secretary silent in the first half hour, scores 0.41 against everyone and becomes SPK_05. By block 7 there are eleven file speakers for nine people. SPK_09 holds 14 seconds and SPK_10 holds 22, both just under threshold against SPK_02: one director who moved closer to the microphone after a break. A final merge pass, which merges above-threshold pairs when the smaller holds under a minute of speech, folds both into SPK_02.
The lesson generalises: conditions drift in long sessions, and small file speakers close to an existing one are usually splits. Log file speakers per hour and the share of speech held by speakers under one minute; a rise in either is the earliest sign of trouble.
From tags to names: enrollment and identification
Users want names, not SPEAKER_03. Speaker identification supplies them for enrolled people: enrollment stores the normalised mean embedding of a consenting person's sample as a voiceprint, and identification compares each file speaker's centroid with the voiceprints.
This is an open-set problem: most speakers may not be enrolled, so the system must be able to answer no one. Use a threshold on the best score plus a margin over the second best, chosen from data. Collect target trials (people against their own voiceprint) and impostor trials (against others') on audio like yours, plot both distributions, and pick the threshold for a low false acceptance rate. A wrong name is far worse than none, so accept that some enrolled people stay unnamed.
def identify(centroid, speech_s, voiceprints, threshold, margin=0.05, min_speech_s=20.0):
# Open-set identification. Returns an enrolled person id, or None for "unknown".
# centroid and voiceprints are unit vectors from the SAME embedding model version.
if speech_s < min_speech_s or not voiceprints:
return None # too little audio to name anyone
scores = sorted(((float(centroid @ vp), pid) for pid, vp in voiceprints.items()), reverse=True)
best, pid = scores[0]
second = scores[1][0] if len(scores) > 1 else -1.0
if best >= threshold and best - second >= margin:
return pid
return None # wrong names cost more than no namesRaw scores shift with recording conditions, so a threshold tuned on headsets misbehaves on speakerphones. Adaptive symmetric score normalisation (AS-norm), which rescales each score against a cohort of other speakers, often lets one threshold hold across conditions; measure false acceptance per channel type with and without it. Enroll on the devices people actually use, require a minimum of clean speech, and keep the duration floor at identification time.
Linking the same person across recordings
Without enrollment you can still find that Monday's anonymous speaker is Thursday's. Store each file speaker's centroid and duration in a per-tenant speaker library and match new speakers with the same open-set rule, so a user labels a person once and the name propagates.
This is harder than linking blocks: microphones, rooms and codecs differ, voices change, and every library entry adds a chance of a false match. Treat human-confirmed labels as anchors that automatic matches never override, suggest matches for confirmation until measured false acceptance is low enough, and keep libraries scoped, per tenant always and per team where possible.
Joining speakers to the transcript
The transcript comes from a separate ASR pass, in parallel, and the join gives each word to the speaker overlapping it most (the companion article covers the rules). Both passes must read the same normalised file, and the join should be its own stage so it can be rerun when either side improves. Live captions need an incremental design around streaming ASR, where labels can change after display.
An evaluation harness you can trust
Diarization error rate (DER) is missed speech plus false alarm plus speaker confusion, divided by total reference speech. The number means nothing without its settings, so the harness fixes them in code: whether a forgiveness collar around reference boundaries is applied (older evaluations commonly forgave 0.25 seconds each side, which is collar=0.5 in pyannote because its collar is the total width; many recent benchmarks use none) and whether overlapped speech is scored. pyannote.metrics implements the standard computation and reports the components separately.
from pyannote.core import Annotation
from pyannote.database.util import load_rttm
from pyannote.metrics.diarization import DiarizationErrorRate
COLLAR, SKIP_OVERLAP = 0.0, False # fixed settings, printed with every report
ref = load_rttm("eval/reference.rttm") # dict: file id -> Annotation
hyp = load_rttm("runs/2026-10-01/system.rttm")
metric = DiarizationErrorRate(collar=COLLAR, skip_overlap=SKIP_OVERLAP)
for uri, reference in sorted(ref.items()):
d = metric(reference, hyp.get(uri, Annotation(uri=uri)), detailed=True)
t = d["total"]
print(f"{uri}: DER={d['diarization error rate']:.3f} miss={d['missed detection']/t:.3f} "
f"fa={d['false alarm']/t:.3f} conf={d['confusion']/t:.3f}")
print(f"overall DER={abs(metric):.3f} collar={COLLAR} skip_overlap={SKIP_OVERLAP}")Read the components: missed speech points at VAD or overlap handling, false alarm at noise passing VAD, confusion at embeddings, clustering or linking. Read per-file numbers, since a mean can hide one recording type at three times the rest. Also compare file speaker counts with the reference, because a split speaker holding a few seconds barely moves DER but ruins per-speaker summaries, and report correct names, wrong names and unknowns separately for identification.
Build references from your own audio, twenty to forty minutes per recording type including overlap, version them, never tune on the files you report, and rerun the harness for every model, threshold or block change.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Too many speakers on long files | Conditions drift and one person splits across blocks | Overlap votes, final merge pass for small speakers, alert on speakers per hour |
| Two people merged | Similar voices, low threshold, noisy centroids from short block speakers | Stricter threshold for small speakers; pass speaker-count bounds when known |
| Wrong name in a transcript | Threshold tuned on different audio, no margin rule | Per-channel thresholds, AS-norm, margin and duration floor; default to unknown |
| Words shifted by a fixed offset | ASR and diarization read differently trimmed audio | Normalise once; both read the same file |
| Identification collapses after an upgrade | Embedding model changed; stored voiceprints come from the old model | Version voiceprints by model; re-embed from consented enrollment audio |
Voiceprints are biometric data
An embedding that can recognise a person is a biometric identifier. Under the EU GDPR, biometric data processed to uniquely identify a person is a special category under Article 9 and needs an explicit legal basis such as explicit consent. The Illinois Biometric Information Privacy Act names voiceprints and requires written consent and a published retention schedule, and other jurisdictions have similar rules. This is not legal advice; involve counsel. The engineering consequences are clear, though.
- Keep anonymous diarization, which stores no embeddings, separate from identification and linking, which are opt-in per tenant and per person.
- Store consent with each voiceprint and delete it when consent is withdrawn; delete per-file embeddings with the recording.
- Never match across tenants, and protect the library like a credential store.
- Log every identification with its score so a disputed attribution can be audited.
Operating the pipeline
Monitor per job: speech ratio, file speakers per hour, link similarities, unknown rate and real-time factor, alerting on distribution shifts. An embedding model upgrade invalidates every stored voiceprint and library centroid, so keep consented enrollment audio and re-embed before switching.
What to do next
- Check whether you need diarization at all: if each participant has a channel, use the channel.
- Build a job pipeline that stores every stage's output and records model versions and parameters per stage.
- Count embeddings on your longest file, then choose a block length and overlap that bound worker memory.
- Implement block linking with overlap votes and a final merge pass, and log file speakers per hour.
- Build the harness: RTTM references from your own audio, fixed collar and overlap settings, per-file components.
- If you need names, collect target and impostor trials, set a low-false-acceptance threshold with a margin, and default to unknown.
- Design consent, retention and deletion for voiceprints before the first one is stored.