A wake word detector listens to every second of audio in a room so that the rest of a voice system can sleep. It must react to its phrase in a fraction of a second, in noise, from across the room, from many voices and accents, while almost never firing on anything else, and it must do this within the power budget of a chip that runs all day. Those requirements push wake word detection toward an architecture of its own: a cascade of small streaming models with careful decision logic, rather than a speech recognizer looking for a word.
This article builds that architecture from the metrics upward, with working code for the front end, the decision logic and the evaluation sweep. It assumes the neighbouring components exist; voice activity detection and streaming ASR have their own articles.
The two numbers that define the product
A detector makes two kinds of mistake. A false reject is a missed wake: the user said the phrase and nothing happened. It is measured as a rate over a set of recorded positive utterances. A false accept is a wake with no wake word: the device lights up during a conversation or a TV show. Because negatives are not countable utterances but continuous audio, false accepts are measured per hour of audio, usually written FA/h.
The two trade against each other through the decision threshold, and the scale of the negative side is what makes the problem hard. A device in a busy home may hear many hours of speech and media a day. At 1 FA/h that is several unwanted wakes daily per device, which users experience as both annoying and a privacy breach, since a false accept usually means audio is sent somewhere. Products therefore aim for false-accept rates far below one per hour and accept a few percent of false rejects to get there. Pick your targets explicitly before building anything, because every later decision is a move along that curve.
The cascade architecture
No single model meets both a microwatt-class power budget and a very low FA/h. The standard answer is a cascade. A tiny first-stage model runs continuously on a DSP or microcontroller, tuned for high recall: it fires on the wake word almost every time and on a modest amount of other audio. Only when it fires does the device wake its application processor, which runs a larger second-stage model over the buffered audio. Some systems add a third check on a server, re-scoring the clip before a session continues.
Three supporting pieces make the cascade work. Echo cancellation subtracts what the device itself is playing (see acoustic echo cancellation), so its own TTS or music does not wake it. A ring buffer retains the last second or two of audio, so the verifier can re-examine the phrase and the recognizer receives the start of the command. And the decision logic, not the model, decides when a score becomes a wake.
The streaming front end
Most keyword models consume log-mel energies: 16 kHz audio cut into 25 ms windows every 10 ms, a magnitude spectrum, a bank of 40 or so triangular mel filters, and a logarithm. The front end must be streaming: it receives 10 ms hops of audio and emits one feature frame per hop, carrying the window overlap as state. Noise suppression may run before it (see real-time noise suppression), but train and test with the same chain, because the model learns the front end's artefacts.
import numpy as np
class LogMelStream:
def __init__(self, sr=16000, win=400, hop=160, n_fft=512, n_mels=40):
self.win, self.hop, self.n_fft = win, hop, n_fft
self.window = np.hanning(win).astype(np.float32)
self.buf = np.zeros(win, np.float32)
# any standard mel filterbank, e.g. librosa.filters.mel(sr=sr, n_fft=n_fft, n_mels=n_mels)
self.fb = mel_filterbank(sr, n_fft, n_mels) # (n_mels, n_fft//2 + 1)
def push(self, hop_samples):
"""Feed exactly `hop` new samples; return one log-mel frame."""
self.buf = np.concatenate([self.buf[self.hop:], hop_samples])
spec = np.abs(np.fft.rfft(self.buf * self.window, self.n_fft)) ** 2
return np.log(self.fb @ spec + 1e-6)On a microcontroller the same pipeline runs in fixed point with a precomputed filterbank, and the model typically sees a sliding window of the last one to one and a half seconds of frames, enough to cover the phrase.
Models small enough to run forever
Chen, Parada and Heigold (ICASSP 2014) showed a small feed-forward network over stacked frames, predicting keyword sub-units against a filler class, could beat a keyword-filler HMM baseline. Later work moved to convolutional models; the Hello Edge study (Zhang et al., 2017) compared DNN, CNN, recurrent and depthwise-separable CNN (DS-CNN) architectures under microcontroller memory and compute limits and found DS-CNN gave the best accuracy for the budget. Public benchmarks commonly use the Google Speech Commands dataset (Warden, 2018), though a product wake word needs its own data.
Two engineering points matter more than the architecture choice. The model must be streaming: causal convolutions or recurrent state let it produce a score every frame without recomputing the whole window. And it must be quantized, usually to 8-bit integers, to fit the memory and power of the always-on core; quantize with calibration data that includes far-field and noisy audio, not only clean positives.
Training data dominates results. Positives need many speakers, accents, distances, rooms and speaking styles. Negatives need hundreds or thousands of hours of general speech, television, music and household noise, plus hard negatives: words that sound like the wake phrase. Augmentation multiplies both: convolve with room impulse responses, mix noise at a range of signal-to-noise ratios, vary speed and gain. Because negatives vastly outnumber positives at run time, weight or sample the classes so the model does not learn to say no to everything.
Decision logic: from frame scores to a wake
A per-frame posterior is jittery; thresholding it directly gives both misses and double-fires. Chen et al. smoothed posteriors over a short window and took the maximum of the smoothed score over a longer window as the confidence. Production logic adds a minimum number of frames above threshold, a refractory period after each wake, and suppression while the device is in a session.
from collections import deque
class WakeDecider:
def __init__(self, threshold=0.85, smooth=30, confirm=5, lockout=150):
self.th, self.confirm, self.lockout = threshold, confirm, lockout
self.hist = deque(maxlen=smooth) # 30 frames = 300 ms at a 10 ms hop
self.above, self.cooldown = 0, 0
def step(self, p_keyword, suppressed=False):
self.hist.append(p_keyword)
if self.cooldown:
self.cooldown -= 1
return False
score = sum(self.hist) / len(self.hist)
self.above = self.above + 1 if score >= self.th else 0
if self.above >= self.confirm and not suppressed:
self.above, self.cooldown = 0, self.lockout # 1.5 s refractory
self.hist.clear()
return True
return FalseThe detector should also report where the phrase started, from the frame index where the smoothed score crossed a lower threshold. The session uses that to cut the wake word out of the audio sent to the recognizer, or to align it, and the verifier uses it to know which slice of the ring buffer to re-score.
Self-wake, media and broadcast triggers
The device's own output is the worst negative source: its speaker is centimetres from its microphones and it may literally say the wake word, for example when reading a message aloud. Echo cancellation with the playback signal as reference removes most of it; on top of that, many products raise the threshold or suppress detection when their own TTS is speaking the phrase, and duck playback volume as soon as stage one fires so the verifier hears the user more clearly.
Television and advertising are the other systematic source: one broadcast containing the phrase triggers many devices at once. Generic defences are a stricter verifier and server-side checks that notice many near-identical triggers arriving at the same moment. Treat any such server check as a privacy-sensitive design that needs its own review.
Evaluating with threshold sweeps
Evaluation needs two sets: positive clips (with labels), and long continuous negative recordings that represent real homes, measured in hours. Run the full streaming chain, including the decision logic, over both. Each threshold yields one point: false reject rate on positives and wakes per hour on negatives. Plot the curve and pick the operating point from your targets.
def sweep(scores_pos, neg_streams, neg_hours, thresholds, decider_cls=WakeDecider):
rows = []
for th in thresholds:
missed = 0
for clip in scores_pos:
d = decider_cls(threshold=th) # one decider per clip, so state carries
missed += not any(d.step(p) for p in clip)
fas = 0
for stream in neg_streams:
d = decider_cls(threshold=th)
fas += sum(d.step(p) for p in stream)
rows.append((th, missed / len(scores_pos), fas / neg_hours))
return rows # (threshold, FRR, FA/h)Note that any() stops at the first wake, which is what you want for positives. Two common evaluation mistakes: positives recorded close to the microphone in quiet rooms, which flatter the false reject rate, and negative sets too small to measure the target. To estimate 0.1 FA/h with any confidence you need many tens of hours of negative audio at minimum, and far more to distinguish 0.05 from 0.1.
Worked example: budgeting a two-stage cascade
The following operating points are illustrative, not measurements of any product. Stage one runs at 5 FA/h with a 1% false reject rate. Stage two, run only on stage-one triggers, rejects 98% of those false accepts and passes 97% of true wakes. The cascade's false accept rate is 5 × 0.02 = 0.1 FA/h, and its recall is 0.99 × 0.97 = 0.96, a 4% false reject rate.
The power budget follows from the same numbers: the application processor wakes about five times an hour for false triggers plus once per real use, so the expensive model runs for seconds per hour instead of continuously. If stage one drifts to 20 FA/h, perhaps because a new device has a noisier microphone, the system FA/h quadruples to 0.4 and battery life drops, even though stage two is unchanged. Monitor stage-one trigger rates in the field for exactly this reason.
Operating it in the field
- Log decisions, scores and stage outcomes as metadata, not audio, unless the user has opted in; the detector exists to avoid recording.
- Track stage-one trigger rate, stage-two rejection rate and session abandonment per hardware model and firmware version; a rising rejection rate is a false-accept leak upstream.
- Ship model and threshold updates over the air behind staged rollouts, and keep the previous model to roll back.
- Re-run the full sweep whenever the microphone, enclosure, echo canceller or front end changes; the model learned the old chain.
- Consider optional speaker verification for personal devices, with the enrolment embedding kept on the device.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Device wakes on its own speech | Echo reference misaligned or missing | Fix AEC delay alignment; suppress during own TTS |
| Misses from across the room | Training positives too close-talk | Far-field recordings, room impulse augmentation |
| Double wakes on one utterance | No refractory period | Lockout after each wake, clear smoothing history |
| False accepts on a similar word | Missing hard negatives | Collect confusable words, retrain, raise stage-two threshold |
| Command start cut off | Ring buffer too short or wrong start index | Longer pre-roll, report keyword start time |
| Field FA/h far above lab | Negative set unrepresentative | Test on long recordings of real homes and media |
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Threshold | Lower: fewer misses, more false wakes | Higher: fewer false wakes, more frustrated users |
| Stage-one size | Tiny: lowest power, more triggers to verify | Larger: fewer triggers, more always-on power |
| Server verification | Lowest FA/h | Latency, network dependence, privacy cost |
| Custom phrases | User choice | Little phrase-specific training data, weaker accuracy |
What to do next
- Set FA/h and false reject targets and write them down.
- Build a negative test set of many hours of real-home audio and a far-field positive set.
- Implement the streaming front end and decision logic above, then sweep thresholds over both sets.
- Add a second stage if one model cannot reach the FA/h target within your power budget.
- Wire the playback reference into echo cancellation and test self-wake with the device's own TTS.
- Ship with field telemetry on stage trigger rates, and re-sweep after every hardware or front-end change.