Automatic speech recognition turns a waveform into text. Modern systems do it with one neural network trained end to end, which makes ASR look simple from the outside: audio in, words out. Inside, the architecture still has distinct stages with distinct failure modes, and most production problems come from misunderstanding one of them: a sample-rate mismatch in the front-end, a decoder family that hallucinates on silence, a word error rate that looks worse than it is because of punctuation.
This article explains the architecture from the signal up: how audio becomes features, what the encoder does, how the three decoder families differ, how decoding and language models fit in, how to measure accuracy honestly and how to run long recordings in batch. Low-latency live captioning adds its own constraints and is covered separately in streaming ASR architecture; here the focus is the model and the offline system.
The shape of an ASR system
Every current end-to-end recogniser has the same skeleton. A front-end converts audio into a sequence of spectral frames. An encoder, a deep network, turns those frames into a shorter sequence of contextual vectors that represent what was said around each moment. A decoder maps encoder vectors to text tokens, and a search procedure picks the most likely token sequence, optionally helped by a language model. Finally, text normalization turns tokens into the written form users expect.
The tokens are usually subword units, such as a vocabulary of a few hundred to a few thousand SentencePiece or byte-pair pieces, rather than characters or whole words. Subwords keep sequences short and let the model spell words it never saw in training.
Front-end: from waveform to log-mel frames
Speech models almost always consume log-mel spectrograms. The audio is resampled to 16 kHz, cut into overlapping windows of 25 ms (400 samples) every 10 ms (160 samples), transformed with an FFT, pooled into 80 mel-spaced frequency bands that mimic the ear's resolution, and log-compressed. The result is about 100 frames per second, each an 80-number vector. Some models use more bands; Whisper large-v3, for example, uses 128 mel bins where earlier Whisper models used 80.
import soundfile as sf
import torch, torchaudio
data, sr = sf.read("call.wav", dtype="float32", always_2d=True) # (samples, channels)
wav = torch.from_numpy(data).mean(dim=1).unsqueeze(0) # downmix to mono: (1, samples)
if sr != 16000:
wav = torchaudio.functional.resample(wav, sr, 16000)
mel = torchaudio.transforms.MelSpectrogram(
sample_rate=16000, n_fft=400, win_length=400, # 25 ms window
hop_length=160, n_mels=80)(wav) # 10 ms hop, 80 mel bins
logmel = torch.log(mel.clamp(min=1e-10)) # (1, 80, frames): ~100 frames per second
# models also apply their own normalisation (per-utterance mean/variance or fixed scaling)The front-end is where the most avoidable accuracy losses happen. Feed 8 kHz telephone audio upsampled to 16 kHz into a model trained on wideband audio and the upper half of the spectrum is empty, which the model has rarely seen. Feed 44.1 kHz audio without resampling and every frequency is shifted. Stereo files with one silent channel, clipping and aggressive noise suppression all show up here; resampling and noise suppression are worth understanding for that reason. Always match the feature pipeline the model was trained with, down to the normalisation.
Encoder: subsampling, then context
One hundred frames per second is more resolution than text needs, so encoders start with convolutional subsampling, typically reducing the frame rate by 4 (one encoder frame per 40 ms) and sometimes by 8. A 10-second clip becomes 1,000 feature frames and 250 encoder frames. This matters because self-attention cost grows with the square of sequence length.
The dominant encoder design is the Conformer: Transformer blocks with a convolution module added, so each layer models both local acoustic patterns and long-range context. Whisper uses a plain Transformer encoder over fixed 30-second windows: 3,000 feature frames, halved by a stride-2 convolution to 1,500 positions. Encoders are most of the parameters and most of the compute, and they are where pre-training on large unlabelled or weakly labelled audio pays off.
Three decoder families
The hard problem in ASR is alignment: 250 encoder frames must become perhaps 30 tokens, and nobody labelled which frames belong to which token. The three families solve that differently.
CTC (connectionist temporal classification) adds a blank symbol and predicts one label per encoder frame, independently. Training sums over every frame-level path that collapses to the reference text. Decoding collapses repeated labels and removes blanks. CTC is fast, non-autoregressive and naturally streamable, but because frames are predicted independently it has no internal language model, so it benefits a lot from an external one.
def ctc_greedy(log_probs, vocab, blank=0):
"""log_probs: (frames, vocab) from the CTC head. Collapse repeats, then drop blanks."""
best = log_probs.argmax(dim=-1).tolist()
out, prev = [], None
for t in best:
if t != prev and t != blank:
out.append(vocab[t])
prev = t
return "".join(out)
# frames: h h _ e _ l l _ l _ o o
# collapse repeats -> h _ e _ l _ l _ o ; drop blanks -> "hello"
# the blank between the two l's is what lets CTC spell a double letterTransducers (RNN-T and its variants) add a prediction network that sees the tokens emitted so far and a joint network that combines it with the current encoder frame. At each step the model either emits a token or emits blank, meaning move to the next frame. That removes CTC's independence assumption while keeping monotonic, frame-synchronous decoding, which is why transducers dominate on-device and streaming recognisers.
Attention encoder-decoders (AED), including Whisper, use an autoregressive Transformer decoder that cross-attends to all encoder frames and generates tokens like a language model. They are usually the most accurate on long, well-formed utterances and can learn extra tasks through special tokens, such as language identification, translation and timestamps. The price is that alignment is soft: nothing forces the decoder to stay on the audio, so it can repeat itself, skip content, or write fluent text for silence.
| CTC | Transducer | Attention encoder-decoder | |
|---|---|---|---|
| Alignment | Monotonic, per frame | Monotonic, per frame | Learned, soft |
| Output dependence | None between frames | Conditioned on previous tokens | Full autoregressive |
| Streaming | Natural | Natural, the usual choice | Needs chunking tricks |
| Typical failure | Misspellings, weak grammar | Deletions at high latency limits | Hallucination, repetition loops |
| Decode cost | Lowest | Moderate | Highest per token |
Many production systems combine them: a shared encoder trained with a CTC loss and an attention loss together, with CTC scores used during attention decoding to keep hypotheses aligned to the audio.
Decoding and language-model fusion
Greedy decoding takes the best token at each step. Beam search keeps the best few partial hypotheses, typically 4 to 16, and usually gains a small but real amount of accuracy. Its bigger value is that it gives you a place to add knowledge the acoustic model lacks.
The most common addition is shallow fusion with an external language model trained on text from your domain. Each hypothesis is scored as log p_asr(y|x) + lambda log p_lm(y) + beta |y|, where lambda weights the language model (often 0.1 to 0.5) and beta is a length bonus that counteracts the language model's preference for short outputs. Tune both on a development set from your own traffic, never on the test set. For names and product terms, contextual biasing boosts a list of phrases during search; the streaming article covers biasing in detail.
Two-pass systems generate an n-best list with a fast first pass and rescore it with a larger model. That is often the cheapest way to add a big language model, including an LLM, without slowing every beam step.
Text normalization and measuring WER honestly
Word error rate is (substitutions + deletions + insertions) divided by the number of reference words, computed from a word-level edit distance. It can exceed 100 percent when a model inserts many words. It is the standard metric, and it is easy to get wrong.
import re
def normalise(s):
s = s.lower()
s = re.sub(r"[^\w\s']", " ", s) # strip punctuation, keep apostrophes
return s.split()
def wer(ref, hyp):
r, h = normalise(ref), normalise(hyp)
d = list(range(len(h) + 1)) # one-row Levenshtein over words
for i in range(1, len(r) + 1):
prev, d[0] = d[0], i
for j in range(1, len(h) + 1):
cur = min(d[j] + 1, # deletion
d[j - 1] + 1, # insertion
prev + (r[i - 1] != h[j - 1])) # substitution or match
prev, d[j] = d[j], cur
return d[len(h)] / max(len(r), 1)
print(wer("Turn the kitchen lights off at nine.",
"turn the kitchen light off at night")) # 2 / 7 = 0.286In the example, the reference has 7 words and the hypothesis has two substitutions, lights to light and nine to night, so WER is 2/7, or 28.6 percent. Without the normalise step the capital T and the full stop would each count as errors too and inflate the score. The opposite problem is just as common: a model that writes 9 pm scored against a reference that says nine p m gets penalised for being correct.
So compare models only under one fixed normalization, applied identically to references and hypotheses, and report which one you used. Report WER by slice as well: accent, channel type, noise level and domain vocabulary. A single corpus-wide number hides the slices users complain about. For downstream tasks, such as extracting an order number, measure the task directly; a one-word error can be harmless or fatal depending on the word.
Long-form and batch transcription
Real recordings are minutes to hours long, and encoders are trained on short segments. The standard offline pipeline segments first, transcribes segments in parallel and stitches the results.
- Run voice activity detection to find speech regions and skip silence and music, which saves compute and removes the inputs most likely to make AED models hallucinate.
- Merge speech regions into segments near the model's preferred length, 30 seconds for Whisper, cutting in pauses rather than mid-word.
- Batch segments of similar length on the GPU. Throughput scales well with batch size because the encoder dominates cost.
- Transcribe with previous-segment text as a prompt only when you have checked it helps; it improves consistency of names but lets one hallucination propagate.
- Stitch segments, align words to timestamps (with CTC alignment or the model's own timestamp tokens), then run diarization if you need speaker labels.
Measure throughput as real-time factor, processing time divided by audio duration. Batch GPU transcription commonly runs far below 0.1, meaning an hour of audio takes a few minutes, but measure it on your own hardware, model size and batch shape before planning capacity.
Failure modes in production
- Hallucination on non-speech. AED models can emit fluent text for silence, music or noise, often stock phrases common in their training captions. Gate with VAD, discard segments with a low average log-probability or high no-speech probability, and flag outputs that are highly repetitive.
- Repetition loops. The decoder repeats a phrase until it hits the token limit. Detect with a compression-ratio check on the output, and retry the segment with sampling at a higher temperature.
- Wrong language. Automatic language identification on a short or noisy first segment picks the wrong language and the whole file follows. Set the language when you know it.
- Domain vocabulary. Drug names, product codes and surnames come out as common words. Use biasing, a domain language model or fine-tuning on a few hours of in-domain labelled audio.
- Channel mismatch. Narrowband telephony, far-field microphones and heavy compression all raise WER sharply. Evaluate on audio from the real channel, not clean headset recordings.
- Timestamp drift. Word times slide by hundreds of milliseconds across segment boundaries, which breaks subtitles and search. Re-align words after stitching.
Choosing an architecture
For live, low-latency or on-device recognition, pick a transducer or CTC model and read the streaming guide. For offline transcription of varied, long audio where accuracy matters most and cost is secondary, a large attention encoder-decoder with VAD gating is a strong default. For high-volume, cost-sensitive batch work on a narrow domain, a CTC or transducer model fine-tuned on your data plus a domain language model often matches a larger general model at a fraction of the compute.
Whatever you choose, keep three things constant across comparisons: the evaluation audio, the normalization and the decoding settings. Most claims that model A beats model B dissolve once those are held fixed.
What to do next
- Collect 1 to 5 hours of real audio from your production channel and have it transcribed carefully, with a written style guide for numbers and punctuation.
- Write one normalization function and use it for every WER you report, broken down by at least three slices.
- Verify the front-end: sample rate, channel handling and feature settings exactly as the model expects.
- Benchmark one CTC or transducer model and one attention encoder-decoder on that set, with VAD gating on both.
- Add hallucination guards: no-speech and log-probability thresholds plus a repetition check.
- Try a domain language model or phrase biasing for names and terms before paying for fine-tuning.
- Measure real-time factor at your intended batch size and plan capacity from that number.