An audio encoder turns a waveform into a sequence of vectors that a decoder, a classifier or a language model can use. Whisper, wav2vec 2.0, HuBERT and their descendants all have one, and most speech and audio-language systems spend a large share of their GPU time inside it. The encoders differ in their front end, in how they were trained and in the input shapes they accept, and those differences decide how you batch, how much padding you pay for, and what goes wrong in production.
This article explains audio encoders from the waveform up: how samples become frames, what the two dominant front ends compute, how to estimate encoder FLOPs from the frame count, and how fixed versus variable input shapes change batching, precision and streaming. It works through the cost of transcribing a short clip, lists the common failure modes, and ends with a checklist for running encoders efficiently.
From samples to frames
Speech models almost always take 16 kHz mono audio: 16,000 samples per second, which covers frequencies up to 8 kHz, enough for intelligible speech. A transformer cannot attend over 16,000 positions per second, so the first job of any encoder is to reduce the rate to tens of frames per second while keeping what matters.
The classic front end is a log-mel spectrogram. A short-time Fourier transform takes a 25 ms window (400 samples) every 10 ms (160 samples), giving 100 frames per second. Each frame's power spectrum is projected onto a bank of triangular mel filters, spaced to mimic the ear's frequency resolution, and the result is log-compressed. Whisper uses 80 mel bins in most checkpoints and 128 in large-v3. Its own preprocessing is a few lines of PyTorch, and running it on the GPU removes a common CPU bottleneck:
import torch, torch.nn.functional as F
N_FFT, HOP, SR, CHUNK = 400, 160, 16_000, 30
window = torch.hann_window(N_FFT, device="cuda")
def log_mel(audio, mel_filters): # audio: (B, samples) float32 on GPU
audio = F.pad(audio, (0, SR * CHUNK - audio.shape[-1])) # pad to 30 s
stft = torch.stft(audio, N_FFT, HOP, window=window, return_complex=True)
power = stft[..., :-1].abs() ** 2 # drop the extra frame: 3000 frames
mel = mel_filters @ power # (n_mels, 201) @ (B, 201, 3000)
log_spec = torch.clamp(mel, min=1e-10).log10()
peak = log_spec.amax(dim=(-2, -1), keepdim=True)
log_spec = torch.maximum(log_spec, peak - 8.0) # 80 dB dynamic range
return (log_spec + 4.0) / 4.0Keep this in float32. The log of small powers and the clamp at 1e-10 lose precision badly in fp16, and the cost of the front end is small next to the transformer. The mel filter matrix must match the checkpoint: Whisper ships its own filters, and a librosa or torchaudio bank with different normalisation produces inputs the model never saw.
The other front end skips the spectrogram. wav2vec 2.0 and HuBERT feed the raw waveform into a stack of seven 1-D convolutions with 512 channels, kernel widths 10, 3, 3, 3, 3, 2, 2 and strides 5, 2, 2, 2, 2, 2, 2. The total stride is 320 samples (20 ms) and the receptive field is 400 samples (25 ms), so a clip of L samples yields floor((L - 400) / 320) + 1 frames, which is 49 frames for one second. The network learns its own filter bank instead of using a fixed one.
Encoder families
Whisper (Radford et al., 2022) is an encoder-decoder trained with supervision on a large corpus of transcribed audio. Its encoder applies two 1-D convolutions with kernel 3 to the log-mel input, the first with stride 1 and the second with stride 2, then adds sinusoidal position embeddings and runs a stack of pre-norm transformer layers. Every input is exactly 30 seconds: 3,000 mel frames become 1,500 encoder positions, one per 20 ms. Shorter audio is padded to 30 seconds; longer audio is cut into windows. The large configuration has 32 encoder layers of width 1,280 with 20 heads.
wav2vec 2.0 (Baevski et al., 2020) is self-supervised. During pre-training, spans of the feature-encoder output are masked, and the transformer must identify the true quantised feature for each masked frame among distractors, a contrastive objective. Relative position comes from a wide grouped convolution rather than embeddings. The Base model has 12 layers of width 768 and about 95M parameters; Large has 24 layers of width 1,024 and about 317M. It is then fine-tuned, usually with a CTC head, on labelled speech.
HuBERT (Hsu et al., 2021) uses the same backbone but a different target: it clusters features offline with k-means, first MFCCs and then hidden states of an earlier model, and trains the transformer to predict the cluster ID of masked frames. Those discrete units turned out to be useful well beyond ASR, as tokens for speech language models and speech resynthesis.
Conformer encoders, common in production ASR, interleave convolution modules with self-attention so local acoustic patterns are modelled cheaply. They usually take log-mel input with heavier subsampling, often to 40 ms per frame, which cuts attention cost further.
Where the GPU time goes
Encoder cost follows from frame count T, width d and layer count L. In each layer the projections and MLP have about 12d2 weights, and a matrix multiply costs two FLOPs per weight per token, so the dense part costs about 24 L d2 T FLOPs. Attention scores and the weighted sum add about 4 L T2 d. The arithmetic is easy to run yourself:
def encoder_flops(T, d, L):
dense = 24 * L * d * d * T
attn = 4 * L * T * T * d
return dense, attn
# Whisper large encoder: always 1,500 positions
print([f / 1e12 for f in encoder_flops(1500, 1280, 32)]) # ~[1.89, 0.37] TFLOP
# wav2vec 2.0 Base on a 3 s clip: 149 frames
print([f / 1e9 for f in encoder_flops(149, 768, 12)]) # ~[25.3, 0.8] GFLOPTwo observations follow. First, Whisper large spends about 2.3 TFLOP per 30-second window whatever the clip length. A 3-second voice command pays the same encoder cost as 30 seconds of speech, so 90 percent of the work is spent on padding. The released checkpoints were trained on full windows, and simply feeding a shorter spectrogram tends to degrade output, so projects that trim the encoder input usually fine-tune for it and check accuracy carefully. Second, at these lengths the dense layers dominate and attention is a minority of FLOPs, so fused attention kernels such as FlashAttention help mainly through memory traffic and kernel count rather than raw arithmetic. For wav2vec models the convolutional front end is also significant: its first layer runs at the full sample rate with 512 output channels, so it is memory-bandwidth heavy and worth profiling separately.
Decoding changes the balance for encoder-decoder ASR. Whisper runs the encoder once per window, then the decoder generates tokens one at a time while cross-attending to all 1,500 encoder states. The cross-attention keys and values are computed once per window and cached. Decoding is latency-bound and small-batch; encoding is a large, regular, compute-bound batch. Profile them separately, using the methods in GPU profiling, and you will usually find they want different batch sizes.
Batching, shapes and precision
Whisper's fixed shape is a gift for the encoder. Every input is (n_mels, 3000), so any set of windows stacks into one batch without masks, and the encoder graph has one shape. That makes it a good candidate for CUDA graph capture or torch.compile with static shapes, both of which remove per-kernel launch overhead. Long recordings split into windows give you the batch directly.
wav2vec 2.0 and HuBERT take variable-length input, so batching needs padding and the transformer needs to ignore padded frames. Batching a 2-second clip with a 20-second clip makes the short one cost the same as the long one. The standard fix is length bucketing: sort pending requests by duration and form batches of similar length under a frame budget.
def frames_w2v(n_samples):
return max(0, (n_samples - 400) // 320 + 1)
def bucket(requests, max_frames=40_000, max_batch=64):
"""Group variable-length clips so padding stays small. requests: (id, n_samples)."""
batches, cur, longest = [], [], 0
for rid, n in sorted(requests, key=lambda r: r[1]):
f = frames_w2v(n)
new_longest = max(longest, f)
if cur and (new_longest * (len(cur) + 1) > max_frames or len(cur) == max_batch):
batches.append(cur)
cur, new_longest = [], f
cur.append(rid)
longest = new_longest
if cur:
batches.append(cur)
return batchesThe budget is frames times batch size, which bounds activation memory, and the attention cost per batch then scales with the longest clip squared. Measure padding waste as one minus real frames over padded frames; with bucketing it should sit in single-digit percent. How this fits a serving scheduler is covered in GPU batching strategies.
Precision follows the usual rules for transformers. Run the encoder in bf16 or fp16 with float32 accumulation, keep the front end and normalisation statistics in float32, and compare outputs against a float32 reference on a held-out set rather than trusting a few examples. Training and fine-tuning use standard mixed precision; the wav2vec models are known to be sensitive early in training, so keep loss scaling enabled with fp16 and watch for overflow.
Streaming and representations
Whisper was not designed for streaming: it sees a 30-second window and its decoder emits text for the whole window. Live captioning systems built on it run it repeatedly on a growing or sliding buffer and reconcile the overlapping outputs, which multiplies encoder cost by the number of re-runs. If latency matters more than language coverage, a model trained for streaming, typically a Conformer or transducer with chunked or limited left-context attention, is the better tool. Those encoders process each new chunk with a bounded cache of past state, so cost per second of audio is constant.
Self-supervised encoders are also used purely for representations: speaker verification, emotion, audio classification, or as the audio tower feeding an audio-language model. Then you pick which layer to read. Different layers carry different information, with intermediate layers often more phonetic and later layers more tied to the pre-training target, so probe several layers on your task, or learn a weighted sum of all layers, rather than assuming the last is best. In audio-language models, the encoder output is usually pooled or strided further, then projected into the language model's embedding space; the frame rate after that adapter decides how many tokens each second of audio costs the language model.
Failure modes
- Wrong sample rate. 44.1 kHz or 48 kHz audio passed as 16 kHz plays as slowed, low-pitched speech and the model produces fluent nonsense. Resample once, at ingestion, and assert the rate at the encoder.
- Stereo and clipping. Average channels to mono; check peak levels. Heavily clipped audio degrades recognition more than mild noise.
- Padding masks done wrong. Some wav2vec 2.0 checkpoints were trained without attention masks and expect zero-padded audio instead; others expect a mask. Hugging Face processors signal this with
return_attention_mask. Follow the checkpoint's processor configuration, or outputs shift with batch composition. - Silence and hallucination. Whisper can emit plausible text for silence or music, especially in padded windows. Run voice activity detection first, drop windows with no speech, and use the no-speech probability and log-probability thresholds that common Whisper decoders expose.
- CPU-bound front end. If the GPU is idle between batches, feature extraction or audio decoding on the CPU is the bottleneck. Decode in parallel workers and compute mel features on the GPU.
- Window boundaries. Cutting long audio at fixed 30-second marks splits words. Cut at silences found by VAD, or overlap windows and merge using timestamps.
What to do next
- Write down your encoder's frame rate and compute frames per request for your real duration distribution.
- Estimate FLOPs per request with the formula above and compare with measured throughput to find your utilisation.
- If you run Whisper on short clips, measure what fraction of encoder work is padding and decide whether a smaller, streaming or fine-tuned model is a better fit.
- For variable-length encoders, add duration bucketing under a frame budget and track padding waste.
- Move mel extraction onto the GPU in float32, and verify it matches the reference preprocessing bit for bit within tolerance.
- Capture the fixed-shape encoder in a CUDA graph or compile it with static shapes, then measure the gain.
- Add VAD in front of the encoder and assert sample rate and channel count at the boundary.
- Profile encoder and decoder separately and size their batches independently.