Serving speech-to-text looks simple: audio goes in, text comes out. In practice an ASR service spends most of its engineering on everything around the model: decoding arbitrary uploads, cutting hours of audio into pieces the model can handle, batching those pieces so the GPU stays busy, stitching results back with correct timestamps, and catching the characteristic failures of encoder-decoder models, which will happily invent text for silence. The model architecture itself, CTC, transducer and attention decoders, is covered in ASR architecture, and real-time streaming design in streaming ASR.
This article is about the serving system. It uses Whisper-style encoder-decoder models as the running example because they dominate open-source deployments and have the most surprising cost profile, and it notes where CTC and transducer models behave differently. You will leave with a request pipeline, a chunking routine, a capacity model based on real-time factor, and a list of the failures to guard against.
Two workloads, two pools
Two workloads share the name ASR and want different systems. Batch transcription handles files: call recordings, podcasts, meetings, video subtitles. The goal is throughput and cost per audio hour, latency is measured in seconds to minutes, and the server is free to reorder, chunk and batch. Streaming recognition handles live audio: voice assistants, captions, call agents. The goal is low latency to partial and final results, each session holds state on a particular worker, and the batch is whatever set of streams has audio ready at each step.
Do not serve both from one pool without thinking. A long file arriving in a streaming pool can occupy a worker for minutes, and streaming sessions in a batch pool get no latency guarantee. Separate pools with separate queues and autoscaling signals are the usual answer.
The request pipeline
A batch request passes through seven stages, shown below. Decoding and resampling convert whatever arrived (MP3, M4A, Opus in WebM, 8 kHz telephony) into 16 kHz mono samples, which is what nearly all open models expect. Voice activity detection finds speech regions (see voice activity detection). The chunk packer merges adjacent regions into windows no longer than the model accepts. The batcher groups windows from many requests, the GPU worker runs them, the stitcher maps window-relative timestamps back to the original file, and guards reject output that looks hallucinated.
Where the GPU time goes
Whisper's encoder always consumes a 30-second window: shorter audio is padded, and the encoder emits 1,500 frames regardless (the frame arithmetic is in audio encoders). So a 3-second voice note costs as much encoder time as 30 seconds of speech. The decoder then generates tokens autoregressively, one forward pass per token, attending to those 1,500 frames through cross-attention. Thirty seconds of fast speech is around 100 tokens, so on a busy GPU the decoder, not the encoder, usually dominates wall time; this is why distilled and turbo variants keep the large encoder and shrink the decoder (large-v3-turbo has 4 decoder layers instead of 32).
Three serving consequences follow. First, pack short segments together into full windows when the product allows it, because padding is wasted encoder work. Second, batch decoding across many windows, since a single sequence uses a small fraction of the GPU. Third, the decoder's self-attention cache grows with tokens while the cross-attention cache is fixed per window, so memory per active window is predictable, which makes admission control simple. CTC models are different: the encoder is the whole cost, there is no autoregressive loop, and variable-length batching by duration bucket is the main lever.
Long-form audio: chunking and stitching
Long audio must be cut. The original Whisper approach is sequential: transcribe 30 seconds, move the window to the last timestamp, condition on the previous text, repeat. It is accurate at boundaries but strictly serial, so one hour of audio is 120 dependent decoder runs, and an error such as a repetition loop propagates. The batched approach cuts with VAD first, packs speech into windows of up to 30 seconds at silence boundaries, and transcribes all windows of a file in parallel. This is what faster-whisper's BatchedInferencePipeline does. A minimal packer:
def pack_segments(speech, max_s=30.0, max_gap=1.0):
"""speech: sorted list of (start, end) seconds from VAD.
Returns windows (start, end) each at most max_s long, cut only at silences."""
windows, cur = [], None
for start, end in speech:
while end - start > max_s: # one region longer than a window
if cur: windows.append(cur); cur = None
windows.append((start, start + max_s)) # hard cut: mark for overlap stitching
start += max_s
if cur and end - cur[0] <= max_s and start - cur[1] <= max_gap:
cur = (cur[0], end) # extend through a short pause
else:
if cur: windows.append(cur)
cur = (start, end)
if cur: windows.append(cur)
return windows
print(pack_segments([(0.4, 9.8), (10.3, 22.0), (23.5, 29.0), (31.0, 70.0)]))
# [(0.4, 22.0), (23.5, 29.0), (31.0, 61.0), (61.0, 70.0)]In the example the 1.5-second pause before 23.5 exceeds max_gap, so a new window starts there, and the 39-second region at the end is cut hard at 61.0. Hard cuts can split a word; give them a second or two of overlap and drop duplicated words at the seam by timestamp. Each window's timestamps are relative to its start, so the stitcher adds the window offset back. Because packed windows drop the long silences, they also remove the inputs most likely to produce hallucinated text.
Batching and serving stacks
The batcher's job is to keep the GPU full without letting any request wait too long. Admit work by audio seconds, not request count, since one request can be 5 seconds or 5 hours. A simple policy: form a batch when it reaches the target size or the oldest item has waited a maximum delay, and interleave windows from different requests so one long file cannot starve short ones.
import asyncio, time
class WindowBatcher:
def __init__(self, run_batch, max_batch=16, max_wait=0.05):
self.q, self.run = asyncio.Queue(), run_batch
self.max_batch, self.max_wait = max_batch, max_wait
async def submit(self, window_audio):
fut = asyncio.get_running_loop().create_future()
await self.q.put((time.monotonic(), window_audio, fut))
return await fut
async def loop(self):
while True:
first = await self.q.get()
batch, deadline = [first], first[0] + self.max_wait
while len(batch) < self.max_batch:
timeout = deadline - time.monotonic()
if timeout <= 0: break
try: batch.append(await asyncio.wait_for(self.q.get(), timeout))
except asyncio.TimeoutError: break
texts = await asyncio.to_thread(self.run, [b[1] for b in batch])
for (_, _, fut), t in zip(batch, texts):
fut.set_result(t)You rarely need to write this yourself. vLLM serves Whisper behind an OpenAI-compatible /v1/audio/transcriptions endpoint and batches decoding with its continuous-batching scheduler (see continuous batching); install it with pip install vllm[audio] and note the upload cap set by VLLM_MAX_AUDIO_CLIP_FILESIZE_MB, 25 MB by default. Triton Inference Server suits multi-model pipelines, and its sequence batcher routes streaming chunks with a correlation ID so each session's state stays on one model instance (Triton in depth).
vllm serve openai/whisper-large-v3-turbo
curl http://localhost:8000/v1/audio/transcriptions \
-F model=openai/whisper-large-v3-turbo \
-F file=@meeting.wav -F language=en
Streaming sessions: sticky state, batched steps
A streaming pool has different physics. Each session sends small chunks, typically 100 to 500 ms of audio, and the model keeps per-session state between them: encoder caches for a streaming transducer, or a rolling audio buffer and the previous hypothesis for a Whisper-style model run on a sliding window. That state pins the session to one worker for its lifetime, so the load balancer must route by session, not by request. Use a session ID in the connection, WebSocket or gRPC bidirectional stream, and have the router keep a session-to-worker map, or hash the ID consistently so a router restart does not scatter live sessions.
Batching still happens, but across sessions at each step: every few tens of milliseconds the worker gathers the sessions that have a new chunk ready and runs one batched forward pass. The batch size is therefore the number of concurrently speaking sessions, and the per-step latency budget caps it. Measure the largest session count at which p95 time-to-partial stays under target, and admit new sessions only while a worker is below it; rejecting a new call is far better than slowing every live caption. When a worker must go away, stop admitting, let its sessions end or migrate at an utterance boundary, and only then terminate it.
Capacity planning with RTFx
ASR throughput is measured as inverse real-time factor, RTFx: seconds of audio processed per second of wall time. Measure it on your own audio, with your own VAD, batch size and precision, because silence ratio and speech rate change it substantially. Then capacity planning is arithmetic. Suppose your daily intake is 20,000 audio hours, peaking at three times the average rate, and one GPU sustains an RTFx of 400 at the batch size that meets your latency target. Average load is 20,000 times 3,600 divided by 86,400, about 833 audio seconds per second; peak is 2,500. That needs 2,500 / 400 = 6.25 GPUs at peak, so 8 with headroom for a GPU failure and deploys. Measure RTFx against raw intake seconds with VAD running, so silence removal is already priced in; do not subtract it a second time.
Autoscale on queued audio seconds divided by fleet RTFx, which estimates how long the backlog will take to drain, rather than on GPU utilisation, which saturates near 100 percent long before latency becomes a problem. For streaming pools, capacity is concurrent sessions per GPU at your partial-result latency target, and sessions must be drained, not killed, when scaling in.
Failure modes
- Hallucination on silence and music. Encoder-decoder models produce fluent text for non-speech. Remove silence with VAD, and use the no-speech probability (openai-whisper's default threshold is 0.6) to discard windows.
- Repetition loops. The decoder repeats a phrase until the token limit. Cap tokens per window and reject output whose gzip compression ratio exceeds a threshold (2.4 by default in openai-whisper) or whose average log-probability is below -1.0, then retry with temperature fallback.
- Wrong sample rate or channels. 8 kHz telephony fed as 16 kHz plays at double speed; stereo calls summed to mono mix two speakers. Resample explicitly and transcribe channels separately when they carry different speakers.
- Language misdetection. Detection on a short or noisy first window picks the wrong language for the whole file. Pass the language when the caller knows it.
- Head-of-line blocking. A ten-hour upload processed serially delays everyone queued behind it. Chunk first, interleave windows across requests, and cap per-request concurrency.
- Timestamp drift. Off-by-one window offsets shift every subtitle after a seam. Test stitching with synthetic audio whose word times are known.
Trade-offs
Sequential long-form decoding gives the best boundary accuracy and the worst latency; VAD-batched chunking is many times faster but depends on VAD quality and seam handling. Larger batches raise throughput and cost per hour falls, but tail latency rises. Turbo and distilled models cut decoder cost sharply for a small accuracy loss that you must measure on your domain. CTC and transducer models are cheaper and naturally streaming, while Whisper-style models are more robust across accents and noise. Measure word error rate after text normalisation on your own audio before trusting any leaderboard.
What to do next
- Build a test set of 2 to 5 hours of your real audio with reference transcripts, including silence, music and crosstalk.
- Put explicit decode, resample and VAD stages in front of the model and log the speech ratio.
- Measure RTFx and p95 latency for batch sizes 1, 4, 8 and 16 on your target GPU.
- Choose a stack: vLLM's transcription endpoint for Whisper at scale, faster-whisper for simple workers, Triton for multi-model or streaming pipelines.
- Add the no-speech, compression-ratio and log-probability guards and count how often each fires.
- Size the fleet from peak audio seconds per second divided by RTFx, and autoscale on backlog drain time.