A text-to-speech model that sounds excellent in a notebook can still fail in production, because users do not judge a voice interface by waveform quality alone. They judge it by how long the silence lasts before the voice starts and whether it ever stutters once it has. Those are serving properties. They depend on how text is cut into pieces, how requests are batched on the GPU, how audio is decoded in chunks, and how much buffer the client keeps, more than on the model architecture.
This article is about running TTS on GPUs as a service. The model side (text front ends, acoustic models, vocoders and codec language models) is explained in the TTS architecture article; here those components are treated as stages with costs. The goal is that you can set the right SLOs, estimate how many concurrent voices a GPU can carry, build a streaming pipeline that does not click or underrun, and cancel work instantly when a user interrupts.
The two numbers: time to first audio and real-time factor
Two numbers define a streaming TTS service. Time to first audio (TTFA) is the delay between the request (or the first text arriving) and the first playable audio reaching the client. Real-time factor (RTF) is the time taken to synthesise a piece of audio divided by that audio's duration. An RTF of 0.25 means four seconds of speech take one second to produce.
The subtle point is that RTF must hold per stream, continuously, not just on average across the GPU. When the client's playout buffer empties before the next chunk arrives, the user hears a gap, called an underrun. So the operational SLOs are: TTFA at p95, the fraction of streams with any underrun, and per-stream RTF at p99 measured over sliding windows of a second or two. Average RTF is a capacity metric, not a quality metric; the serving SLO article covers how to set percentile targets once you know which number is which.
The serving pipeline and where it runs
Most modern TTS services can be described as four stages with very different hardware profiles. Some models fuse stages, but the cost structure is the same.
| Stage | Runs on | Shape of the work | Batching behaviour |
|---|---|---|---|
| Text front end | CPU | normalisation, segmentation, phonemisation | embarrassingly parallel per request |
| Token generator | GPU | autoregressive: one step per audio frame or token group | continuous batching, KV cache, like an LLM |
| Decoder or vocoder | GPU | parallel convolution or transformer over a chunk of frames | fixed-size chunks batch well |
| Audio encoder | CPU | PCM framing or Opus encoding | cheap per stream, but scales with concurrency |
Non-autoregressive acoustic models skip the step-by-step generator and process a whole segment in one pass, so segment length directly sets their TTFA.
Cutting text for streaming
Streaming starts with how text is cut. Synthesising a whole paragraph before speaking wastes the user's time; synthesising word by word destroys prosody, because intonation depends on where the sentence goes. The usual compromise is to segment at sentence boundaries, with a short first segment so speech starts early, and to cap segment length so a run-on sentence does not stall the stream.
import re
SENT_END = re.compile(r"(?<=[.!?])\s+")
def segments(text_stream, first_max=60, max_chars=220):
"""Yield speakable segments from an incremental text stream (e.g. LLM tokens)."""
buf, first = "", True
for piece in text_stream:
buf += piece
while True:
limit = first_max if first else max_chars
parts = SENT_END.split(buf, maxsplit=1)
if len(parts) == 2 and parts[0]:
seg, buf = parts
elif len(buf) > limit:
cut = max(buf.rfind(",", 0, limit), buf.rfind(" ", 0, limit))
cut = cut if cut > 0 else limit
seg, buf = buf[:cut + 1], buf[cut + 1:]
else:
break
first = False
yield seg.strip()
if buf.strip():
yield buf.strip()When the text itself comes from an LLM, this segmenter sits on the LLM's token stream, and the voice can start after the first clause rather than after the whole answer. Two front-end failures recur in production: text normalisation errors ("Dr." read as "drive", "1/2" read as a date, currency and units in the wrong order) and abbreviations that the sentence splitter mistakes for sentence ends. Keep a regression set of tricky strings for your domain and run it on every front-end change; these bugs are audible and embarrassing, and the model cannot fix them.
Token generation under a frame budget
Codec-language-model TTS generates discrete audio tokens with a transformer decoder, conditioned on text and a voice prompt. To a serving engine this looks almost exactly like LLM decoding: a prefill over the conditioning, then one decode step per audio frame, each step reading the KV cache. That means the techniques in continuous batching apply directly: admit new utterances between steps, evict finished ones immediately, and page the KV cache so long utterances do not fragment memory.
The difference from text is the rate requirement. A neural codec has a fixed frame rate; EnCodec's 24 kHz model, for example, produces 75 frames per second of audio, each frame carrying several codebook tokens. If a model emits one frame per decode step (predicting the codebooks of a frame together, or with a delay pattern that staggers them), a stream needs 75 steps per second of speech just to keep up. That converts the RTF requirement into a hard ceiling on step time: one step must take less than about 13.3 ms at whatever batch size you run. Codecs with lower frame rates relax that ceiling proportionally, which is one reason codec design (see the neural codec article) matters so much for serving.
A text stream may slow slightly under load; a TTS stream that slows underruns. So cap batch size at the largest value whose p99 step time stays under the frame budget with margin, and queue new utterances past that cap.
Chunked decoding without clicks
The decoder turns frames into samples. Decoding a whole utterance at once gives the best quality but delays the first sample until the last frame exists. Streaming decodes chunks of frames as they arrive, and that introduces boundary artifacts: convolutional decoders need context on both sides, so a chunk decoded alone produces a click or a slight discontinuity where it meets the next. The standard remedy is to decode each chunk with some left context from the previous chunk, discard the samples that context produces, and crossfade a short overlap region.
import numpy as np
def stream_decode(frames_iter, decode, hop, chunk=20, left_ctx=6, fade=4):
"""decode(frames[n, ...]) -> samples of length n * hop. Yields playable audio."""
hist, ctx, tail = [], 0, None # ctx: leading frames already played
ramp = np.linspace(0.0, 1.0, fade * hop, dtype=np.float32)
for f in frames_iter:
hist.append(f)
if len(hist) < ctx + chunk + fade:
continue
audio = decode(np.stack(hist))[ctx * hop:] # drop context samples
if tail is not None: # crossfade into previous chunk
audio[:fade * hop] = tail * (1 - ramp) + audio[:fade * hop] * ramp
yield audio[:chunk * hop]
tail = audio[chunk * hop:(chunk + fade) * hop].copy()
hist = hist[ctx + chunk - left_ctx:] # keep left_ctx + fade frames
ctx = left_ctx
if len(hist) > ctx: # flush the remainder
audio = decode(np.stack(hist))[ctx * hop:]
if tail is not None:
audio[:fade * hop] = tail * (1 - ramp) + audio[:fade * hop] * ramp
yield audioChunk size is the main dial. Small chunks cut TTFA but cost more GPU launches per second of audio and more redundant context computation; larger chunks are efficient but delay the start. A common pattern is a small first chunk followed by larger steady-state chunks. Fixed chunk sizes also batch well across streams, because every stream's decode call has the same shape, which makes the decoder a good candidate for CUDA graph capture or a compiled engine.
Worked example: how many voices fit on one GPU
Capacity planning ties the stages together. The numbers below are illustrative, chosen to make the arithmetic visible; measure your own model before relying on any of them.
| Quantity | Assumption | Result |
|---|---|---|
| Frame budget | 75 frames/s, one frame per step | 13.3 ms per step |
| Measured step time | batch 48: 9 ms p50, 11 ms p99 | fits, about 17% margin at p99 |
| Measured step time | batch 64: 12 ms p50, 15 ms p99 | p99 misses: underruns under load |
| Generator capacity | batch cap 48 | 48 simultaneous speaking streams |
| Decoder cost | 20-frame chunk, batch 48: 6 ms every 20 steps | about 3% extra GPU time |
| Talk ratio | voice agent speaks 35% of session time | about 137 concurrent sessions per GPU |
| TTFA | front end 10 ms + prefill 30 ms + 24 frames x 9 ms + decode 6 ms + network 30 ms | about 290 ms |
Two lessons fall out. First, the batch cap is set by the p99 step time against the frame budget, not by memory or by average throughput: batch 64 would produce more audio in aggregate and still sound worse. Second, in conversational products sessions are mostly silent on the TTS side, so the session count a GPU supports is the speaking-stream cap divided by the talk ratio, provided admission control queues the rare bursts. The TTFA line also shows where to cut: the 24 frames the first decoded chunk needs (20 plus 4 of crossfade lookahead) dominate, so a smaller first chunk is the cheapest TTFA win.
Client buffers, formats and transport
The client is part of the latency budget. It needs a playout buffer: begin playing only after a small amount of audio has accumulated, then keep the buffer topped up. Starting with zero buffer gives the lowest TTFA and the highest underrun risk; starting with 500 ms gives smooth playback and a sluggish feel. Adaptive buffers, as used in real-time audio generally and described in the real-time audio article, grow the buffer after an underrun and shrink it slowly when delivery is steady.
On the wire, raw PCM is simplest: 24 kHz, 16-bit mono is 48,000 bytes per second, which is fine inside a data centre and heavy on mobile networks. Opus in an Ogg or WebM container cuts that by an order of magnitude at speech-quality bitrates but costs CPU per stream on the server, which must be in the capacity plan. Avoid WAV for streaming: its header declares the total data length, which is unknown when the first bytes are sent. Interactive agents usually use WebSockets, which carry text up, audio down and cancellation on one connection.
Barge-in and cancellation
In a voice agent the user will interrupt. When speech is detected on the microphone while the agent is talking, everything downstream of that decision must stop within a few hundred milliseconds: the client flushes its playout buffer, the server aborts the generator sequence (freeing its batch slot and KV cache), drops queued decoder chunks and stops the upstream LLM if it is still producing text. A partial cancel, where the GPU keeps generating audio nobody will hear, wastes exactly the capacity the worked example rationed so carefully.
async def on_barge_in(session):
session.client.send({"type": "flush_audio"}) # stop playback immediately
for req_id in session.active_tts_requests():
await tts_engine.abort(req_id) # frees generator slot + KV cache
decoder_queue.drop(req_id) # pending chunks never decoded
if session.llm_request:
await llm_engine.abort(session.llm_request)
session.record_spoken_text() # what the user actually heardThe last line matters: the agent's memory should contain what the user heard, not what it intended to say. Played-out sample counts per segment, reported by the client, make that mapping exact.
Caching, voices, failure modes and trade-offs
Caching. Many products repeat phrases: greetings, confirmations, menu prompts. Cache the encoded audio keyed by normalised text, voice and synthesis settings; a hit has near-zero TTFA and costs no GPU. For cloned or custom voices, cache the speaker embedding or prompt KV prefix per voice so each request does not recompute it.
Many voices on one GPU. Voices implemented as conditioning inputs share one model and batch freely. Voices implemented as fine-tuned weights fragment batches unless the engine supports adapter batching; plan replica pools per voice family accordingly.
Abuse controls. Voice cloning is a fraud vector. Require consent records for cloned voices, rate-limit cloning endpoints and consider watermarking output.
| Failure | Symptom | Fix |
|---|---|---|
| Batch above frame budget | stutter only at peak load | cap batch by p99 step time; queue new utterances |
| Chunk boundary artifacts | periodic clicks | left context plus crossfade; test with sine sweeps and silence |
| Over-long first segment | long silence before speech | short first segment, then sentence-sized ones |
| Normalisation bugs | numbers, dates, units misread | domain regression set; locale-aware normaliser |
| Partial cancel | capacity drops after interruptions | abort generator, decoder queue and LLM together |
| CPU encoder saturation | GPU idle, TTFA rising | count Opus encode cores in capacity plans |
The trade-offs run along one axis: latency against efficiency and quality. Shorter segments and chunks start speech sooner but cost prosody and GPU efficiency; larger batches raise throughput but threaten the per-step frame budget; bigger client buffers remove stutter but add delay to every turn of the conversation.
What to do next
- Instrument TTFA, per-stream RTF over sliding windows and the underrun rate from the client, and alert on underruns rather than average RTF.
- Measure your generator's step time at p99 across batch sizes and set the batch cap below the frame budget with at least 15% margin.
- Implement chunked decoding with left context and crossfade, then listen to chunk boundaries on a sine sweep and on silence to confirm there are no clicks.
- Add sentence segmentation with a short first segment, and a front-end regression set of numbers, dates, units and abbreviations from your domain.
- Wire barge-in end to end and verify that the generator's running-sequence count drops within one step of an interruption.
- Add a phrase cache for repeated prompts and measure its hit rate before scaling hardware.