A text-to-speech (TTS) system turns a few bytes of text into tens of thousands of audio samples per second. At 24 kHz, one second of speech is 24,000 samples, and nothing in the text says how fast to speak, where to pause, which syllable to stress or what the voice should sound like. TTS is therefore a one-to-many problem: many waveforms are correct for the same sentence. Every architecture is a way of breaking that gap into steps a model can learn.
This article follows the pipeline from raw text to a stream of audio. It covers the text front end, the main acoustic model families, the intermediate representations that connect them to a waveform, vocoders and neural codecs, and the streaming and serving design that decides the latency users actually hear. It closes with evaluation, failure modes, trade-offs and a checklist. It does not rank specific commercial voices or models, which change monthly.
The pipeline at a glance
Almost every system has the same four stages, even where a single network implements several of them. The front end turns written text into speakable units. The acoustic model decides timing and prosody and produces an intermediate acoustic representation. The vocoder or codec decoder turns that representation into a waveform. Post-processing and streaming deliver it at the right sample rate and loudness, in chunks, without audible seams. The key architectural choice is the intermediate representation: a mel spectrogram or discrete neural-codec tokens.
The text front end
Text normalization expands everything that is written differently from how it is spoken. "$12.50" becomes "twelve dollars and fifty cents", "3 kg" becomes "three kilograms", "Dr." becomes "doctor" or "drive" depending on context, and "1998" is a year in one sentence and a quantity in the next. This stage causes a large share of the errors users notice, because a perfect voice reading the wrong words is still wrong. Production systems use weighted rule grammars, a trained model, or both, with a large test set of tricky cases.
import re
UNITS = {"km": "kilometres", "kg": "kilograms", "%": "percent"}
def say_number(n: int) -> str:
ones = "zero one two three four five six seven eight nine ten eleven twelve thirteen fourteen fifteen sixteen seventeen eighteen nineteen".split()
tens = "_ _ twenty thirty forty fifty sixty seventy eighty ninety".split()
if n < 20: return ones[n]
if n < 100: return tens[n // 10] + ("" if n % 10 == 0 else "-" + ones[n % 10])
if n < 1000:
rest = n % 100
return ones[n // 100] + " hundred" + ("" if rest == 0 else " and " + say_number(rest))
raise ValueError("extend for larger numbers, or route to a proper verbalizer")
def normalize(text: str) -> str:
text = re.sub(r"\$(\d+)\.(\d\d)\b",
lambda m: f"{say_number(int(m[1]))} dollars and {say_number(int(m[2]))} cents", text)
text = re.sub(r"(\d+)\s?(km|kg|%)",
lambda m: f"{say_number(int(m[1]))} {UNITS[m[2]]}", text)
text = re.sub(r"\b(\d+)\b", lambda m: say_number(int(m[1])), text)
return text
print(normalize("Pay $12.50 for 3 kg, 40% off."))
# Pay twelve dollars and fifty cents for three kilograms, forty percent off.This toy normalizer shows the shape of the problem: ordered rewrites, where the currency rule must run before the bare-number rule. After normalization, grapheme-to-phoneme (G2P) conversion maps words to phonemes, using a pronunciation lexicon first and a learned model for words that are not in it. Many recent models skip explicit phonemes and consume characters or subword tokens directly. That removes a component, but it moves pronunciation errors into the acoustic model, where they are harder to fix. Keep a way to force a pronunciation, such as SSML phoneme tags or a custom lexicon, for names and product terms.
Acoustic model families
The acoustic model must solve alignment: how many output frames each input unit lasts. The main families differ mainly in how they do that.
| Family | Example | Alignment | Character |
|---|---|---|---|
| Autoregressive with attention | Tacotron 2 | learned attention, one frame at a time | natural prosody; can skip or repeat words; sequential |
| Non-autoregressive with durations | FastSpeech 2 | duration predictor trained on forced alignments | fast and robust; needs an aligner; prosody from pitch and energy predictors |
| End-to-end | VITS | monotonic alignment search during training | outputs a waveform directly; VAE, flows and adversarial training |
| Codec language model | VALL-E | implicit, by next-token prediction | strong zero-shot voice cloning; autoregressive failure modes return |
| Flow matching | F5-TTS | implicit; text padded to the length of the speech | non-autoregressive, parallel; quality depends on sampling steps |
Autoregressive attention models learn alignment from the data. When the attention slips, the model skips a word, repeats a phrase or babbles, and long or unusual inputs make this more likely. Non-autoregressive models such as FastSpeech 2 predict an explicit duration per phoneme, expand the encoder outputs to frame length and generate all frames in parallel. That makes them fast and nearly immune to skipping, at the cost of an external aligner at training time and somewhat flatter prosody unless pitch and energy are modelled. VITS combines a conditional variational autoencoder, normalizing flows and a GAN-trained decoder, and produces a waveform directly.
F5-TTS, from 2024, is a fully non-autoregressive flow-matching model built on a Diffusion Transformer. It was trained on 100K hours of public multilingual data, and its authors report a real-time factor of 0.15 and an inference-time sampling schedule they call Sway Sampling. Flow matching itself is explained in flow matching.
Mel spectrograms and vocoders
A mel spectrogram is a short-time Fourier transform whose frequency axis is warped to the mel scale and pooled into typically 80 bands, stored as log magnitudes. It keeps what the ear cares about and discards phase. That is why it is easy to predict, and also why a separate model is needed to reconstruct a waveform.
import torch, torchaudio
wav, sr = torchaudio.load("speech.wav") # mono, 22,050 Hz in this example
mel = torchaudio.transforms.MelSpectrogram(
sample_rate=sr, n_fft=1024, win_length=1024, hop_length=256,
n_mels=80, f_min=0.0, f_max=8000.0, power=1.0)(wav)
log_mel = torch.log(torch.clamp(mel, min=1e-5)) # what acoustic models predict
print(log_mel.shape) # [1, 80, frames]; frames = about sr / 256 = 86 per secondWith a hop of 256 samples at 22,050 Hz, the acoustic model produces about 86 frames per second, and the vocoder must upsample each frame to 256 samples. The analysis parameters are a contract between the two models: a vocoder trained on one hop, window, band count or frequency range produces noise when fed another. Record them with every checkpoint.
Neural vocoders replaced signal-processing reconstruction because they sound far better. WaveNet showed the quality ceiling but generated sample by sample. GAN vocoders such as HiFi-GAN generate all samples in parallel with transposed convolutions, and are trained against discriminators that look at the waveform at several periods and scales. They are the usual choice behind mel-based acoustic models.
Codec language models
A neural audio codec, covered in neural audio codecs, compresses audio into a few parallel streams of discrete tokens using residual vector quantization. EnCodec's 24 kHz model, for example, emits frames at 75 Hz with 1,024-entry codebooks, and 8 codebooks give 6 kbps. That turns speech into something a language model can predict. VALL-E (Wang et al., 2023) treats TTS as conditional language modelling over these codes: given phonemes and a 3-second recording of an unseen speaker as an acoustic prompt, it continues the codes in that voice. It was trained on 60K hours of English speech. Its design splits the work: an autoregressive model predicts the first, coarse codebook token by token, and a non-autoregressive model fills in the remaining codebooks in parallel.
The token budget explains the architecture. Eight codebooks at 75 Hz are 600 tokens per second, which would be slow to generate autoregressively, so only the first 75 per second are generated sequentially. The benefits are strong zero-shot cloning and prosody that follows the prompt. The costs are the old autoregressive failure modes (skips, repeats and runaway generation) and a cloning capability that needs consent and abuse controls.
Streaming and the latency budget
Two numbers describe TTS speed. Real-time factor (RTF) is synthesis time divided by audio duration: at RTF 0.1, ten seconds of audio take one second to make, and RTF must stay well below 1 per stream under load. Time to first audio (TTFA) is the delay from receiving text to the first sound at the listener, and it is what users perceive in voice agents. Streaming synthesis splits text at clause or sentence boundaries and synthesizes each piece as soon as it arrives, so that TTFA depends on the first chunk rather than the whole reply.
Worked example. A voice agent streams LLM tokens into TTS. The first clause, about eight words, is complete 180 ms after the LLM starts. The front end takes 5 ms. A non-autoregressive acoustic model at a per-stream RTF of 0.05 needs 50 ms for that clause's roughly one second of audio. The vocoder at RTF 0.02 adds 20 ms. Encoding, network and a 60 ms client jitter buffer add about 90 ms. TTFA is 180 + 5 + 50 + 20 + 90 = 345 ms. The largest item is waiting for text, so the biggest win is a shorter first chunk. An autoregressive codec model with its first chunk at RTF 0.3 would add about 250 ms, which is the architectural price of zero-shot cloning in this budget.
import numpy as np
def stream_tts(text_chunks, frontend, acoustic, vocoder, speaker, xfade=256):
"""Yield audio as soon as each clause is synthesized, crossfading chunk seams."""
tail = None
for clause in text_chunks: # split on sentence or clause boundaries
units = frontend(clause) # normalize + G2P
feats = acoustic(units, speaker=speaker) # mel frames or codec tokens
audio = vocoder(feats) # float32 PCM
if tail is not None: # linear crossfade hides the seam
ramp = np.linspace(0.0, 1.0, xfade, dtype=np.float32)
audio[:xfade] = audio[:xfade] * ramp + tail * (1.0 - ramp)
tail = audio[-xfade:].copy()
yield audio[:-xfade] # hold back the tail for the next seam
if tail is not None:
yield tailCrossfading a few milliseconds of overlap hides level jumps at the seams, but prosody still resets at each chunk unless the model is conditioned on the previous chunk. The capture, buffering and playback side of the budget is covered in real-time audio architecture.
Serving architecture
TTS serving looks like LLM serving with an extra stage. Non-autoregressive models batch naturally, because all frames of a request are generated at once, so batch requests of similar length together. Autoregressive codec models need continuous batching like any decoder, plus a parallel stage for the remaining codebooks. Run the vocoder on the same GPU to avoid copying features between devices, or on CPU when the acoustic model is the bottleneck. Cache audio for frequent fixed prompts, keyed by text, voice, model version and settings.
Output format is part of the contract. Generate at the model's native rate. If clients need a different rate, resample once, with a proper polyphase filter, and apply loudness normalization so voices and chunks sit at a consistent level.
Evaluating TTS
Subjective listening tests remain the reference: mean opinion score (MOS) for naturalness, comparative tests such as MUSHRA or A/B preference, and similarity ratings for cloned voices. They are slow and noisy, so pipelines add automatic proxies. Intelligibility is measured by transcribing the output with an ASR model and computing word error rate against the input, which catches skips, repeats and mispronunciations. Speaker similarity is the cosine similarity of speaker-verification embeddings between the output and the reference voice. Predicted MOS models such as UTMOS give a cheap naturalness estimate that is good for tracking regressions and poor for comparing very different systems. Keep a fixed set of hard test sentences and listen to samples from every release.
Failure modes
- Normalization errors. Wrong expansion of numbers, dates, abbreviations or currency. They are frequent, very noticeable and fixable only in the front end.
- Alignment failures. Skipped words, repeats or endless babble in autoregressive models, especially on long or repetitive input. Cap output length relative to input length and check ASR WER.
- Parameter mismatch. A vocoder fed features with the wrong hop, band count or normalization produces buzz or noise.
- Chunk seams. Clicks, level jumps or prosody resets at every streaming boundary.
- Out-of-domain voices. Cloning from noisy or reverberant prompts copies the noise, because codec models reproduce the acoustic environment of the prompt too.
- Misuse. Voice cloning without consent. Require verified consent for custom voices, log synthesis, and consider watermarking output.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Phonemes vs characters as input | controllable pronunciation | lexicon and G2P to maintain |
| Non-autoregressive vs autoregressive | speed, robustness, easy batching | flatter prosody unless modelled |
| Mel plus vocoder vs codec tokens | mature, cheap, well understood | weaker zero-shot cloning |
| Small first chunk | lower TTFA | worse prosody on the opening phrase |
| Fewer flow-matching steps | lower RTF | quality drop |
What to do next
- Write down your latency target as TTFA and per-stream RTF at peak concurrency, then choose the model family that can meet it.
- Build a normalization test set from real text in your domain, including numbers, dates, currency, names and abbreviations, and run it in CI.
- Add a custom lexicon or SSML phoneme overrides for product names and proper nouns.
- Record the audio front-end parameters (sample rate, hop, window, band count) with every acoustic model and vocoder checkpoint, and assert that they match at load time.
- Automate ASR-WER, speaker similarity and predicted-MOS checks on a fixed hard-text set, and listen to samples before each release.
- Implement clause-level streaming with crossfaded seams, measure TTFA end to end, and require consent for any custom voice.