A text-to-speech (TTS) system turns a few bytes of text into tens of thousands of audio samples per second. At 24 kHz, one second of speech is 24,000 samples, and nothing in the text says how fast to speak, where to pause, which syllable to stress or what the voice should sound like. TTS is therefore a one-to-many problem: many waveforms are correct for the same sentence. Every architecture is a way of breaking that gap into steps a model can learn.

This article follows the pipeline from raw text to a stream of audio. It covers the text front end, the main acoustic model families, the intermediate representations that connect them to a waveform, vocoders and neural codecs, and the streaming and serving design that decides the latency users actually hear. It closes with evaluation, failure modes, trade-offs and a checklist. It does not rank specific commercial voices or models, which change monthly.

Advertisement

The pipeline at a glance

A text-to-speech pipeline: text to units to acoustic representation to waveformRaw text or SSMLfrom app or LLMNormalizationnumbers, dates, unitsG2P or tokenizerphonemes or charsAcoustic modeldurations, prosodyIntermediatemel frames or codec tokensVocoder or decoderHiFi-GAN, codecPost-processingresample, loudnessStream to clientchunks, crossfadeSpeaker and styleembedding or promptEach stage can fail on its own: normalization errors are wrong words, acoustic errors are skips or repeats,vocoder errors are buzz or metallic artefacts, and streaming errors are clicks and gaps between chunks.
A TTS pipeline. Classic systems use every box; end-to-end models merge the acoustic model and vocoder, but the front end and streaming stages remain.

Almost every system has the same four stages, even where a single network implements several of them. The front end turns written text into speakable units. The acoustic model decides timing and prosody and produces an intermediate acoustic representation. The vocoder or codec decoder turns that representation into a waveform. Post-processing and streaming deliver it at the right sample rate and loudness, in chunks, without audible seams. The key architectural choice is the intermediate representation: a mel spectrogram or discrete neural-codec tokens.

The text front end

Text normalization expands everything that is written differently from how it is spoken. "$12.50" becomes "twelve dollars and fifty cents", "3 kg" becomes "three kilograms", "Dr." becomes "doctor" or "drive" depending on context, and "1998" is a year in one sentence and a quantity in the next. This stage causes a large share of the errors users notice, because a perfect voice reading the wrong words is still wrong. Production systems use weighted rule grammars, a trained model, or both, with a large test set of tricky cases.

import re

UNITS = {"km": "kilometres", "kg": "kilograms", "%": "percent"}

def say_number(n: int) -> str:
    ones = "zero one two three four five six seven eight nine ten eleven twelve thirteen fourteen fifteen sixteen seventeen eighteen nineteen".split()
    tens = "_ _ twenty thirty forty fifty sixty seventy eighty ninety".split()
    if n < 20: return ones[n]
    if n < 100: return tens[n // 10] + ("" if n % 10 == 0 else "-" + ones[n % 10])
    if n < 1000:
        rest = n % 100
        return ones[n // 100] + " hundred" + ("" if rest == 0 else " and " + say_number(rest))
    raise ValueError("extend for larger numbers, or route to a proper verbalizer")

def normalize(text: str) -> str:
    text = re.sub(r"\$(\d+)\.(\d\d)\b",
                  lambda m: f"{say_number(int(m[1]))} dollars and {say_number(int(m[2]))} cents", text)
    text = re.sub(r"(\d+)\s?(km|kg|%)",
                  lambda m: f"{say_number(int(m[1]))} {UNITS[m[2]]}", text)
    text = re.sub(r"\b(\d+)\b", lambda m: say_number(int(m[1])), text)
    return text

print(normalize("Pay $12.50 for 3 kg, 40% off."))
# Pay twelve dollars and fifty cents for three kilograms, forty percent off.

This toy normalizer shows the shape of the problem: ordered rewrites, where the currency rule must run before the bare-number rule. After normalization, grapheme-to-phoneme (G2P) conversion maps words to phonemes, using a pronunciation lexicon first and a learned model for words that are not in it. Many recent models skip explicit phonemes and consume characters or subword tokens directly. That removes a component, but it moves pronunciation errors into the acoustic model, where they are harder to fix. Keep a way to force a pronunciation, such as SSML phoneme tags or a custom lexicon, for names and product terms.

Advertisement

Acoustic model families

The acoustic model must solve alignment: how many output frames each input unit lasts. The main families differ mainly in how they do that.

FamilyExampleAlignmentCharacter
Autoregressive with attentionTacotron 2learned attention, one frame at a timenatural prosody; can skip or repeat words; sequential
Non-autoregressive with durationsFastSpeech 2duration predictor trained on forced alignmentsfast and robust; needs an aligner; prosody from pitch and energy predictors
End-to-endVITSmonotonic alignment search during trainingoutputs a waveform directly; VAE, flows and adversarial training
Codec language modelVALL-Eimplicit, by next-token predictionstrong zero-shot voice cloning; autoregressive failure modes return
Flow matchingF5-TTSimplicit; text padded to the length of the speechnon-autoregressive, parallel; quality depends on sampling steps

Autoregressive attention models learn alignment from the data. When the attention slips, the model skips a word, repeats a phrase or babbles, and long or unusual inputs make this more likely. Non-autoregressive models such as FastSpeech 2 predict an explicit duration per phoneme, expand the encoder outputs to frame length and generate all frames in parallel. That makes them fast and nearly immune to skipping, at the cost of an external aligner at training time and somewhat flatter prosody unless pitch and energy are modelled. VITS combines a conditional variational autoencoder, normalizing flows and a GAN-trained decoder, and produces a waveform directly.

F5-TTS, from 2024, is a fully non-autoregressive flow-matching model built on a Diffusion Transformer. It was trained on 100K hours of public multilingual data, and its authors report a real-time factor of 0.15 and an inference-time sampling schedule they call Sway Sampling. Flow matching itself is explained in flow matching.

Mel spectrograms and vocoders

A mel spectrogram is a short-time Fourier transform whose frequency axis is warped to the mel scale and pooled into typically 80 bands, stored as log magnitudes. It keeps what the ear cares about and discards phase. That is why it is easy to predict, and also why a separate model is needed to reconstruct a waveform.

import torch, torchaudio

wav, sr = torchaudio.load("speech.wav")                  # mono, 22,050 Hz in this example
mel = torchaudio.transforms.MelSpectrogram(
    sample_rate=sr, n_fft=1024, win_length=1024, hop_length=256,
    n_mels=80, f_min=0.0, f_max=8000.0, power=1.0)(wav)
log_mel = torch.log(torch.clamp(mel, min=1e-5))            # what acoustic models predict
print(log_mel.shape)   # [1, 80, frames]; frames = about sr / 256 = 86 per second

With a hop of 256 samples at 22,050 Hz, the acoustic model produces about 86 frames per second, and the vocoder must upsample each frame to 256 samples. The analysis parameters are a contract between the two models: a vocoder trained on one hop, window, band count or frequency range produces noise when fed another. Record them with every checkpoint.

Neural vocoders replaced signal-processing reconstruction because they sound far better. WaveNet showed the quality ceiling but generated sample by sample. GAN vocoders such as HiFi-GAN generate all samples in parallel with transposed convolutions, and are trained against discriminators that look at the waveform at several periods and scales. They are the usual choice behind mel-based acoustic models.

Codec language models

A neural audio codec, covered in neural audio codecs, compresses audio into a few parallel streams of discrete tokens using residual vector quantization. EnCodec's 24 kHz model, for example, emits frames at 75 Hz with 1,024-entry codebooks, and 8 codebooks give 6 kbps. That turns speech into something a language model can predict. VALL-E (Wang et al., 2023) treats TTS as conditional language modelling over these codes: given phonemes and a 3-second recording of an unseen speaker as an acoustic prompt, it continues the codes in that voice. It was trained on 60K hours of English speech. Its design splits the work: an autoregressive model predicts the first, coarse codebook token by token, and a non-autoregressive model fills in the remaining codebooks in parallel.

The token budget explains the architecture. Eight codebooks at 75 Hz are 600 tokens per second, which would be slow to generate autoregressively, so only the first 75 per second are generated sequentially. The benefits are strong zero-shot cloning and prosody that follows the prompt. The costs are the old autoregressive failure modes (skips, repeats and runaway generation) and a cloning capability that needs consent and abuse controls.

Streaming and the latency budget

Two numbers describe TTS speed. Real-time factor (RTF) is synthesis time divided by audio duration: at RTF 0.1, ten seconds of audio take one second to make, and RTF must stay well below 1 per stream under load. Time to first audio (TTFA) is the delay from receiving text to the first sound at the listener, and it is what users perceive in voice agents. Streaming synthesis splits text at clause or sentence boundaries and synthesizes each piece as soon as it arrives, so that TTFA depends on the first chunk rather than the whole reply.

Worked example. A voice agent streams LLM tokens into TTS. The first clause, about eight words, is complete 180 ms after the LLM starts. The front end takes 5 ms. A non-autoregressive acoustic model at a per-stream RTF of 0.05 needs 50 ms for that clause's roughly one second of audio. The vocoder at RTF 0.02 adds 20 ms. Encoding, network and a 60 ms client jitter buffer add about 90 ms. TTFA is 180 + 5 + 50 + 20 + 90 = 345 ms. The largest item is waiting for text, so the biggest win is a shorter first chunk. An autoregressive codec model with its first chunk at RTF 0.3 would add about 250 ms, which is the architectural price of zero-shot cloning in this budget.

import numpy as np

def stream_tts(text_chunks, frontend, acoustic, vocoder, speaker, xfade=256):
    """Yield audio as soon as each clause is synthesized, crossfading chunk seams."""
    tail = None
    for clause in text_chunks:                 # split on sentence or clause boundaries
        units = frontend(clause)               # normalize + G2P
        feats = acoustic(units, speaker=speaker)   # mel frames or codec tokens
        audio = vocoder(feats)                 # float32 PCM
        if tail is not None:                   # linear crossfade hides the seam
            ramp = np.linspace(0.0, 1.0, xfade, dtype=np.float32)
            audio[:xfade] = audio[:xfade] * ramp + tail * (1.0 - ramp)
        tail = audio[-xfade:].copy()
        yield audio[:-xfade]                   # hold back the tail for the next seam
    if tail is not None:
        yield tail

Crossfading a few milliseconds of overlap hides level jumps at the seams, but prosody still resets at each chunk unless the model is conditioned on the previous chunk. The capture, buffering and playback side of the budget is covered in real-time audio architecture.

Serving architecture

TTS serving looks like LLM serving with an extra stage. Non-autoregressive models batch naturally, because all frames of a request are generated at once, so batch requests of similar length together. Autoregressive codec models need continuous batching like any decoder, plus a parallel stage for the remaining codebooks. Run the vocoder on the same GPU to avoid copying features between devices, or on CPU when the acoustic model is the bottleneck. Cache audio for frequent fixed prompts, keyed by text, voice, model version and settings.

Output format is part of the contract. Generate at the model's native rate. If clients need a different rate, resample once, with a proper polyphase filter, and apply loudness normalization so voices and chunks sit at a consistent level.

Evaluating TTS

Subjective listening tests remain the reference: mean opinion score (MOS) for naturalness, comparative tests such as MUSHRA or A/B preference, and similarity ratings for cloned voices. They are slow and noisy, so pipelines add automatic proxies. Intelligibility is measured by transcribing the output with an ASR model and computing word error rate against the input, which catches skips, repeats and mispronunciations. Speaker similarity is the cosine similarity of speaker-verification embeddings between the output and the reference voice. Predicted MOS models such as UTMOS give a cheap naturalness estimate that is good for tracking regressions and poor for comparing very different systems. Keep a fixed set of hard test sentences and listen to samples from every release.

Failure modes

  • Normalization errors. Wrong expansion of numbers, dates, abbreviations or currency. They are frequent, very noticeable and fixable only in the front end.
  • Alignment failures. Skipped words, repeats or endless babble in autoregressive models, especially on long or repetitive input. Cap output length relative to input length and check ASR WER.
  • Parameter mismatch. A vocoder fed features with the wrong hop, band count or normalization produces buzz or noise.
  • Chunk seams. Clicks, level jumps or prosody resets at every streaming boundary.
  • Out-of-domain voices. Cloning from noisy or reverberant prompts copies the noise, because codec models reproduce the acoustic environment of the prompt too.
  • Misuse. Voice cloning without consent. Require verified consent for custom voices, log synthesis, and consider watermarking output.

Trade-offs

ChoiceGainCost
Phonemes vs characters as inputcontrollable pronunciationlexicon and G2P to maintain
Non-autoregressive vs autoregressivespeed, robustness, easy batchingflatter prosody unless modelled
Mel plus vocoder vs codec tokensmature, cheap, well understoodweaker zero-shot cloning
Small first chunklower TTFAworse prosody on the opening phrase
Fewer flow-matching stepslower RTFquality drop

What to do next

  1. Write down your latency target as TTFA and per-stream RTF at peak concurrency, then choose the model family that can meet it.
  2. Build a normalization test set from real text in your domain, including numbers, dates, currency, names and abbreviations, and run it in CI.
  3. Add a custom lexicon or SSML phoneme overrides for product names and proper nouns.
  4. Record the audio front-end parameters (sample rate, hop, window, band count) with every acoustic model and vocoder checkpoint, and assert that they match at load time.
  5. Automate ASR-WER, speaker similarity and predicted-MOS checks on a fixed hard-text set, and listen to samples before each release.
  6. Implement clause-level streaming with crossfaded seams, measure TTFA end to end, and require consent for any custom voice.
Key takeaway: A TTS system is a front end that turns text into speakable units, an acoustic model that decides timing and prosody, a vocoder or codec decoder that produces the waveform, and a streaming layer that delivers it. The architectural choices are the input units, how alignment is handled (attention, durations, monotonic search or implicit) and whether the intermediate representation is a mel spectrogram or codec tokens. Those choices set robustness, cloning ability and latency. Put engineering effort into normalization, the feature contract between models and the streaming budget, and measure intelligibility automatically on every release.