A music generator turns a description such as "slow lo-fi hip hop with warm electric piano and vinyl crackle" into tens of seconds of audio that sounds like a coherent piece. That is a harder problem than speech synthesis. Music is polyphonic, often stereo, has structure over minutes rather than sentences, and listeners notice timing errors of a few milliseconds and pitch errors of a few cents.
This article explains how current systems are put together and why. It covers the representation problem that shapes everything else, the two generator families in use, how prompts and melodies steer the output, how to produce clips longer than the model's window, and how to serve, evaluate and safeguard a generator. The codec that produces audio tokens is covered in depth in neural audio codecs, and the closely related speech problem in text-to-speech architecture.
Why nobody generates raw samples
Start with arithmetic. Thirty seconds of stereo audio at 32 kHz is 32,000 x 30 x 2 = 1,920,000 samples. A transformer that predicted one sample per step would need nearly two million sequential steps for one clip, and its attention would have to relate a drum hit at second 2 to the same pattern at second 28, which is 832,000 samples later.
Every practical system therefore splits the job in two. A compression model maps audio to a much shorter sequence, either discrete tokens or continuous latent vectors, at roughly 20 to 100 frames per second. The generator works only in that compressed space, where 30 seconds might be 1,500 frames. A decoder then restores the full sample rate. The quality ceiling of the whole system is the quality of that decoder: if the compression model cannot reconstruct a crisp cymbal, no generator on top of it can produce one.
Two representations: discrete tokens and continuous latents
Discrete codec tokens. A neural codec such as EnCodec or SoundStream encodes each audio frame into a vector and quantizes it with residual vector quantization (RVQ). The first codebook captures the coarse content, each further codebook quantizes what the previous ones missed, and the frame becomes a small stack of integers. MusicGen uses a 32 kHz EnCodec model with 4 codebooks at 50 Hz, so one second of audio is 50 frames of 4 tokens, or 200 tokens. Tokens suit transformer language models directly: cross-entropy loss, sampling with temperature and top-k, and the whole LLM serving stack all carry over.
Continuous latents. A variational autoencoder compresses the waveform into a sequence of real-valued vectors instead. There is no quantization error, and latents suit diffusion models, which learn to denoise continuous values. Stability AI's Stable Audio models follow this route: a waveform autoencoder, a diffusion model over its latents, and text and timing conditioning.
Semantic and acoustic tokens. Some systems use two token types. Google's MusicLM, for example, first generates tokens from a self-supervised model that capture musical content and structure, and then generates codec tokens that capture timbre and recording detail, conditioned on the first stream.
Two generator families
Autoregressive token models predict the next frame of tokens given all previous frames and the conditioning. MusicGen is a single-stage transformer decoder of this type, released at 300M, 1.5B and 3.3B parameters and trained on 20K hours of licensed music. OpenAI's Jukebox was an earlier, much slower hierarchical design. Autoregressive models extend naturally: you can feed an existing clip as a prefix and continue it, which makes continuation, looping and long-form generation straightforward.
Latent diffusion models start from noise shaped like the whole latent sequence and refine all frames together over a number of denoising steps. Every step sees the full clip, so global structure such as a consistent tempo is easier, and generation time depends on the number of steps rather than the clip length. The trade is that the length is fixed when generation starts, and editing a region means inpainting rather than simply continuing.
| Property | Autoregressive over tokens | Latent diffusion |
|---|---|---|
| Generation order | Frame by frame, left to right | All frames at once, over denoising steps |
| Latency driver | Clip length times per-step cost | Number of steps times per-step cost |
| Continuation | Natural: use the clip as a prefix | Needs inpainting or outpainting |
| Streaming output | Possible, frames arrive in order | Not until the final step |
| Typical weakness | Drift and repetition over long clips | Fixed length, smeared transients |
| Sampling knobs | Temperature, top-k, top-p, guidance | Steps, sampler, guidance |
Codebook interleaving: turning a stack into a sequence
A token model must decide how to order K codebooks per frame. Three patterns are common. Flattening writes all K tokens of frame 1, then all of frame 2, and so on. It is exact but makes the sequence K times longer. Parallel prediction emits all K tokens of a frame in one step from separate heads. It is fast but ignores the dependency of a finer codebook on the coarser ones in the same frame. The delay pattern described in the MusicGen paper shifts codebook k by k steps, so at step t the model predicts codebook 1 of frame t, codebook 2 of frame t-1, and so on. Each fine token is then predicted after its coarser neighbours in the same frame are known, while the number of steps stays close to the number of frames.
For stereo, the audiocraft documentation describes encoding the left and right channels separately and interleaving their codebooks as [1_L, 1_R, 2_L, 2_R, ...]. The sketch below builds a delay pattern for a mono stream, using a special token for positions that do not exist yet.
def delay_pattern(codes, pad_id):
"""codes: list of K lists, each of length T (codebook k, frame t).
Returns K lists of length T + K - 1 in which codebook k is shifted right by k steps."""
K, T = len(codes), len(codes[0])
out = [[pad_id] * (T + K - 1) for _ in range(K)]
for k in range(K):
for t in range(T):
out[k][t + k] = codes[k][t]
return out
def undo_delay(pattern, K):
"""Inverse: read codebook k from offset k and drop the padding at both ends."""
T = len(pattern[0]) - (K - 1)
return [pattern[k][k:k + T] for k in range(K)]
# 4 codebooks, 50 frames per second, 30 seconds: 1,500 frames, 1,503 generation steps,
# versus 6,000 steps when the 4 codebooks are flattened.
Conditioning and classifier-free guidance
Text reaches the generator through a pretrained text encoder. MusicGen's paper uses a frozen T5 encoder, and the generator attends to its output embeddings with cross-attention at every layer. Melody conditioning, available in the musicgen-melody checkpoint, extracts a chromagram from reference audio, a frame-by-frame summary of energy in each of the twelve pitch classes. It keeps the harmonic and melodic outline while discarding timbre, so the model can re-orchestrate a hummed tune. Diffusion systems such as Stable Audio also condition on timing values, the start offset and total length of the clip, which lets one model generate different durations and helps it place endings.
Almost every system then applies classifier-free guidance. During training the condition is randomly dropped for a fraction of examples, so the same network learns both a conditional and an unconditional prediction. At inference the two are combined to push the output further towards the prompt:
# One sampling step with classifier-free guidance, for a token model.
logits_cond = model(prefix, condition=text_embedding)
logits_uncond = model(prefix, condition=None) # same weights, empty condition
logits = logits_uncond + guidance * (logits_cond - logits_uncond)
next_tokens = sample_top_k(logits / temperature, k=250)Guidance above 1 trades diversity for prompt adherence. Too low and the output ignores the prompt; too high and it becomes repetitive and harsh.
Worked example: a 30-second clip with MusicGen
The audiocraft library exposes MusicGen through a small API. The documentation shows MusicGen.get_pretrained, set_generation_params(duration=...), generate(descriptions) and generate_with_chroma for melody conditioning. A minimal script looks like this:
import torchaudio
from audiocraft.models import MusicGen
from audiocraft.data.audio import audio_write
model = MusicGen.get_pretrained("facebook/musicgen-melody")
model.set_generation_params(duration=30)
descriptions = ["slow lo-fi hip hop, warm electric piano, vinyl crackle, 80 bpm"]
melody, sr = torchaudio.load("hummed_tune.wav") # the melody to follow
wav = model.generate_with_chroma(descriptions, melody[None], sr)
for i, one in enumerate(wav):
audio_write(f"clip_{i}", one.cpu(), model.sample_rate, strategy="loudness")Trace the numbers. The model produces 30 x 50 = 1,500 frames, or a little over 1,500 decoder steps with the delay pattern. Each step runs the transformer once per guidance branch over a growing key-value cache. The EnCodec decoder then turns 1,500 x 4 tokens into 960,000 samples at 32 kHz. Doubling the duration roughly doubles the steps, and attention cost grows faster than that, which is one reason models are trained on fixed-length windows.
To go beyond the training window, generate in overlapping chunks: keep the last several seconds of the previous chunk as the prefix of the next, generate the continuation, and cross-fade at the seam. The prefix carries tempo and key forward, but not larger structure. The model does not know it is in the second chorus, so long pieces tend to wander. Systems that need song-level form usually add a planning layer, such as a section list with per-section prompts, above the generator.
Training data and the pipeline behind the model
The data pipeline decides what the model can do and what legal risk it carries. A typical pipeline ingests licensed tracks, resamples them to the codec rate (see audio resampling), cuts them into training windows, and attaches text. Text comes from catalogue metadata such as genre, instruments, mood and tempo, sometimes expanded into captions by a tagging or captioning model. Near-duplicate removal matters twice: duplicates waste compute, and a track seen many times is more likely to be memorized and reproduced.
Serving a music generator
Serving looks like LLM serving with a very long output and an audio decoder at the end. The operational decisions are these:
- Queue, do not block. A 30-second clip takes seconds to tens of seconds of GPU time. Accept the request, return a job ID, and deliver the result by polling or callback.
- Batch across requests. Token models batch well with continuous batching; group requests with similar durations so short clips do not wait on long ones. Batch the two guidance branches as one forward pass.
- Stream when you can. An autoregressive model can decode and send audio in chunks as frames arrive, so the listener hears the start while the rest is generated. A diffusion model cannot do this.
- Normalize loudness on the way out. Generated clips vary widely in level; normalize to a target loudness as described in loudness normalization so a playlist of outputs does not jump in volume.
Evaluating output
No single metric captures musical quality, so teams combine several. Fréchet Audio Distance (FAD) compares the distribution of embeddings of generated clips with that of real music; lower is better, but it depends strongly on which embedding model is used, so compare numbers only within one setup. A classifier-based KL divergence checks whether generated clips get the same tag distributions as reference clips with the same prompt. A text-audio similarity score from a joint embedding model such as CLAP measures prompt adherence. None of these hears a missed downbeat or a chord that clashes, so every serious evaluation also runs blind listening tests that rate overall quality and prompt match separately.
Safeguards
A music generator can reproduce training material and imitate identifiable artists, so safeguards are part of the architecture, not an afterthought. On the input side, filter or rewrite prompts that name specific artists or songs when your policy or licence requires it. On the output side, compare generated audio against an index of the training set using audio fingerprinting or embedding similarity, and block or regenerate clips that match too closely. Embed a watermark in every output so generated audio can be identified later; Meta's AudioSeal is one published example of a neural audio watermark designed to survive common edits. Keep provenance records linking each output to the model version and request.
Failure modes
| Symptom | Likely cause | Mitigation |
|---|---|---|
| Tempo drifts or the beat stumbles | Autoregressive drift, prefix too short | Longer prefix for continuation, tempo in the prompt, beat-tracker check with regeneration |
| Metallic, smeared cymbals and transients | Codec or VAE reconstruction limits | Better codec, more codebooks, a higher-rate decoder; this is a ceiling, not a sampling bug |
| Output ignores the prompt | Guidance too low, prompt vocabulary unlike training captions | Raise guidance, rewrite prompts into catalogue-style tags |
| Loops the same bar forever | Low temperature, top-k too small | Raise temperature or top-k, add a repetition check |
| Audible seam between chunks | No cross-fade, key or tempo change at the boundary | Overlap and cross-fade, keep more context in the prefix |
| Recognizable melody from a real song | Memorization of duplicated training data | Deduplicate training data, similarity check on outputs |
What to do next
- Generate the same prompt with an autoregressive model and a latent diffusion model, and write down where each fails.
- Compute the frame and token budget for your target duration before choosing a model: frames per second, codebooks, and steps with and without the delay pattern.
- Sweep guidance and temperature on 20 fixed prompts with fixed seeds, and pick defaults with a small blind listening test.
- Build a chunked continuation loop with overlap and cross-fade, and test it on a 2-minute piece.
- Add loudness normalization, a watermark and a training-set similarity check to the output path before anyone outside the team hears the results.
- Log model version, seed and every sampling parameter with each output so any clip can be reproduced.