Data order is the training decision people make without noticing. A loader reads shards in the order they were written, a fine-tuning set is the concatenation of five source files, a streaming pipeline shuffles inside a buffer that is far smaller than one source, and the run trains on what amounts to a sequence of topics rather than a mixture. Nothing crashes. The loss curve just has more structure than it should, and the model is quietly better at whatever it saw last.

This article treats ordering as an engineering problem with two halves. Shuffling is about making each batch a fair sample of the data you meant to train on. Curriculum is about deliberately changing that sample over time. Both have to be deterministic so that a resumed or repeated run sees the same order. Filtering, deduplication and static mixture weights are covered in pretraining data for small language models; this page starts after those decisions are made.

Advertisement

Why order matters to SGD

Stochastic gradient descent replaces the gradient of the loss over the whole dataset with the gradient over one minibatch. That estimate is unbiased only if the minibatch is a random sample of the dataset. When consecutive batches come from one source, each step pushes the weights toward that source, the next block of a different source pushes them back, and the optimizer spends its updates oscillating. The adaptive moment estimates in Adam make this worse, because they adapt to the gradient statistics of the current block and then meet a different distribution.

There is also a recency effect. The learning rate schedule in most runs decays toward zero, so what the model sees in the final stretch is learned at small step sizes and is not overwritten afterwards. In an unshuffled run this is an accident; in an annealing curriculum it is the design. Either way, the end of the data stream has outsized influence, and you should decide what is there rather than inherit it from the file listing.

The shuffling levels

A map-style dataset, where any example can be fetched by index, can be shuffled perfectly: draw a random permutation of the indices for each epoch. Pretraining corpora rarely work like that. They are stored as many immutable shards and read sequentially for throughput, so a full permutation would turn reads into random seeks. The practical answer is a two-level shuffle: permute the shard order, read several shards concurrently, and pass documents through a shuffle buffer that emits a random element each time it is full.

Two-level shuffle with a step-indexed mixture scheduleImmutable shardsweb, code, books, curatedShard permutationseed + epochInterleave k shardsper-source streamsShuffle buffersize B documentsMixture schedule w(step)pure function, loggedsource weightsPack + batchrank r takes slice rTraining steploss by source, per batchCheckpoint cursorseed, epoch, shard, offsetsaveShuffle randomness comes only from (seed, epoch); the schedule depends only on step.Both are deterministic, so a resumed run reproduces the order it would have seen.Buffer size B bounds how far apart two examples can be and still land in the same batch.
Shard order and buffer draws are seeded from (seed, epoch); the mixture schedule is a function of the step. Neither depends on wall-clock or worker timing.

The buffer is the part people misjudge. A buffer of B documents can only mix documents that arrive within about B positions of each other. If a shard contains a million documents from one domain and the buffer holds ten thousand, the output is still a long run of that domain. Interleaving k shards before the buffer fixes this: each buffer fill draws from k different shards, so locality inside one shard no longer becomes locality in the batches. The packing step that turns documents into fixed-length sequences should come after the shuffle, so that one packed sequence does not concatenate neighbours from the same source file.

import random

def two_level_stream(shards, seed, epoch, k=8, buffer_size=10_000):
    """shards: list of shard readers, each yielding documents in stored order."""
    rng = random.Random(hash((seed, epoch)))   # same order on every rank
    order = shards[:]
    rng.shuffle(order)                         # level 1: shard order
    open_ = [iter(s) for s in order[:k]]       # read k shards at once
    pending = order[k:]
    buf = []
    while open_ or buf:
        if open_:
            i = rng.randrange(len(open_))
            try:
                buf.append(next(open_[i]))
            except StopIteration:
                open_.pop(i)
                if pending:
                    open_.append(iter(pending.pop()))
                continue
        if len(buf) >= buffer_size or (not open_ and buf):
            j = rng.randrange(len(buf))        # level 2: buffer shuffle
            buf[j], buf[-1] = buf[-1], buf[j]
            yield buf.pop()
Advertisement

Epochs and distributed ranks

Two mistakes are common when the same data is read more than once or by several processes. The first is reusing one permutation for every epoch. PyTorch's DistributedSampler derives its permutation from its seed and the epoch number, and its documentation says to call set_epoch at the start of each epoch; without it, every epoch has the same order. The Hugging Face streaming IterableDataset works the same way: shuffle(seed, buffer_size) sets up the shard and buffer shuffle, and set_epoch changes the effective seed per epoch.

The second mistake is giving each rank its own random seed for the permutation. Every rank must compute the same global order and then take a disjoint slice of it, otherwise ranks can see overlapping examples and some data is never read. Seed the permutation from the run seed and the epoch only, and let rank decide which slice to take.

# PyTorch map-style data: the sampler must be told the epoch
sampler = torch.utils.data.distributed.DistributedSampler(
    dataset, num_replicas=world_size, rank=rank, shuffle=True, seed=1234)
loader = torch.utils.data.DataLoader(dataset, batch_size=16, sampler=sampler)
for epoch in range(num_epochs):
    sampler.set_epoch(epoch)        # without this, every epoch repeats one order
    for batch in loader:
        ...

# Hugging Face streaming data: buffer shuffle plus epoch reseeding
ds = load_dataset("json", data_files=files, split="train", streaming=True)
ds = ds.shuffle(seed=1234, buffer_size=10_000)
for epoch in range(num_epochs):
    ds.set_epoch(epoch)             # effective seed becomes seed + epoch
    for example in ds:
        ...

Ordering in fine-tuning

Fine-tuning sets are small, often trained for several epochs, and assembled from sources that differ sharply: a few thousand tool-call examples, a larger chat set, a safety set. Concatenating them without a shuffle produces textbook catastrophic forgetting inside a single epoch: the model gets good at the tool-call format and then loses it during the chat block. Shuffle the combined set every epoch, and if one source is small but important, upsample it rather than placing it at the end.

Length grouping is the one deliberate departure from a full shuffle that many fine-tuning stacks offer; the Hugging Face Trainer calls it group_by_length. It batches examples of similar length together to waste fewer padding tokens, while keeping randomness at the level of groups. It is a throughput optimization with an ordering cost, since short and long examples now arrive in different batches. Packing, described in the SFT guide, often makes it unnecessary.

Curriculum learning: what the evidence supports

Curriculum learning, proposed for neural networks by Bengio and colleagues in 2009, orders training from easy to hard in the way people are taught. It is intuitive, and the evidence is mixed. A large study by Wu, Dyer and Neyshabur in 2021 found that with a standard training budget, easy-to-hard orderings gave at best marginal gains over random order, while curricula did help when the compute budget was cut short or the labels were noisy. Treat any curriculum as a hypothesis that needs an ablation, not as free accuracy.

Three forms have held up well enough to appear in published training recipes. Sequence-length warmup, studied by Li and colleagues in 2022, starts pretraining on short sequences and increases length over the early steps, which they reported improved stability at large batch sizes and learning rates. Quality annealing places a small amount of high-quality data at the end of pretraining while the learning rate decays: the Llama 3 report describes annealing on curated data, MiniCPM pairs its warmup-stable-decay schedule with higher-quality data in the decay phase, and OLMo 2 uses a separate mid-training stage on a curated mix. Mixture schedules change source weights over the run, for example raising code later once basic language modelling is in place.

What these share is that the ordering is coarse. None of them sorts individual examples by a difficulty score; they change the distribution in a few phases, and inside each phase the data is still shuffled.

Schedules as pure functions of the step

Implement every curriculum as a function of the global step, not of elapsed time or of which shards happen to be open. That makes the schedule easy to plot before the run, easy to log during it and trivially resumable. The sampler asks the schedule for the current weights, picks a source, and draws the next document from that source's shuffled stream.

def mixture_weights(step, total_steps):
    """Pure function of step: easy to log, test and resume."""
    anneal_start = int(0.9 * total_steps)
    base = {"web": 0.62, "code": 0.18, "books": 0.12, "curated": 0.08}
    if step < anneal_start:
        return base
    # final 10%: upweight curated data while the LR decays to zero
    return {"web": 0.35, "code": 0.20, "books": 0.10, "curated": 0.35}

def max_seq_len(step, warmup_steps=2_000, short=512, full=2_048):
    """Sequence-length warmup: start short, grow linearly to full context."""
    if step >= warmup_steps:
        return full
    return short + (full - short) * step // warmup_steps

def next_source(step, total_steps, rng):
    w = mixture_weights(step, total_steps)
    return rng.choices(list(w), weights=list(w.values()))[0]

Two rules keep schedules honest. Align phase boundaries with the learning-rate schedule, because an annealing phase only works if the learning rate is decaying through it. And account for repetition: if the curated source holds 1% of the tokens but is sampled at 8% for most of the run and 35% at the end, it will be seen many times. Check the implied number of epochs per source so that annealing data is not memorized.

Reproducibility and resume

A run that crashes and resumes must continue the same order. If the order depends on seed, epoch and step, the data cursor is small: seed, epoch, the shard permutation position, offsets within the open shards, the buffer contents or a way to rebuild them, and the random number generator state. The checkpoint contents that make the rest of the run resumable are covered in the SLM pretraining loop. Losing the data cursor is a common silent bug: the run resumes from the first shard, the first part of the data is seen twice and the end is never seen.

Buffer state is the awkward part. Saving ten thousand documents per rank is possible but bulky; the simpler approach is to make the buffer refill deterministic from the shard offsets and replay the last fill on resume, accepting a few thousand repeated documents. Whatever you choose, test it: kill a short run, resume it and compare the sequence of document IDs with an uninterrupted run.

Worked example: a 350M-parameter model on 20B tokens

A team trains a 350M-parameter model on 20 billion tokens from four sources, each read through its own two-level stream: filtered web text, code, books and a small curated set of textbook-style and instruction-like documents. Shards are 1 GB and each holds one source. The plan is a seeded shard permutation per epoch, eight shards open at once, a ten-thousand-document buffer per rank, packing to 2,048 tokens after the buffer, sequence-length warmup from 512 to 2,048 tokens over the first 2,000 steps, and a final 10% of steps on the annealing mixture while the learning rate decays to zero.

Before committing compute, they run two small ablations at 10% of the budget: one with the schedule and one with a fixed mixture, same seed and same total tokens per source. They compare held-out loss per source and a small downstream suite, and keep the anneal only if it helps on the evaluations they care about without hurting the rest. During the full run they log the source histogram of every batch and plot loss per source; the shuffle is working when the histogram is stable from batch to batch and the per-source losses fall smoothly rather than in steps.

Diagnostics

  • Per-batch source histogram: wide swings mean the shuffle is not mixing.
  • Loss broken down by source: a sawtooth that repeats with a period equal to a shard length means shards are being read in blocks.
  • Loss at epoch boundaries: a drop exactly at the boundary suggests the epoch reused the previous order.
  • Document-ID overlap between ranks for one epoch: should be zero.
  • Implied epochs per source under the schedule: anything well above a few for small sources is a memorization risk.

Failure modes and trade-offs

  • Buffer too small for the shard layout. The batches look random in a sample of ten but are clustered over thousands of steps.
  • Packing before shuffling. Each sequence holds consecutive documents from one source file.
  • Same order every epoch. The missing set_epoch call; fine for one epoch, harmful for many.
  • Unablated curriculum. A schedule copied from a paper at a different scale, never compared with random order.
  • Lost cursor on resume. The start of the data is repeated and the end is skipped.

The trade-offs are throughput against randomness and simplicity against control. Larger buffers and more open shards mix better but cost memory and file handles. Length grouping saves padding but correlates batches. A curriculum can help at small budgets but adds a schedule to tune and an ablation to pay for. For curated fine-tuning data, see the training data curation pipeline.

Key takeaway: <p><strong>What to do next.</strong> Make each batch a fair sample of the intended mixture with a two-level seeded shuffle, change that mixture only through a step-indexed schedule, and prove the order is reproducible across a resume. Treat curricula as hypotheses, and keep only the phases an ablation supports.</p><ol><li>Check your shard layout and set the number of interleaved shards and buffer size so no source dominates long runs of batches.</li><li>Seed the shuffle from run seed and epoch only, and call set_epoch every epoch.</li><li>Shuffle before packing, and shuffle combined fine-tuning sets every epoch.</li><li>Express any curriculum as mixture_weights(step) and plot it against the learning-rate schedule.</li><li>Compute implied epochs per source under that schedule.</li><li>Save the data cursor in checkpoints and run a kill-and-resume test.</li><li>Log per-batch source histograms and per-source loss, and run a small ablation before adopting a curriculum.</li></ol>