Every training run is a bet on a dozen numbers: learning rate, warmup, batch size, weight decay, dropout, schedule length. Get the learning rate wrong by a factor of three and a run that should converge diverges, or crawls. Hyperparameter search is how teams turn that bet into a measurement, and on GPUs it is expensive enough that the search strategy matters as much as the model.

This article compares the five strategies that cover nearly all practical use: grid search, random search, Bayesian optimization, early-stopping methods (successive halving, Hyperband and ASHA), and population-based training. For each it shows the algorithm as pseudocode, what it costs in GPU hours, and how it fails. It closes with working Optuna code, a note on transferring results from small proxy models, and guidance on choosing a method. The service architecture behind a tuning platform is covered in Hyperparameter Optimization Architecture.

Advertisement

What is actually worth tuning

Hyperparameters are not equally important. For training with Adam-family optimizers, the peak learning rate matters far more than anything else, and its effect is multiplicative, so it must be searched on a log scale. Next come the ones that interact with it: batch size, warmup length and the decay schedule. Weight decay, Adam's beta2, dropout and label smoothing usually matter less, and many architecture choices matter only when you are changing the model itself. For fine-tuning, the learning rate and the number of epochs dominate; for LoRA add rank and alpha.

Design the search space before choosing an algorithm. Use log-uniform sampling for anything that spans orders of magnitude (learning rate, weight decay), linear sampling for bounded fractions (warmup fraction, dropout), and categorical choices for discrete options. Make ranges wide enough that the optimum is inside them: if the best trials pile up at an edge, the range was wrong. Fix everything you do not intend to study, including the seed schedule and the evaluation set, so that differences between trials come from the hyperparameters rather than from the harness.

Grid search: exhaustive and wasteful

Grid search evaluates every combination of a fixed list of values per dimension. Its cost is the product of the list sizes: five values for each of four hyperparameters is 625 runs. Worse, when only one dimension matters, as the learning rate usually does, those 625 runs test just five distinct learning rates, and the remaining runs repeat them with irrelevant variations.

Grid search still suits one or two dimensions: it is reproducible and plots cleanly, as in a learning-rate sweep over a power-of-two ladder. Beyond two dimensions, prefer random search.

Advertisement

Random search: the baseline to beat

Random search samples each hyperparameter independently from its distribution. Bergstra and Bengio (JMLR, 2012) showed that because it tests a new value of every dimension in every trial, it covers the important dimensions far better than a grid of the same size. There is also a simple guarantee: if the top 5 percent of the space counts as good enough, the probability that at least one of n independent trials lands there is 1 - 0.95n. Sixty trials give about 95 percent, regardless of how many dimensions there are.

Random search is trivially parallel, since trials do not depend on each other, and it is robust to noisy metrics. That makes it the right default for a first search and the baseline every smarter method must beat.

# Random search: sample every dimension independently, on the right scale.
import math, random

SPACE = {
    "lr":           ("log",  1e-5, 1e-3),
    "warmup_frac":  ("lin",  0.0, 0.1),
    "weight_decay": ("log",  1e-3, 0.3),
    "beta2":        ("choice", [0.95, 0.98, 0.999]),
}

def sample(space, rng):
    cfg = {}
    for name, spec in space.items():
        kind = spec[0]
        if kind == "log":
            lo, hi = math.log(spec[1]), math.log(spec[2])
            cfg[name] = math.exp(rng.uniform(lo, hi))
        elif kind == "lin":
            cfg[name] = rng.uniform(spec[1], spec[2])
        else:
            cfg[name] = rng.choice(spec[1])
    return cfg

rng = random.Random(0)
trials = [sample(SPACE, rng) for _ in range(60)]  # 1 - 0.95**60 = 0.954

Bayesian optimization: learn where to look

Bayesian optimization fits a cheap surrogate model of the objective to the trials seen so far and uses an acquisition function to choose the next point, balancing exploitation (points predicted to be good) against exploration (points where the surrogate is uncertain). Classic implementations use a Gaussian process with expected improvement. The Tree-structured Parzen Estimator (TPE, Bergstra and colleagues, 2011), the default sampler in Optuna and the basis of Hyperopt, instead models two densities, one over good trials and one over the rest, and proposes points where their ratio is highest. TPE handles categorical and conditional parameters naturally and scales to many trials.

The catch is sequencing. A surrogate is useful only after it has seen results, so Bayesian methods start with random trials (often ten to twenty), and they gain most when trials are few and expensive. With 64 GPUs running at once, a sequential proposer becomes a bottleneck; implementations propose batches, for example by treating pending trials as if they had returned a guessed value (the constant liar heuristic). In practice, a Bayesian method on a handful of important dimensions usually matches random search's result in fewer trials, and the gap narrows as parallelism grows.

Early stopping: successive halving, Hyperband and ASHA

Most bad configurations are visibly bad early: their loss is higher after one epoch and stays higher. Successive halving exploits that. Start n configurations on a small budget, keep the best 1/eta, give survivors eta times the budget, and repeat until a few run to completion. With eta = 3 and budgets of 1, 3, 9 and 27 units (a unit might be 1,000 steps), 81 configurations become 27, then 9, then 3. When survivors resume from checkpoints, the total cost is 81 + 27 x 2 + 9 x 6 + 3 x 18 = 243 units, against 2,187 units to train all 81 to completion: about nine times cheaper for the same number of explored configurations.

The risk is killing a slow starter, such as a low learning rate that looks worse early but finishes better. Hyperband (Li and colleagues, JMLR 2018) hedges by running several successive-halving brackets with different starting budgets. ASHA (Li and colleagues, MLSys 2020) removes the synchronization barrier: instead of waiting for all 81 trials to finish rung 0, any free worker promotes a trial as soon as it ranks in the top 1/eta of the results that rung has seen so far, and otherwise starts a new configuration. GPUs never idle waiting for stragglers, which is why ASHA is the standard scheduler for large clusters.

Successive halving, eta = 3: keep the top third at every rungrung: 1 units81 trials run81 configs x 1 unitrung: 3 units27 trials runtop 27 resume to 3rung: 9 units9 trials runtop 9 resume to 9rung: 27 units3 trials runtop 3 train to full 27promote 1/3promote 1/3promote 1/3ASHA: promote asynchronouslya free GPU takes any promotable trial, else a new oneCost with checkpoint resume243 units versus 2,187 for 81 full runs
Successive halving with eta = 3. Each rung trains the survivors of the previous rung to three times the budget. ASHA makes promotion asynchronous, so a free GPU never waits for a whole rung to finish.
# ASHA (asynchronous successive halving), framework-agnostic pseudocode.
# rungs[k] holds (trial_id, metric) pairs that finished budget r_min * eta**k.

def get_job(rungs, eta, max_rung):
    # Look for a promotion from the highest rung downward.
    for k in reversed(range(max_rung)):
        done = sorted(rungs[k], key=lambda t: t.metric)        # lower is better
        top = done[: len(done) // eta]                          # top 1/eta so far
        for t in top:
            if not t.promoted:
                t.promoted = True
                return ("resume", t.trial_id, k + 1)          # train to next budget
    return ("new", sample_config(), 0)                        # nothing promotable

def worker_loop(gpu):
    while budget_left():
        kind, trial, k = get_job(rungs, eta=3, max_rung=3)
        state = load_checkpoint(trial) if kind == "resume" else init(trial)
        metric = train_until(state, units=r_min * 3**k, gpu=gpu)
        save_checkpoint(trial, state)
        rungs[k].append(Result(trial, metric))

Population-based training: tune while training

Population-based training (PBT, Jaderberg and colleagues, DeepMind, 2017) runs a population of models in parallel and periodically lets weak members copy strong ones. At each interval, the bottom fraction of the population loads the weights, optimizer state and hyperparameters of a member in the top fraction (exploit) and then perturbs the hyperparameters, for example multiplying the learning rate by 0.8 or 1.2 (explore). Training continues from the copied checkpoint, so no compute restarts from zero.

The result is not a single best configuration but a schedule: the learning rate a member followed over time, recorded in its lineage. That suits problems where the best value changes during training, such as reinforcement learning and GAN training, and it uses a fixed number of GPUs for the lifetime of the search. It is harder to reproduce, because the winning schedule is path-dependent, and it is greedy: an interval that is too short rewards hyperparameters that look good immediately, such as a low learning rate that drops the loss quickly but stalls later.

# Population-based training: exploit (copy a better member), then explore (perturb).
population = [Member(cfg=sample_config(), weights=init()) for _ in range(16)]

for interval in range(num_intervals):
    for m in population:                        # in parallel, one member per GPU
        train(m, steps=2000)
        m.score = evaluate(m)                   # held-out loss on a fixed eval set
    ranked = sorted(population, key=lambda m: m.score)     # lower is better
    quarter = len(ranked) // 4
    for loser in ranked[-quarter:]:
        winner = random.choice(ranked[:quarter])
        loser.weights = copy(winner.weights)    # also copy optimizer state and step
        loser.opt_state = copy(winner.opt_state)
        loser.cfg = dict(winner.cfg)
        loser.cfg["lr"] *= random.choice([0.8, 1.2])
        loser.history.append((interval, winner.id, loser.cfg["lr"]))

Running trials on GPUs

Tuning cost is set by how trials map to hardware. Small models waste a full GPU on one trial; run several per device with separate processes, or partition an A100 or H100 with MIG so each trial gets isolated memory and compute. Large models need several GPUs per trial, so the scheduler must allocate whole groups and parallelism falls accordingly. Checkpoints for resumption add storage and I/O: a 7B model with optimizer state is roughly 100 GB per checkpoint, and ASHA may keep hundreds, so keep only rung-boundary checkpoints for trials still eligible for promotion and delete the rest.

Evaluation must be cheap, fixed and consistent. A promotion decision is only as good as the metric behind it: use the same held-out set for every trial, big enough that the difference between trials exceeds the noise, and avoid metrics that depend on generation length or sampling temperature. Measure seed noise once by running the default configuration with three seeds; any difference smaller than that spread is not a real difference.

Working code with Optuna

Optuna combines a sampler, which proposes configurations, with a pruner, which stops unpromising ones. The example below pairs TPE with successive-halving pruning: each trial reports validation loss after every epoch, and the pruner stops it if it ranks poorly among trials at the same step. Shared storage lets many workers attach to one study, one script per GPU: SQLite for a few processes on one machine, MySQL or PostgreSQL across nodes. HyperbandPruner is a drop-in alternative when you are unsure of the minimum useful budget.

import optuna

def objective(trial):
    lr = trial.suggest_float("lr", 1e-5, 1e-3, log=True)
    wd = trial.suggest_float("weight_decay", 1e-3, 0.3, log=True)
    warmup = trial.suggest_float("warmup_frac", 0.0, 0.1)
    model, opt, sched = build(lr=lr, weight_decay=wd, warmup_frac=warmup)
    for epoch in range(27):
        train_one_epoch(model, opt, sched)
        val_loss = evaluate(model)
        trial.report(val_loss, step=epoch)
        if trial.should_prune():
            raise optuna.TrialPruned()
    return val_loss

study = optuna.create_study(
    direction="minimize",
    sampler=optuna.samplers.TPESampler(seed=0),
    pruner=optuna.pruners.SuccessiveHalvingPruner(min_resource=1, reduction_factor=3),
    storage="sqlite:///hpo.db", study_name="sft-7b-lr", load_if_exists=True,
)
study.optimize(objective, n_trials=80)
print(study.best_params)

Transferring results from small models

For large pretraining runs, no strategy can afford many full-size trials. The standard approach is to tune a small proxy and transfer. Under standard parameterization the optimal learning rate shifts with width, so a proxy's optimum is wrong for the target. The maximal update parameterization (muP; Yang and colleagues, 2022) rescales initialization and per-layer learning rates so that the optimum stays roughly stable as width grows, which makes tuning a small model meaningful for a large one; see muP and learning-rate transfer. Batch size also moves the optimum: larger batches tolerate higher learning rates up to a critical batch size, so retune when you change the global batch substantially. Schedules are covered in learning rate schedules.

Failure modes

  • Optimum at the edge of the range. The best trials cluster at a bound; widen the range and rerun rather than reporting the edge.
  • Overfitting the validation set. Hundreds of trials selected on one split overfit it; confirm the winner on a separate test split.
  • Pruning slow starters. Aggressive early stopping removes configurations with long warmups or low learning rates; set the minimum budget beyond warmup, or use Hyperband brackets.
  • Noise mistaken for signal. Differences smaller than seed variance are not findings; rerun the top three with new seeds.
  • Mismatched budgets. The best learning rate for 2,000 steps is usually higher than for 20,000 with the same schedule; tune at the schedule length you will use, or scale the schedule with it.
  • PBT collapse. The population converges to one lineage early and stops exploring; lengthen the interval and limit how many members may copy the same winner.
  • Leaked checkpoints. Resume-based schedulers fill disks; garbage-collect checkpoints of trials that can no longer be promoted.

Choosing a method

With a few dimensions and plenty of parallel GPUs, run random search with ASHA pruning: it is simple, robust and keeps the cluster busy. With expensive trials and limited parallelism, use TPE or a Gaussian-process method with a pruner, and seed it with a few random trials. When the best setting changes over training, or you have a fixed pool of GPUs for a long run, use PBT and keep its lineage so the schedule can be replayed. For multi-billion-parameter pretraining, tune a muP proxy and transfer. Use grid search only for one- or two-dimensional sweeps you intend to plot.

Whatever the method, budget the search explicitly, for example as a fixed fraction of the final training run's GPU hours, and stop when the best result has not improved over the last quarter of trials. The largest savings rarely come from a cleverer optimizer: they come from a tight space, a cheap reliable metric and early stopping.

Key takeaway: Tune the learning rate first and on a log scale, fix everything you are not studying, and measure seed noise before trusting any difference. Random search is the baseline, Bayesian methods such as TPE save trials when trials are expensive and few, successive halving and ASHA stop bad runs early and cut the cost of exploring 81 configurations about ninefold in the worked example, and PBT turns the search into a learned schedule. On GPUs, pack small trials, garbage-collect checkpoints, and for very large models tune a muP proxy and transfer rather than searching at full scale.