Hyperparameter optimization (HPO) is usually taught as an algorithm question: grid or random, Bayesian or bandit. In production it is mostly a systems question. A tuning run is a fleet of training jobs that share one search state, compete for GPUs, crash at awkward moments and feed results to a sampler that must choose the next configuration while half the fleet is still running. Get the plumbing wrong and you spend the budget re-running dead trials, handing near-identical configurations to idle workers, or crowning a winner that was only lucky.

This article is about that plumbing. The companion page on tuning methods in practice covers grid, random, Bayesian optimization, ASHA and population-based training. Here those methods are pluggable parts, and the subject is the architecture around them: study and trial records, the ask/tell interface, the trial state machine, parallel workers, pruning as a protocol, surviving worker death, and keeping tuning away from your test set. Code uses Optuna, but every idea carries over to Ray Tune, Katib or your own controller.

Advertisement

The objects every HPO system has

Strip any tuning framework down and you find the same six objects. The search space is a function that turns a trial handle into a concrete configuration: which parameters exist, their ranges, their scales, and which ones only exist when another has a certain value. A study is one optimization run over one space toward one objective. A trial is one evaluation of one configuration, with a state, its parameters, any intermediate values it reported and its final value. The sampler reads the history of trials and proposes parameters for the next one. The pruner reads intermediate values and decides whether a running trial should stop early. Storage holds all of it durably so that many processes on many machines see one consistent history.

The design rule that makes this work is that workers are stateless. A worker asks storage for a trial, trains, reports and tells storage the result. That is what lets you add workers mid-run and lose a node without losing the study.

An HPO system: one shared study record, many stateless workers, pluggable sampler and prunerSearch spaceversioned codeSamplerask: next paramsPrunerstop or continueStudy storagetrials, params, values, heartbeatsWorker 1train, reportWorker 2train, reportWorker Ntrain, reportGPU schedulerleases, preemptionRequeue jobstale trialsModel registrywinning checkpointdefinesreads all trialsreads curvesreport, tellenqueuebest trialWorkers hold no search state. Everything the sampler and pruner know comes from the storage record.
Samplers and pruners are pure functions of the trial history in storage. Workers, the requeue job and the registry all talk to the same record.

The trial state machine

A trial moves through a small set of states, and most operational bugs are a trial stuck in the wrong one. In Optuna the states are WAITING (enqueued with fixed parameters, not started), RUNNING, COMPLETE, PRUNED and FAIL. Other frameworks use different names for the same ideas.

StateEntered whenWhat the sampler does with it
WAITINGYou enqueue a configuration by hand or a requeue job doesNothing yet; the next ask picks it up before sampling new parameters
RUNNINGA worker asks and starts trainingTreats it as pending; with constant liar it pretends a bad value so others avoid the region
COMPLETEThe objective returned a valueFull evidence for the model of the objective
PRUNEDThe pruner stopped itPartial evidence; samplers such as TPE can use the last reported value
FAILAn exception, or a stale heartbeatIgnored; a failed trial teaches the sampler nothing

Two transitions need care. A worker that is killed by preemption or an out-of-memory kill never reports, so its trial stays RUNNING forever unless something notices. And FAIL conflates very different causes: a preempted node is worth retrying, while a configuration that diverges to NaN will fail again on every retry. Record the cause as a trial attribute so later code can tell them apart.

Advertisement

Ask and tell: decoupling the search from the training

The cleanest interface between a search algorithm and the rest of the system is ask/tell. ask() creates a trial and returns a handle you can draw parameters from; tell() records the outcome. Nothing else is shared. Optuna's study.optimize(objective) is a convenience loop around these two calls. When each trial is its own cluster job, call them directly from a controller, as below. The training jobs then need only the trial number and the storage URL, so they can report intermediate values and check for pruning on their own.

# controller.py: ask/tell when trials run as cluster jobs instead of in-process
study = optuna.load_study(study_name="resnet50-cls-v3", storage=STORAGE)
inflight = {}
while budget.remaining_gpu_hours() > 0 or inflight:
    while budget.remaining_gpu_hours() > 0 and scheduler.free_gpus() and len(inflight) < MAX_PARALLEL:
        trial = study.ask()                       # the sampler sees RUNNING trials too
        params = search_space(trial)
        inflight[trial.number] = scheduler.launch(image=IMAGE, params=params, trial=trial.number)
    for number, job in scheduler.poll(inflight):  # jobs report intermediate values to STORAGE
        if job.status == "succeeded":
            study.tell(number, job.final_metric)
        elif job.status == "pruned":
            study.tell(number, state=optuna.trial.TrialState.PRUNED)
        else:
            study.tell(number, state=optuna.trial.TrialState.FAIL)
            job_log.record(number, job.exit_reason)  # preempted, OOM, NaN loss, bad config
        del inflight[number]

The controller owns the parallelism limit, the GPU budget and the mapping from exit reasons to trial states, but no memory of past results. If it crashes, a new instance loads the study and carries on.

Parallel suggestion and the constant liar

Sequential Bayesian optimization assumes one trial at a time: observe, update, propose. With 16 workers, 15 trials are in flight whenever one more is requested. If the sampler ignores them, it proposes nearly the same point it proposed to the last idle worker, because its model has not changed. You pay for 16 GPUs and learn about one region.

The common fix is the constant liar heuristic: when proposing, treat every running trial as if it had already returned a poor value, so the model is pushed away from regions already being explored. Optuna's TPESampler has a constant_liar parameter for this; recent releases default it to on, so set it explicitly and your behavior will not change with a library upgrade. It also keeps a trial that died without reporting from attracting further suggestions.

Parallelism still has a cost: each proposal is made with less information than a sequential run would have had. Keep concurrent trials well below the total budget. Sixteen workers for a 400-trial study is reasonable; 128 workers for a 200-trial study is random search with overhead. Spend surplus GPUs on more seeds per configuration or a wider ASHA bottom rung instead.

Pruning is a reporting protocol

Pruning only works if every trial reports comparable numbers at comparable steps. The worker calls trial.report(value, step) after each evaluation and trial.should_prune() right after; the pruner compares this trial's value at this step with other trials' values at the same step. Three things break that comparison silently.

  • Steps in different units. If batch size is tuned and trials report per optimizer step, step 5 means different amounts of training. Report epochs or examples seen.
  • Noisy intermediate metrics. A score on 500 validation examples swings between evaluations, and the pruner kills good trials on a bad draw. Use a larger fixed slice or a moving average.
  • Slow starters. Long warmup makes some configurations look worse early and better late. Put the first rung after warmup, via min_resource in Optuna's SuccessiveHalvingPruner.

The worker below shows the whole protocol, plus the attributes that make results reproducible.

# worker.py: start as many copies as you have GPUs; they coordinate only through STORAGE
import optuna

STORAGE = optuna.storages.RDBStorage(
    url="postgresql://hpo:secret@hpo-db:5432/hpo",
    heartbeat_interval=60,   # this worker writes a heartbeat for its running trial every 60 s
    grace_period=180,        # no heartbeat for 180 s means the trial is stale
)
SPACE_VERSION = "v3"         # bump on any change to search_space(); new version, new study

def search_space(trial):
    return {
        "lr": trial.suggest_float("lr", 1e-5, 3e-3, log=True),
        "weight_decay": trial.suggest_float("weight_decay", 1e-6, 1e-1, log=True),
        "warmup_frac": trial.suggest_float("warmup_frac", 0.0, 0.1),
        "batch_size": trial.suggest_categorical("batch_size", [64, 128, 256]),
    }

def objective(trial):
    cfg = search_space(trial)
    trial.set_user_attr("space_version", SPACE_VERSION)
    trial.set_user_attr("git_sha", GIT_SHA)
    trial.set_user_attr("data_snapshot", DATA_SNAPSHOT)
    model, opt, start = build_or_resume(cfg, ckpt_key=params_hash(cfg))
    for epoch in range(start, EPOCHS):
        train_one_epoch(model, opt)
        val = evaluate(model, split="val")        # the test split is never touched here
        trial.report(val, step=epoch)             # step = epochs completed, same unit for all trials
        if trial.should_prune():
            raise optuna.TrialPruned()
        save_checkpoint(params_hash(cfg), epoch)  # a retry of the same params resumes here
    return val

study = optuna.create_study(
    study_name=f"resnet50-cls-{SPACE_VERSION}",
    storage=STORAGE,
    direction="maximize",
    sampler=optuna.samplers.TPESampler(n_startup_trials=20, constant_liar=True),
    pruner=optuna.pruners.SuccessiveHalvingPruner(min_resource=3, reduction_factor=3),
    load_if_exists=True,                          # every worker attaches to the same study
)
study.optimize(objective, timeout=48 * 3600)      # wall-clock budget per worker

Storage, heartbeats and dead trials

A shared relational database is the usual storage for multi-machine studies. Every ask reads the trial history and every report writes a row, so give it connection pooling and keep one study per experiment so history reads stay short. Optuna also offers file-based journal storage, whose class names changed during the 4.x series; check the docs for your installed version.

Heartbeats are how dead workers are noticed. With heartbeat_interval set, a running trial's timestamp is refreshed in the background; once it is older than grace_period the trial is stale and Optuna moves it to FAIL. Optuna's automatic-retry callback is being renamed across recent releases, so check your version's RDBStorage reference. An explicit requeue job is easier to reason about, because it decides which failures deserve a retry:

# requeue.py: run from one place on a schedule, never from every worker
from optuna.trial import TrialState

study = optuna.load_study(study_name="resnet50-cls-v3", storage=STORAGE)
optuna.storages.fail_stale_trials(study)          # experimental API: expired heartbeat -> FAIL
trials = study.get_trials(deepcopy=False)
already = {t.user_attrs.get("retry_of") for t in trials}
for t in trials:
    if t.state != TrialState.FAIL or t.number in already:
        continue
    if failure_kind(t) in {"config_error", "nan_loss", "oom"}:   # deterministic: retrying wastes GPUs
        continue
    if t.user_attrs.get("retry_of") is not None:                 # one retry only
        continue
    study.enqueue_trial(t.params, user_attrs={"retry_of": t.number})

Requeueing with the same parameters only saves compute if the training code can resume. Key checkpoints by a hash of the parameters, not by trial number, so that the retry, which gets a new number, finds the checkpoint the dead trial wrote.

The search space is versioned code

A study's history is only meaningful under the space that produced it. Widen a range halfway through and the sampler mixes trials from two distributions; rename a parameter and old trials stop informing new ones. Version the space, store the version on every trial, and start a new study when it changes, warm-started by enqueueing the best few old configurations.

Put scale parameters such as learning rate on a log scale, fix anything you have no reason to tune, and for large models tune a small proxy and transfer: hyperparameter scaling transfer explains when that works.

Worked example: budget arithmetic for 8 GPUs and 48 hours

Say a full trial is 27 epochs and an epoch takes 8 minutes on one GPU, so a full trial is 3.6 GPU-hours. Eight GPUs for 48 hours is 384 GPU-hours, which buys about 106 full trials with no pruning, or 2,880 epochs.

Now add successive halving with rungs at 3, 9 and 27 epochs and a reduction factor of 3. Every started trial runs 3 epochs; one in three runs 6 more; one in nine runs the last 18. Expected cost per started trial is 3 + 6/3 + 18/9 = 7 epochs, against 27. The same 2,880 epochs now start about 411 trials, of which about 45 finish: roughly four times the exploration, provided early rankings predict late ones. Check that once: run 20 random configurations to the end and see whether the top third at epoch 3 contains the eventual best.

Then reserve budget the pruner never sees: re-run the top five configurations with three fresh seeds each, 54 GPU-hours here, before promoting anything. The best of hundreds of noisy trials is partly luck.

Keeping tuning honest

Every trial is a peek at the validation set, and hundreds of peeks overfit it. Tune only on validation. Touch the test split once, for the chosen configuration, and never feed its score into another study. With scarce data, use nested cross-validation. Record the data snapshot on each trial, because a study run on two data versions is two studies. The MLOps architecture page covers the lineage records that make this queryable.

Failure modes

  • Zombie trials. Preempted workers leave trials in RUNNING; the study looks busy and, without constant liar, the region around them is never revisited. Fix with heartbeats and a requeue job.
  • Retry storms. A configuration that runs out of memory is retried forever. Classify failures and retry only transient ones, once.
  • Duplicate suggestions. Many workers, no constant liar, so the GPUs explore one point. Visible as clusters of near-identical parameters with overlapping start times.
  • Silent space drift. Someone edits the space in place; old and new trials mix. Version the space and refuse to attach to a study whose version differs.
  • Lucky winners. The top trial does not reproduce. Re-run top-k with fresh seeds before promotion.

Trade-offs and build or buy

ChoiceGainCost
In-process optimize loopLeast code; pruning just worksTrial lifetime tied to the worker process
Ask/tell controller with cluster jobsIsolation, retries, per-trial resourcesYou own scheduling and exit-reason mapping
Aggressive pruningSeveral times more configurations per GPU-hourKills late bloomers; needs a rank-correlation check
High parallelismShorter wall-clock timeLess informed proposals per trial

Export the study history with each model release, next to the config the training pipeline records, so the next tuning run starts from evidence.

What to do next

  1. Write the search space as one versioned function and store its version, the git SHA and the data snapshot on every trial.
  2. Move to shared storage with heartbeats on, and add a requeue job that retries transient failures once.
  3. Set constant_liar explicitly and keep concurrent trials well below the total trial budget.
  4. Report intermediate values in a unit that is fixed across configurations, and put the first pruning rung after warmup.
  5. Run the 20-configuration rank-correlation check once before trusting early stopping on a new problem.
  6. Reserve budget to re-run the top five configurations with fresh seeds before anything is promoted.
  7. Remove test-split evaluation from the objective; compute it once, for the chosen configuration only.
Key takeaway: An HPO system is a shared trial history with stateless workers around it. Samplers and pruners are swappable functions of that history; what makes or breaks a tuning run is the architecture: an ask/tell boundary, a trial state machine whose failures are classified, heartbeats and a requeue job for dead workers, constant liar for parallel proposals, a reporting unit that keeps pruning comparisons fair, a versioned search space, and a hard wall between the validation split you tune on and the test split you report. Budget the run with rung arithmetic, and keep some compute aside to re-run the winners with fresh seeds.