Hyperparameter optimization (HPO) is usually taught as an algorithm question: grid or random, Bayesian or bandit. In production it is mostly a systems question. A tuning run is a fleet of training jobs that share one search state, compete for GPUs, crash at awkward moments and feed results to a sampler that must choose the next configuration while half the fleet is still running. Get the plumbing wrong and you spend the budget re-running dead trials, handing near-identical configurations to idle workers, or crowning a winner that was only lucky.
This article is about that plumbing. The companion page on tuning methods in practice covers grid, random, Bayesian optimization, ASHA and population-based training. Here those methods are pluggable parts, and the subject is the architecture around them: study and trial records, the ask/tell interface, the trial state machine, parallel workers, pruning as a protocol, surviving worker death, and keeping tuning away from your test set. Code uses Optuna, but every idea carries over to Ray Tune, Katib or your own controller.
The objects every HPO system has
Strip any tuning framework down and you find the same six objects. The search space is a function that turns a trial handle into a concrete configuration: which parameters exist, their ranges, their scales, and which ones only exist when another has a certain value. A study is one optimization run over one space toward one objective. A trial is one evaluation of one configuration, with a state, its parameters, any intermediate values it reported and its final value. The sampler reads the history of trials and proposes parameters for the next one. The pruner reads intermediate values and decides whether a running trial should stop early. Storage holds all of it durably so that many processes on many machines see one consistent history.
The design rule that makes this work is that workers are stateless. A worker asks storage for a trial, trains, reports and tells storage the result. That is what lets you add workers mid-run and lose a node without losing the study.
The trial state machine
A trial moves through a small set of states, and most operational bugs are a trial stuck in the wrong one. In Optuna the states are WAITING (enqueued with fixed parameters, not started), RUNNING, COMPLETE, PRUNED and FAIL. Other frameworks use different names for the same ideas.
| State | Entered when | What the sampler does with it |
|---|---|---|
WAITING | You enqueue a configuration by hand or a requeue job does | Nothing yet; the next ask picks it up before sampling new parameters |
RUNNING | A worker asks and starts training | Treats it as pending; with constant liar it pretends a bad value so others avoid the region |
COMPLETE | The objective returned a value | Full evidence for the model of the objective |
PRUNED | The pruner stopped it | Partial evidence; samplers such as TPE can use the last reported value |
FAIL | An exception, or a stale heartbeat | Ignored; a failed trial teaches the sampler nothing |
Two transitions need care. A worker that is killed by preemption or an out-of-memory kill never reports, so its trial stays RUNNING forever unless something notices. And FAIL conflates very different causes: a preempted node is worth retrying, while a configuration that diverges to NaN will fail again on every retry. Record the cause as a trial attribute so later code can tell them apart.
Ask and tell: decoupling the search from the training
The cleanest interface between a search algorithm and the rest of the system is ask/tell. ask() creates a trial and returns a handle you can draw parameters from; tell() records the outcome. Nothing else is shared. Optuna's study.optimize(objective) is a convenience loop around these two calls. When each trial is its own cluster job, call them directly from a controller, as below. The training jobs then need only the trial number and the storage URL, so they can report intermediate values and check for pruning on their own.
# controller.py: ask/tell when trials run as cluster jobs instead of in-process
study = optuna.load_study(study_name="resnet50-cls-v3", storage=STORAGE)
inflight = {}
while budget.remaining_gpu_hours() > 0 or inflight:
while budget.remaining_gpu_hours() > 0 and scheduler.free_gpus() and len(inflight) < MAX_PARALLEL:
trial = study.ask() # the sampler sees RUNNING trials too
params = search_space(trial)
inflight[trial.number] = scheduler.launch(image=IMAGE, params=params, trial=trial.number)
for number, job in scheduler.poll(inflight): # jobs report intermediate values to STORAGE
if job.status == "succeeded":
study.tell(number, job.final_metric)
elif job.status == "pruned":
study.tell(number, state=optuna.trial.TrialState.PRUNED)
else:
study.tell(number, state=optuna.trial.TrialState.FAIL)
job_log.record(number, job.exit_reason) # preempted, OOM, NaN loss, bad config
del inflight[number]The controller owns the parallelism limit, the GPU budget and the mapping from exit reasons to trial states, but no memory of past results. If it crashes, a new instance loads the study and carries on.
Parallel suggestion and the constant liar
Sequential Bayesian optimization assumes one trial at a time: observe, update, propose. With 16 workers, 15 trials are in flight whenever one more is requested. If the sampler ignores them, it proposes nearly the same point it proposed to the last idle worker, because its model has not changed. You pay for 16 GPUs and learn about one region.
The common fix is the constant liar heuristic: when proposing, treat every running trial as if it had already returned a poor value, so the model is pushed away from regions already being explored. Optuna's TPESampler has a constant_liar parameter for this; recent releases default it to on, so set it explicitly and your behavior will not change with a library upgrade. It also keeps a trial that died without reporting from attracting further suggestions.
Parallelism still has a cost: each proposal is made with less information than a sequential run would have had. Keep concurrent trials well below the total budget. Sixteen workers for a 400-trial study is reasonable; 128 workers for a 200-trial study is random search with overhead. Spend surplus GPUs on more seeds per configuration or a wider ASHA bottom rung instead.
Pruning is a reporting protocol
Pruning only works if every trial reports comparable numbers at comparable steps. The worker calls trial.report(value, step) after each evaluation and trial.should_prune() right after; the pruner compares this trial's value at this step with other trials' values at the same step. Three things break that comparison silently.
- Steps in different units. If batch size is tuned and trials report per optimizer step, step 5 means different amounts of training. Report epochs or examples seen.
- Noisy intermediate metrics. A score on 500 validation examples swings between evaluations, and the pruner kills good trials on a bad draw. Use a larger fixed slice or a moving average.
- Slow starters. Long warmup makes some configurations look worse early and better late. Put the first rung after warmup, via
min_resourcein Optuna'sSuccessiveHalvingPruner.
The worker below shows the whole protocol, plus the attributes that make results reproducible.
# worker.py: start as many copies as you have GPUs; they coordinate only through STORAGE
import optuna
STORAGE = optuna.storages.RDBStorage(
url="postgresql://hpo:secret@hpo-db:5432/hpo",
heartbeat_interval=60, # this worker writes a heartbeat for its running trial every 60 s
grace_period=180, # no heartbeat for 180 s means the trial is stale
)
SPACE_VERSION = "v3" # bump on any change to search_space(); new version, new study
def search_space(trial):
return {
"lr": trial.suggest_float("lr", 1e-5, 3e-3, log=True),
"weight_decay": trial.suggest_float("weight_decay", 1e-6, 1e-1, log=True),
"warmup_frac": trial.suggest_float("warmup_frac", 0.0, 0.1),
"batch_size": trial.suggest_categorical("batch_size", [64, 128, 256]),
}
def objective(trial):
cfg = search_space(trial)
trial.set_user_attr("space_version", SPACE_VERSION)
trial.set_user_attr("git_sha", GIT_SHA)
trial.set_user_attr("data_snapshot", DATA_SNAPSHOT)
model, opt, start = build_or_resume(cfg, ckpt_key=params_hash(cfg))
for epoch in range(start, EPOCHS):
train_one_epoch(model, opt)
val = evaluate(model, split="val") # the test split is never touched here
trial.report(val, step=epoch) # step = epochs completed, same unit for all trials
if trial.should_prune():
raise optuna.TrialPruned()
save_checkpoint(params_hash(cfg), epoch) # a retry of the same params resumes here
return val
study = optuna.create_study(
study_name=f"resnet50-cls-{SPACE_VERSION}",
storage=STORAGE,
direction="maximize",
sampler=optuna.samplers.TPESampler(n_startup_trials=20, constant_liar=True),
pruner=optuna.pruners.SuccessiveHalvingPruner(min_resource=3, reduction_factor=3),
load_if_exists=True, # every worker attaches to the same study
)
study.optimize(objective, timeout=48 * 3600) # wall-clock budget per worker
Storage, heartbeats and dead trials
A shared relational database is the usual storage for multi-machine studies. Every ask reads the trial history and every report writes a row, so give it connection pooling and keep one study per experiment so history reads stay short. Optuna also offers file-based journal storage, whose class names changed during the 4.x series; check the docs for your installed version.
Heartbeats are how dead workers are noticed. With heartbeat_interval set, a running trial's timestamp is refreshed in the background; once it is older than grace_period the trial is stale and Optuna moves it to FAIL. Optuna's automatic-retry callback is being renamed across recent releases, so check your version's RDBStorage reference. An explicit requeue job is easier to reason about, because it decides which failures deserve a retry:
# requeue.py: run from one place on a schedule, never from every worker
from optuna.trial import TrialState
study = optuna.load_study(study_name="resnet50-cls-v3", storage=STORAGE)
optuna.storages.fail_stale_trials(study) # experimental API: expired heartbeat -> FAIL
trials = study.get_trials(deepcopy=False)
already = {t.user_attrs.get("retry_of") for t in trials}
for t in trials:
if t.state != TrialState.FAIL or t.number in already:
continue
if failure_kind(t) in {"config_error", "nan_loss", "oom"}: # deterministic: retrying wastes GPUs
continue
if t.user_attrs.get("retry_of") is not None: # one retry only
continue
study.enqueue_trial(t.params, user_attrs={"retry_of": t.number})Requeueing with the same parameters only saves compute if the training code can resume. Key checkpoints by a hash of the parameters, not by trial number, so that the retry, which gets a new number, finds the checkpoint the dead trial wrote.
The search space is versioned code
A study's history is only meaningful under the space that produced it. Widen a range halfway through and the sampler mixes trials from two distributions; rename a parameter and old trials stop informing new ones. Version the space, store the version on every trial, and start a new study when it changes, warm-started by enqueueing the best few old configurations.
Put scale parameters such as learning rate on a log scale, fix anything you have no reason to tune, and for large models tune a small proxy and transfer: hyperparameter scaling transfer explains when that works.
Worked example: budget arithmetic for 8 GPUs and 48 hours
Say a full trial is 27 epochs and an epoch takes 8 minutes on one GPU, so a full trial is 3.6 GPU-hours. Eight GPUs for 48 hours is 384 GPU-hours, which buys about 106 full trials with no pruning, or 2,880 epochs.
Now add successive halving with rungs at 3, 9 and 27 epochs and a reduction factor of 3. Every started trial runs 3 epochs; one in three runs 6 more; one in nine runs the last 18. Expected cost per started trial is 3 + 6/3 + 18/9 = 7 epochs, against 27. The same 2,880 epochs now start about 411 trials, of which about 45 finish: roughly four times the exploration, provided early rankings predict late ones. Check that once: run 20 random configurations to the end and see whether the top third at epoch 3 contains the eventual best.
Then reserve budget the pruner never sees: re-run the top five configurations with three fresh seeds each, 54 GPU-hours here, before promoting anything. The best of hundreds of noisy trials is partly luck.
Keeping tuning honest
Every trial is a peek at the validation set, and hundreds of peeks overfit it. Tune only on validation. Touch the test split once, for the chosen configuration, and never feed its score into another study. With scarce data, use nested cross-validation. Record the data snapshot on each trial, because a study run on two data versions is two studies. The MLOps architecture page covers the lineage records that make this queryable.
Failure modes
- Zombie trials. Preempted workers leave trials in
RUNNING; the study looks busy and, without constant liar, the region around them is never revisited. Fix with heartbeats and a requeue job. - Retry storms. A configuration that runs out of memory is retried forever. Classify failures and retry only transient ones, once.
- Duplicate suggestions. Many workers, no constant liar, so the GPUs explore one point. Visible as clusters of near-identical parameters with overlapping start times.
- Silent space drift. Someone edits the space in place; old and new trials mix. Version the space and refuse to attach to a study whose version differs.
- Lucky winners. The top trial does not reproduce. Re-run top-k with fresh seeds before promotion.
Trade-offs and build or buy
| Choice | Gain | Cost |
|---|---|---|
| In-process optimize loop | Least code; pruning just works | Trial lifetime tied to the worker process |
| Ask/tell controller with cluster jobs | Isolation, retries, per-trial resources | You own scheduling and exit-reason mapping |
| Aggressive pruning | Several times more configurations per GPU-hour | Kills late bloomers; needs a rank-correlation check |
| High parallelism | Shorter wall-clock time | Less informed proposals per trial |
Export the study history with each model release, next to the config the training pipeline records, so the next tuning run starts from evidence.
What to do next
- Write the search space as one versioned function and store its version, the git SHA and the data snapshot on every trial.
- Move to shared storage with heartbeats on, and add a requeue job that retries transient failures once.
- Set
constant_liarexplicitly and keep concurrent trials well below the total trial budget. - Report intermediate values in a unit that is fixed across configurations, and put the first pruning rung after warmup.
- Run the 20-configuration rank-correlation check once before trusting early stopping on a new problem.
- Reserve budget to re-run the top five configurations with fresh seeds before anything is promoted.
- Remove test-split evaluation from the objective; compute it once, for the chosen configuration only.