A release train is a release process with a fixed timetable: the train departs on schedule, changes that are ready board it, and changes that are not ready wait for the next one. The schedule never moves to wait for a feature. Large software organisations adopted trains to replace release dates negotiated around the slowest change, and the same idea fits agents built with the Agent Development Kit for Java unusually well, because an agent's behaviour changes through many small, interacting edits, prompts, tools, model identifiers, library versions and evaluation sets, whose effects you only see when you evaluate them together.

This page designs a weekly train for an ADK Java agent: the timetable, the cut that freezes a manifest, what kinds of change may board and together with what, the gates the frozen candidate must pass, the hotfix lane, and what happens to a change that misses the train or fails on it. It covers process; for the mechanics it relies on, fingerprinting, canaries and flags, it links to the pages that own them.

Why agents suit a train

Agent changes differ from ordinary code changes in three ways that make a timetable valuable. Their effects are statistical: a one-line instruction edit can move task success by several points, and only a full evaluation run, which takes time and model quota, shows it. Their effects interact: a new tool and an instruction that mentions it are only meaningful together, while a model upgrade can change how every prompt is read. And their cost to evaluate is high enough that running the whole suite on every merge is often unaffordable, while running it once on a frozen candidate is cheap.

Continuous deployment answers these with per-change gates, which work when each change is small and the gates are fast. A train batches changes so the expensive evaluation runs once per departure on exactly what will ship, gives every stakeholder a predictable date, and turns schedule pressure into a simple rule: if your change is not ready, it takes the next train. Teams shipping a single agent with fast evals may not need one; teams with several contributors to one agent usually do.

A weekly agent release train: fixed departure, cars that can be dropped, a separate hotfix laneMonTueWedThuFriMonBoardingmerge to mainCut 12:00 Tuemanifest frozenGateseval suite, soakCanary Thu5% sessionsDeparts Fri100% or holdCars on train 2026.41 (each independently revertible)prompt carinstruction difftool carschema + codeeval carnew cases, thresholdsmodel cardropped: failed evalHotfix laneseverity-gated, one change, full gates, same dayMissed the cut?wait for next train; never delay this oneThe train leaves on time with whatever passed. Cars that fail are removed, not fixed in place.An ADK library upgrade and a model change never ride the same train.
A week on the train: changes board until the cut, the frozen manifest goes through gates and a canary, and the train departs on Friday with whatever passed.

The timetable and the cut

Pick a cadence your evaluation can keep up with. Weekly suits most teams: boarding from Monday, a cut at noon on Tuesday, gates on Tuesday and Wednesday, a canary from Thursday and full traffic on Friday morning (or the following Monday, if nobody watches production at weekends). The cut is the key moment. At the cut, a script reads the head of the release branch, records the exact commit, resolves every moving reference into a fixed one, and writes the manifest. Nothing that changes after the cut affects this train.

# release/train-2026.41.yaml  (generated at the cut, then read-only)
train: "2026.41"
cut_commit: 3f1c9e2
departs: "2026-10-16T09:00:00Z"
agent:
  name: support_agent
  model: "<exact versioned model id>"   # pinned; never an alias that moves
  instruction_file: prompts/support_agent.md
  instruction_sha256: 9b0e...       # fingerprint of the exact text
  tools: [lookupOrder, issueRefund, escalate]
  generate_config: { temperature: 0.2, max_output_tokens: 1024 }
libraries:
  google-adk: "<version from the lock file>"
eval:
  suite: evals/support_v14
  thresholds: { task_success: 0.92, policy_violations: 0, p95_turn_ms: 6000 }
cars:
  - { id: prompt-refund-tone, kind: prompt,  owner: ana,  pr: 812 }
  - { id: tool-escalate-v2,  kind: tool,    owner: raj,  pr: 815 }
  - { id: eval-v14-cases,    kind: eval,    owner: mei,  pr: 817 }
  - { id: model-flash-bump,  kind: model,   owner: tom,  pr: 820, status: dropped }

Resolving moving references is the part teams skip. A model alias that the provider repoints, an instruction loaded from a configuration service, or a library version range in the build file each lets the shipped agent differ from the tested one. The manifest pins the model identifier, fingerprints the instruction text, and copies library versions from the lock file. The service then builds the agent only from the manifest and refuses to start if anything has drifted.

/** Builds the agent only from a frozen manifest, so what was tested is what ships. */
public final class AgentFactory {
  public static LlmAgent fromManifest(Manifest m, ToolRegistry registry) throws IOException {
    String instruction = Files.readString(Path.of(m.agent().instructionFile()));
    String actual = sha256(instruction);
    if (!actual.equals(m.agent().instructionSha256())) {
      throw new IllegalStateException("instruction drifted from train " + m.train()
          + ": expected " + m.agent().instructionSha256() + ", found " + actual);
    }
    return LlmAgent.builder()
        .name(m.agent().name())
        .model(m.agent().model())
        .instruction(instruction)
        .tools(registry.resolve(m.agent().tools()))          // fails on unknown tool names
        .generateContentConfig(m.agent().generateConfig().toGenai())
        .build();
  }
}

Cars and boarding rules

Each change boards as a car with a kind, an owner and a link to its review. Cars are what make the train robust: when one fails, it is reverted and the rest still leave. That only works if cars are independent enough to remove, so the boarding rules exist mainly to keep them so.

Car kindExamplesBoarding rule
promptInstruction text, few-shot examplesAny number; each needs eval cases that exercise it
toolNew tool, schema or behaviour changeRides with the prompt change that uses it, as one car
modelModel identifier or generation configAt most one per train; no ADK upgrade on the same train
libraryADK or google-genai versionRides alone or with prompt cars only
evalNew cases, changed thresholdsBoards one train before the change it guards
configLimits, timeouts, feature flagsPrefer flags outside the train; see progressive rollouts

Two rules deserve explanation. A model change and a library upgrade never share a train because each can shift behaviour across the whole agent, and if the combined candidate regresses you cannot tell which caused it. Eval changes board a train early because a threshold changed on the same train as the behaviour it measures can make a regression look like a pass. Provider model versions are retired on the provider's own schedule, so check its model lifecycle page and plan model cars weeks ahead rather than discovering a shutdown date in an incident.

Gates on the frozen candidate

The frozen candidate passes four gates. The evaluation suite runs on the combined manifest and on each car in isolation, so a failing car can be identified and removed. A soak in a staging environment replays recorded traffic to find latency and cost regressions that small eval sets miss. A canary sends a small share of new sessions to the candidate and compares per-session metrics with production. And a named release captain makes the departure call against written criteria.

# Run on the frozen candidate after the cut. Pseudocode; adapt to your CI.
def run_train(manifest):
    results = {}
    for car in manifest.cars:
        if car.status == "dropped":
            continue
        results[car.id] = evaluate(build_without_other_new_cars(manifest, keep=car))  # isolation
    combined = evaluate(build(manifest))                       # everything that boarded
    failing = [c for c, r in results.items() if not r.meets(manifest.eval.thresholds)]
    if failing:
        for car_id in failing:
            revert_car(manifest, car_id)                       # revert, never patch at the station
        combined = evaluate(build(manifest))
    if not combined.meets(manifest.eval.thresholds):
        return hold_train(manifest, reason="interaction failure", evidence=combined)
    return promote_to_canary(manifest)

Evaluating each car in isolation multiplies eval cost by the number of cars, so many teams evaluate the combined candidate first and bisect only when it fails. Either way, the response to a failing car is to revert it and re-run, never to patch it on the release branch: a patch made under schedule pressure is a new, unevaluated change. If the candidate fails even with every failing car removed, the problem is an interaction; hold the train, and ship nothing rather than something untested.

The hotfix lane and missed trains

Some changes cannot wait a week: a tool leaking data, a prompt causing harmful advice, a model identifier that is about to be shut down. The hotfix lane exists so the train timetable does not get broken for these. Entry is gated by severity and needs the release captain's approval. A hotfix contains one change, is made on the current production manifest rather than on main, runs the same gates with a shortened canary, and is merged back to main immediately so the next train carries it. Most agent hotfixes should be kill switches or flag flips rather than code, because those take effect in seconds and are already tested.

The missed-train policy is equally simple: a change that misses the cut, or is dropped at the gates, waits for the next train. Track how often each team's cars are dropped. A team that regularly misses is telling you its changes are too large or its eval cases arrive too late, and the fix is smaller cars, not a later cut.

Worked example: train 2026.41

Train 2026.41 boards four cars: a tone change to the refund instructions, version two of the escalation tool with its instruction update, fourteen new eval cases, and a bump of the model identifier. The eval cases guarding the model car boarded on train 2026.40, as the rule requires; this train's eval car carries cases for 2026.42. At the cut the script writes the manifest and the factory check passes.

The combined candidate misses the task-success threshold. Bisecting finds that the model car alone causes it: the newer model calls the escalation tool too eagerly. The model car is reverted, the combined candidate passes, staging soak shows a small rise in p95 latency within budget, and the canary runs Thursday on 5 percent of new sessions. The train departs Friday with three cars. The model change gets its own instruction adjustment and boards train 2026.42 alone, as the boarding rules require for model changes.

Measuring the train

  • On-time departures. The share of trains that left on schedule. Below about 90 percent, the timetable is not credible.
  • Cars dropped per train, by kind and team.
  • Lead time from merge to full traffic; with this timetable a change merged just after the cut waits about ten days.
  • Rollbacks after departure and hotfixes per month; both should be rare, or the gates are missing something.

Failure modes

  • Holding the train for a feature. Once a departure waits for one change, every change starts negotiating, and the timetable is gone.
  • Moving references. A model alias or remote prompt changes after the cut, so production differs from what passed.
  • Patching at the station. Fixing a failing car on the release branch ships an untested change.
  • Hotfix lane as a fast lane. Features labelled urgent bypass the gates; police entry by severity.
  • Coupled cars. A tool change and its prompt boarded as separate cars, so dropping one breaks the other.

Trade-offs

ChoiceBenefitCost
Weekly versus daily trainOne eval run per week; predictableUp to a week's wait for small fixes
Per-car isolation evalsFinds the failing car at onceEval cost scales with car count
Train versus continuous deploymentBatches expensive evaluationLarger batches, slower feedback
Flags outside the trainInstant config changesUntrained combinations in production

What to do next

  1. Choose a cadence from how long a full evaluation run takes and how much quota it uses; write the timetable down and name a release captain.
  2. Write a cut script that resolves model aliases, fingerprints instructions and copies library versions into a manifest.
  3. Make the service build its agent only from the manifest, as in the factory above, and refuse to start on drift.
  4. Adopt the boarding rules, especially one model car per train and library upgrades alone.
  5. Define hotfix entry criteria and a missed-train rule, and track on-time departures and dropped cars.
  6. Read ADK Java versioning for manifests and session pinning, agent CI/CD for change-aware gate plans, evals in CI, the canary controller and progressive rollouts and kill switches.
Key takeaway: A release train ships an agent on a fixed timetable: changes board as independent cars, a cut freezes them into a manifest that pins the model identifier, fingerprints the instruction and copies library versions, and the frozen candidate passes evaluation, soak and canary before it departs with whatever passed. Failing cars are reverted, never patched in place, and a change that misses the cut takes the next train. Keep model changes and ADK upgrades on separate trains, board eval changes a train early, build the agent only from the manifest, and use a narrow, severity-gated hotfix lane for what cannot wait.