A release train is a release process with a fixed timetable: the train departs on schedule, changes that are ready board it, and changes that are not ready wait for the next one. The schedule never moves to wait for a feature. Large software organisations adopted trains to replace release dates negotiated around the slowest change, and the same idea fits agents built with the Agent Development Kit for Java unusually well, because an agent's behaviour changes through many small, interacting edits, prompts, tools, model identifiers, library versions and evaluation sets, whose effects you only see when you evaluate them together.
This page designs a weekly train for an ADK Java agent: the timetable, the cut that freezes a manifest, what kinds of change may board and together with what, the gates the frozen candidate must pass, the hotfix lane, and what happens to a change that misses the train or fails on it. It covers process; for the mechanics it relies on, fingerprinting, canaries and flags, it links to the pages that own them.
Why agents suit a train
Agent changes differ from ordinary code changes in three ways that make a timetable valuable. Their effects are statistical: a one-line instruction edit can move task success by several points, and only a full evaluation run, which takes time and model quota, shows it. Their effects interact: a new tool and an instruction that mentions it are only meaningful together, while a model upgrade can change how every prompt is read. And their cost to evaluate is high enough that running the whole suite on every merge is often unaffordable, while running it once on a frozen candidate is cheap.
Continuous deployment answers these with per-change gates, which work when each change is small and the gates are fast. A train batches changes so the expensive evaluation runs once per departure on exactly what will ship, gives every stakeholder a predictable date, and turns schedule pressure into a simple rule: if your change is not ready, it takes the next train. Teams shipping a single agent with fast evals may not need one; teams with several contributors to one agent usually do.
The timetable and the cut
Pick a cadence your evaluation can keep up with. Weekly suits most teams: boarding from Monday, a cut at noon on Tuesday, gates on Tuesday and Wednesday, a canary from Thursday and full traffic on Friday morning (or the following Monday, if nobody watches production at weekends). The cut is the key moment. At the cut, a script reads the head of the release branch, records the exact commit, resolves every moving reference into a fixed one, and writes the manifest. Nothing that changes after the cut affects this train.
# release/train-2026.41.yaml (generated at the cut, then read-only)
train: "2026.41"
cut_commit: 3f1c9e2
departs: "2026-10-16T09:00:00Z"
agent:
name: support_agent
model: "<exact versioned model id>" # pinned; never an alias that moves
instruction_file: prompts/support_agent.md
instruction_sha256: 9b0e... # fingerprint of the exact text
tools: [lookupOrder, issueRefund, escalate]
generate_config: { temperature: 0.2, max_output_tokens: 1024 }
libraries:
google-adk: "<version from the lock file>"
eval:
suite: evals/support_v14
thresholds: { task_success: 0.92, policy_violations: 0, p95_turn_ms: 6000 }
cars:
- { id: prompt-refund-tone, kind: prompt, owner: ana, pr: 812 }
- { id: tool-escalate-v2, kind: tool, owner: raj, pr: 815 }
- { id: eval-v14-cases, kind: eval, owner: mei, pr: 817 }
- { id: model-flash-bump, kind: model, owner: tom, pr: 820, status: dropped }Resolving moving references is the part teams skip. A model alias that the provider repoints, an instruction loaded from a configuration service, or a library version range in the build file each lets the shipped agent differ from the tested one. The manifest pins the model identifier, fingerprints the instruction text, and copies library versions from the lock file. The service then builds the agent only from the manifest and refuses to start if anything has drifted.
/** Builds the agent only from a frozen manifest, so what was tested is what ships. */
public final class AgentFactory {
public static LlmAgent fromManifest(Manifest m, ToolRegistry registry) throws IOException {
String instruction = Files.readString(Path.of(m.agent().instructionFile()));
String actual = sha256(instruction);
if (!actual.equals(m.agent().instructionSha256())) {
throw new IllegalStateException("instruction drifted from train " + m.train()
+ ": expected " + m.agent().instructionSha256() + ", found " + actual);
}
return LlmAgent.builder()
.name(m.agent().name())
.model(m.agent().model())
.instruction(instruction)
.tools(registry.resolve(m.agent().tools())) // fails on unknown tool names
.generateContentConfig(m.agent().generateConfig().toGenai())
.build();
}
}
Cars and boarding rules
Each change boards as a car with a kind, an owner and a link to its review. Cars are what make the train robust: when one fails, it is reverted and the rest still leave. That only works if cars are independent enough to remove, so the boarding rules exist mainly to keep them so.
| Car kind | Examples | Boarding rule |
|---|---|---|
| prompt | Instruction text, few-shot examples | Any number; each needs eval cases that exercise it |
| tool | New tool, schema or behaviour change | Rides with the prompt change that uses it, as one car |
| model | Model identifier or generation config | At most one per train; no ADK upgrade on the same train |
| library | ADK or google-genai version | Rides alone or with prompt cars only |
| eval | New cases, changed thresholds | Boards one train before the change it guards |
| config | Limits, timeouts, feature flags | Prefer flags outside the train; see progressive rollouts |
Two rules deserve explanation. A model change and a library upgrade never share a train because each can shift behaviour across the whole agent, and if the combined candidate regresses you cannot tell which caused it. Eval changes board a train early because a threshold changed on the same train as the behaviour it measures can make a regression look like a pass. Provider model versions are retired on the provider's own schedule, so check its model lifecycle page and plan model cars weeks ahead rather than discovering a shutdown date in an incident.
Gates on the frozen candidate
The frozen candidate passes four gates. The evaluation suite runs on the combined manifest and on each car in isolation, so a failing car can be identified and removed. A soak in a staging environment replays recorded traffic to find latency and cost regressions that small eval sets miss. A canary sends a small share of new sessions to the candidate and compares per-session metrics with production. And a named release captain makes the departure call against written criteria.
# Run on the frozen candidate after the cut. Pseudocode; adapt to your CI.
def run_train(manifest):
results = {}
for car in manifest.cars:
if car.status == "dropped":
continue
results[car.id] = evaluate(build_without_other_new_cars(manifest, keep=car)) # isolation
combined = evaluate(build(manifest)) # everything that boarded
failing = [c for c, r in results.items() if not r.meets(manifest.eval.thresholds)]
if failing:
for car_id in failing:
revert_car(manifest, car_id) # revert, never patch at the station
combined = evaluate(build(manifest))
if not combined.meets(manifest.eval.thresholds):
return hold_train(manifest, reason="interaction failure", evidence=combined)
return promote_to_canary(manifest)Evaluating each car in isolation multiplies eval cost by the number of cars, so many teams evaluate the combined candidate first and bisect only when it fails. Either way, the response to a failing car is to revert it and re-run, never to patch it on the release branch: a patch made under schedule pressure is a new, unevaluated change. If the candidate fails even with every failing car removed, the problem is an interaction; hold the train, and ship nothing rather than something untested.
The hotfix lane and missed trains
Some changes cannot wait a week: a tool leaking data, a prompt causing harmful advice, a model identifier that is about to be shut down. The hotfix lane exists so the train timetable does not get broken for these. Entry is gated by severity and needs the release captain's approval. A hotfix contains one change, is made on the current production manifest rather than on main, runs the same gates with a shortened canary, and is merged back to main immediately so the next train carries it. Most agent hotfixes should be kill switches or flag flips rather than code, because those take effect in seconds and are already tested.
The missed-train policy is equally simple: a change that misses the cut, or is dropped at the gates, waits for the next train. Track how often each team's cars are dropped. A team that regularly misses is telling you its changes are too large or its eval cases arrive too late, and the fix is smaller cars, not a later cut.
Worked example: train 2026.41
Train 2026.41 boards four cars: a tone change to the refund instructions, version two of the escalation tool with its instruction update, fourteen new eval cases, and a bump of the model identifier. The eval cases guarding the model car boarded on train 2026.40, as the rule requires; this train's eval car carries cases for 2026.42. At the cut the script writes the manifest and the factory check passes.
The combined candidate misses the task-success threshold. Bisecting finds that the model car alone causes it: the newer model calls the escalation tool too eagerly. The model car is reverted, the combined candidate passes, staging soak shows a small rise in p95 latency within budget, and the canary runs Thursday on 5 percent of new sessions. The train departs Friday with three cars. The model change gets its own instruction adjustment and boards train 2026.42 alone, as the boarding rules require for model changes.
Measuring the train
- On-time departures. The share of trains that left on schedule. Below about 90 percent, the timetable is not credible.
- Cars dropped per train, by kind and team.
- Lead time from merge to full traffic; with this timetable a change merged just after the cut waits about ten days.
- Rollbacks after departure and hotfixes per month; both should be rare, or the gates are missing something.
Failure modes
- Holding the train for a feature. Once a departure waits for one change, every change starts negotiating, and the timetable is gone.
- Moving references. A model alias or remote prompt changes after the cut, so production differs from what passed.
- Patching at the station. Fixing a failing car on the release branch ships an untested change.
- Hotfix lane as a fast lane. Features labelled urgent bypass the gates; police entry by severity.
- Coupled cars. A tool change and its prompt boarded as separate cars, so dropping one breaks the other.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| Weekly versus daily train | One eval run per week; predictable | Up to a week's wait for small fixes |
| Per-car isolation evals | Finds the failing car at once | Eval cost scales with car count |
| Train versus continuous deployment | Batches expensive evaluation | Larger batches, slower feedback |
| Flags outside the train | Instant config changes | Untrained combinations in production |
What to do next
- Choose a cadence from how long a full evaluation run takes and how much quota it uses; write the timetable down and name a release captain.
- Write a cut script that resolves model aliases, fingerprints instructions and copies library versions into a manifest.
- Make the service build its agent only from the manifest, as in the factory above, and refuse to start on drift.
- Adopt the boarding rules, especially one model car per train and library upgrades alone.
- Define hotfix entry criteria and a missed-train rule, and track on-time departures and dropped cars.
- Read ADK Java versioning for manifests and session pinning, agent CI/CD for change-aware gate plans, evals in CI, the canary controller and progressive rollouts and kill switches.