A hotfix is a production change made outside the normal release rhythm to stop ongoing harm. In a conventional service the hard part is shipping fast without breaking something else. In an agent there is an extra difficulty: the obvious fix, a sentence added to the instruction, changes the behaviour of every conversation in ways you cannot fully test, while the less obvious fixes, a guard in a callback or a flag, are narrow, deterministic and testable in minutes.
This article is about the hotfix itself for an ADK Java agent: when to hotfix rather than roll back or flip a kill switch, how to pick the layer the fix lands in by blast radius, how to reproduce the failure from recorded session events and write the failing test first, what the reduced gate must still cover, what happens to sessions already in flight, and how to forward-port so the next release does not quietly undo the fix. The timetable rules for when a hotfix may board are covered by the release train; this page starts once you have decided to ship one.
Hotfix, rollback or kill switch
Three tools stop production harm, and a hotfix is the most expensive of them, so check the others first.
| Response | Use when | Not enough when |
|---|---|---|
| Kill switch or tool quarantine | One tool or capability is causing harm and the agent can work without it for a while | The broken capability is core to every conversation |
| Rollback | The last release introduced the problem and nothing in it must stay | The bug is old, the cause is external, or the release also carried an urgent fix |
| Hotfix | The cause is known, the fix is small, and the harm continues until it ships | The cause is not yet understood; ship containment first |
Agent incidents are disproportionately external: an upstream API changes a field, a model behind an alias shifts behaviour, a partner starts sending inputs in a new format. Rollback cannot help with any of those, because the previous release has the same exposure. That is why agent teams hotfix more often than their service colleagues, and why the discipline below matters. A useful rule: contain first with the playbook's kill switch if harm is severe, then hotfix calmly rather than racing the harm with an untested change.
The fix ladder
Rank candidate fixes by blast radius, meaning how many conversations and code paths the change can affect, and pick the lowest rung that stops the harm.
- Config or flag. Disable a tool, lower a limit, route a slice of users to a fallback. No build, instant revert, if the knob already exists. Flags read once per invocation make this safe mid-conversation.
- Guard callback. A
beforeToolCallbackSyncthat rejects or corrects a specific bad call. Deterministic, unit-testable without a model, touches only calls that match its condition. - Tool code fix. Correct the integration: parse the renamed field, fix the unit conversion. Deterministic but touches every call to that tool.
- Instruction edit. Changes how the model reasons in every turn, including turns that have nothing to do with the bug. Only testable statistically.
- Model pin or change. Affects everything. Reserve for cases where the model itself is the regression, and pin an explicit version rather than an alias.
The common mistake is to treat the instruction as the smallest change because it is the smallest diff. One added sentence such as "never refund more than the item price" can make the model more hesitant on every refund, more verbose, or more likely to ask clarifying questions. That is a behavioural change across the whole traffic mix, shipped under time pressure, with a reduced gate. A guard callback that enforces the same rule in code has none of those effects.
Reproduce from the session record
Before writing a fix, collect the failing cases from the session store. Recorded events contain the exact function calls the model made, with arguments, which is usually all you need to reproduce a tool-layer bug without calling a model at all.
// Pull the recorded tool calls for a list of incident sessions.
List<Map<String, Object>> collectRefundCalls(BaseSessionService sessions, String app,
List<SessionRef> incident) {
List<Map<String, Object>> calls = new ArrayList<>();
for (SessionRef ref : incident) {
ListEventsResponse resp = sessions.listEvents(app, ref.userId(), ref.sessionId()).blockingGet();
for (Event e : resp.events()) {
for (FunctionCall fc : e.functionCalls()) {
if (fc.name().orElse("").equals("issue_refund")) {
calls.add(Map.of("session", ref.sessionId(),
"args", fc.args().orElse(Map.of())));
}
}
}
}
return calls; // write to a JSON fixture under src/test/resources/incidents/
}Save the fixture with the incident id in its file name and keep it forever. It is the regression test, the evidence for the post-incident review, and the input to the next step. Strip personal data before it reaches the repository; argument maps often contain names, emails or addresses.
Worked example: a guard for over-paid refunds
A worked incident makes the pattern concrete. A payments agent offers partial refunds. A prompt change two weeks earlier made the model pass the order total as the refund amount when a user asked to return one item from a multi-item order. Refunds are over-paid; the bug predates the last three releases, so rollback would revert other wanted changes. The immediate fix sits on rung 2: a guard that refuses any refund larger than the price of the line item named in the call, and tells the model why so it can retry with the right amount.
public final class RefundCapGuard {
private final OrderLookup orders;
private final Flags flags;
public Optional<Map<String, Object>> check(InvocationContext inv, BaseTool tool,
Map<String, Object> args, ToolContext ctx) {
if (!tool.name().equals("issue_refund") || !flags.enabled("hotfix_refund_cap")) {
return Optional.empty();
}
String orderId = (String) args.get("orderId");
String itemId = (String) args.get("itemId");
long requested = ((Number) args.get("amountCents")).longValue();
long itemPrice = orders.linePriceCents(orderId, itemId);
if (requested <= itemPrice) {
return Optional.empty(); // legitimate: run the tool
}
metrics.counter("hotfix_refund_cap_blocked").increment();
return Optional.of(Map.of(
"status", "rejected",
"reason", "Refund exceeds the price of item " + itemId + " (" + itemPrice + " cents). "
+ "Refund only the item being returned."));
}
}
// Wiring: LlmAgent.builder()...beforeToolCallbackSync(refundCapGuard::check)Write the test before the guard, from the incident fixture: every recorded over-paid call must be rejected, and a set of recorded legitimate refunds from the same period must pass. Because the guard is a pure function of its inputs plus a lookup you can stub, this test runs in milliseconds and needs no model. Two more properties belong in it: the guard is a no-op for other tools, and it is a no-op when its flag is off, which is your instant revert. The counter tells you, after shipping, how often the model still tries to over-refund, which is the evidence the real fix needs.
The reduced gate
A hotfix skips most of the release gate, but not the parts that protect against the two ways hotfixes go wrong: fixing the bug while breaking something nearby, and shipping a build that differs from production in ways nobody intended.
| Gate step | Normal release | Hotfix |
|---|---|---|
| Incident regression fixture | n/a | Required, must fail before and pass after |
| Unit and tool tests | All | All; they are fast |
| Evaluation suite | Full | Slice that exercises the changed tool or path |
| Safety and policy suite | Full | Full, always |
| Build inputs | Fresh from main | Production tag plus the fix; same base image and lockfile |
| Rollout | Train ramp | Short canary with the regression probe, then full |
The build-input row is the one teams forget. Branch from the tag that is running in production, not from main, and rebuild with the same dependency lockfile and base image digest. Otherwise the hotfix silently carries two weeks of unreleased main and a newer framework version into production under the label of a one-line fix.
Sessions in flight
An agent hotfix lands in the middle of conversations. Three rules keep that safe.
- No state schema changes. Sessions in flight hold state written by the old code. A hotfix that renames a key or changes a value's shape breaks them. If the fix needs new state, read old and new forms and write only the new.
- Expect old behaviour in history. Conversations in flight contain the model's earlier, wrong tool calls. Models imitate patterns in their own history, so a session that already over-refunded once may try again. A guard handles this; an instruction edit is fighting in-context precedent.
- Make the guard's message actionable. The rejection is returned to the model as the tool's result, mid-conversation. A reason that names the right amount lets the model recover in the same turn instead of apologising to the user.
Forward-port and close out
The most common hotfix failure is not a bad fix. It is a good fix that the next scheduled release removes, because the hotfix branch was never merged back and main still has the old code. Close the loop explicitly.
- Cherry-pick the fix and its regression fixture onto main the same day, as its own change, before the next release cut.
- Make the next release's gate run the incident fixture; if the fix is missing, the fixture fails and the release stops.
- Open the root-cause change: here, the prompt or tool schema that let the model send the order total. Ship it on the normal train with the full evaluation suite.
- Give the guard and its flag an expiry date. When the root fix has been live for an agreed period and the blocked-counter reads zero, remove the guard in a normal release, or keep it deliberately as a permanent invariant and say so.
The regression detection fingerprints for the affected tool are worth reviewing in the post-incident meeting: the prompt change that caused this shifted the distribution of refund amounts two weeks before anyone noticed.
Failure modes
- Hotfix reverted by the next release. Prevented by the forward-port and by the incident fixture in the release gate.
- Guard too broad. It blocks legitimate calls, for example refunds that include shipping. Test it against recorded legitimate calls, not only failing ones, and watch the blocked counter for a jump right after rollout.
- Instruction hotfix with side effects. Satisfaction or task completion drops in unrelated flows. Prefer code; if the instruction must change, run the full evaluation suite, not a slice.
- Accidental upgrade. The hotfix build pulls a newer framework or model alias. Build from the production tag with pinned inputs.
- Unversioned config change. A flag flipped by hand, with no record, becomes tribal knowledge. Treat config hotfixes as changes with review and an audit trail.
- Fix verified only on synthetic data. Replay the recorded incident calls; synthetic examples tend to miss the exact shape that triggered the bug.
Trade-offs
Every hotfix trades verification depth for speed. Lower rungs make that trade cheap: a guard callback can be fully verified in minutes because it is deterministic, while an instruction edit cannot be fully verified in any time a hotfix allows. The second trade is accumulation. Guards are easy to add and tend to pile up, each a small special case the next engineer must understand. Expiry dates, owners and a periodic review keep the layer from becoming an unowned rule engine. The third is speed versus understanding: shipping containment first and the hotfix second costs an extra change but avoids writing a fix for a cause you have only guessed.
What to do next
- Write down your hotfix decision order: containment, rollback, then the fix ladder, and link it from the on-call runbook.
- Make sure every write tool has an existing flag, so rung 1 is available before you need it.
- Build a small library that turns recorded session events into test fixtures, and try it on one past incident.
- Add a guard-callback template with flag check, metric and actionable rejection message, plus its unit test skeleton.
- Script the hotfix branch: from the production tag, same lockfile and base image digest.
- Add an "incident fixtures" step to the release gate so forward-ports are enforced.
- Put an expiry date and owner on every guard currently in production, and review them.