Rolling back a stateless web service is a traffic operation: point the load balancer at the previous revision and the problem is gone. Rolling back an agent is harder, because an agent release is not one thing. It is a binary, an instruction prompt, a model identifier, a set of tool contracts, a shape of session state, and, by the time anyone notices a problem, a pile of things the bad release already did: state it wrote, memories it saved, refunds it issued. Shifting traffic back reverts the first item. The others need deliberate handling, and some of them cannot be reverted at all.

This page treats rollback as an incident decision for an ADK Java agent. It covers which levers exist and how fast each one is, the N-1 rule that makes state survive a revert, what happens to old conversations when a tool disappears, how to contain writes the bad release made, a worked incident, and how to rehearse the whole thing so the first real rollback is not the first attempt. The mechanics of moving traffic between revisions are covered in the deploy pipeline and canary articles; this one is about everything around them.

An agent release is five things

Start by listing what can change between two releases, because each has its own rollback lever and its own speed. If all five ship inside one container image, every rollback is a full binary revert, which is the slowest and bluntest option. Teams that roll back well have pulled the fast-changing parts out so they can be reverted on their own.

PartTypical changeLeverTime to revertReverts data?
Feature flag / configenable a new tool, raise a limitflip the flagsecondsno
Instruction promptreworded policy, new examplesrepoint to previous prompt versionseconds to minutesno
Model identifiernew model or model versionconfig changeminutesno
Binary (code, tools, ADK version)new tool, changed callbackshift traffic to previous revisionminutesno
State / memory schemanew state keys, new memory formatnone: roll forward or migratehoursonly by migration

The last column is the point of the table. None of the levers touch data. A rollback restores the behaviour of the old version, running against whatever the new version left behind. The rest of this article is about making that combination safe. To make the levers real, record every part in a release manifest (an image digest, a prompt version, a model string, a flag snapshot) so the incident question "what changed?" has an exact answer instead of a guess. The versioning article shows how to fingerprint such a manifest.

Choosing the lever

During an incident the first decision is which lever to pull, and the rule is to use the smallest one that removes the change. Diff the manifest of the current release against the last known-good one; if only the prompt changed, revert the prompt and leave the binary alone. A binary revert also reverts every unrelated fix that shipped with it, which can reopen bugs you closed last week.

Rollback decision: smallest lever first, then clean up what the bad release wroteSignal breacheval, tool errors, costWhich part changed?release manifest diffConfig flagsecondsPrompt / modelminutes, no buildBinary revisiontraffic shiftRoll forwardstate not N-1 safeVerify on the old versionsame signals, real turnsContain what was writtenstate, memory, side effectsPostmortemadd a drill caseReverting code restores behaviour; it never reverts data. The right-hand boxes are the part teams forget.
Pick the smallest lever that removes the change, verify on the old version, then contain what the bad release wrote.

Writing the procedure down as code, even if a human runs it, removes debate at 2 a.m.:

enum Lever { FLAG, PROMPT, MODEL, BINARY, ROLL_FORWARD }

static Lever chooseLever(Manifest bad, Manifest good, boolean stateIsNMinus1Safe) {
    if (!bad.flags().equals(good.flags()) && bad.sameExceptFlags(good)) return Lever.FLAG;
    if (bad.sameExceptPrompt(good)) return Lever.PROMPT;
    if (bad.sameExceptModel(good)) return Lever.MODEL;
    // Anything touching code, tools or state shape needs a binary decision.
    return stateIsNMinus1Safe ? Lever.BINARY : Lever.ROLL_FORWARD;
}

Roll forward means shipping a fix on top of the bad release instead of reverting it. It is the right call when the old binary cannot safely read what the new one wrote, and when the fix is small and the signals are degraded rather than catastrophic. If the agent is moving money or deleting data, contain first: turn the dangerous tool off with a flag (which needs no binary decision at all), then decide between revert and fix with the bleeding stopped.

State that survives a rollback: the N-1 rule

Session state in ADK Java is a map of keys to values that the runner persists with the session. A new release that writes a new key, or a new shape under an existing key, creates sessions the old release has never seen. After a rollback, the old binary picks those sessions up on the user's next turn. Whether that works depends on one discipline, the N-1 rule: release N-1 must be able to read anything release N writes. Concretely, three habits make it hold.

  1. Additive changes only, in two steps. To change a shape, first ship a release that can read both old and new shapes but still writes the old one; only the next release starts writing the new shape. This is expand/contract, the same pattern used for database columns.
  2. Tolerant readers. Old code ignores keys and fields it does not know rather than failing.
  3. No lossy write-back. Old code must not read a newer structure, drop the fields it does not understand, and write the truncated version back, because that destroys data the moment you roll forward again.
record CartState(int schema, List<String> skus, Map<String, Object> unknown) {
    static final int SUPPORTED_SCHEMA = 2;

    static CartState read(Map<String, Object> state) {
        if (!(state.get("cart") instanceof Map<?, ?> raw)) return new CartState(SUPPORTED_SCHEMA, List.of(), Map.of());
        int schema = raw.get("schema") instanceof Number n ? n.intValue() : 1;
        List<String> skus = raw.get("skus") instanceof List<?> l
                ? l.stream().map(String::valueOf).toList() : List.of();
        Map<String, Object> unknown = new HashMap<>();
        raw.forEach((k, v) -> { if (!Set.of("schema", "skus").contains(k)) unknown.put(String.valueOf(k), v); });
        return new CartState(schema, skus, unknown);
    }

    Map<String, Object> toMap() {
        if (schema > SUPPORTED_SCHEMA) {
            // Written by a newer release: preserve every field we did not understand.
            Map<String, Object> out = new HashMap<>(unknown);
            out.put("schema", schema);
            out.put("skus", skus);
            return out;
        }
        return Map.of("schema", SUPPORTED_SCHEMA, "skus", skus);
    }
}

The unknown map is what makes write-back safe: a rolled-back binary keeps the newer fields intact. Test the rule mechanically rather than by review, as described in the rehearsal section below.

Tools and old conversations

Tools are where agent rollback differs most from service rollback, because the conversation history is part of the input. Suppose release N adds a schedule_callback tool and some sessions call it. After a revert to N-1, those sessions still contain model events with function calls to schedule_callback and the matching function responses, and the history is sent back to the model on the next turn. The model now sees evidence of a tool it is no longer offered, and it may try to call it again.

What ADK Java does next is specific. In adk-java main as of 2026-10-08, Functions.handleFunctionCalls checks each call against the agent's tools; a name that is not present is logged at WARN as Tool not found and skipped, so that call receives no function response. The turn does not crash, but the model asked for something and got nothing back, and what it does next is undefined. Three practical consequences follow.

  • Alert on the Tool not found log line. After a rollback it is the most direct signal that old history is fighting the reverted tool set; in steady state it should be zero.
  • Prefer turning a new tool off with a flag that keeps it registered but returns a polite "this action is temporarily unavailable" result. The model then gets a real response and can tell the user.
  • Never rename a tool and remove the old name in the same release. Keep the old name as an alias for at least one release so that both N-1 and N understand every history the other produced; zero-downtime deploys walks through a rename.

What a rollback cannot undo

The hardest part of an agent rollback is that the bad release acted on the world. It may have issued refunds, sent emails, written long-term memories that will be recalled for months, or set state that steers future turns. None of that is reverted by traffic. You can only contain it, and containment depends on one thing being true before the incident: every write carries the release that made it.

static final String RELEASE = System.getenv().getOrDefault("AGENT_RELEASE", "unknown");

public Map<String, Object> issueRefund(
        @Schema(name = "order_id", description = "Order to refund") String orderId,
        @Schema(name = "amount_cents", description = "Amount in cents") long amountCents,
        ToolContext ctx) {
    String idempotencyKey = ctx.sessionId() + ":" + ctx.functionCallId().orElse(orderId);
    RefundResult r = payments.refund(orderId, amountCents, idempotencyKey,
            Map.of("agent_release", RELEASE, "session_id", ctx.sessionId()));
    return Map.of("status", r.status(), "refund_id", r.id());
}

With that stamp in place, containment becomes a query rather than an archaeology project:

  • Side effects: list every refund, ticket or message tagged with the bad release, and hand the list to the owning team for review or compensation. Compensation is a business action, not an undo; the compensation article covers designing it.
  • Memory: tag memory entries with the release when they are written, and on rollback quarantine entries from the bad window so retrieval excludes them until reviewed. Deleting is irreversible; quarantine first.
  • State: if the bad release wrote harmful values (a wrong discount flag, a corrupted plan), write a one-off repair job keyed on sessions that had turns on that release, rather than resetting all sessions.

Worked example: a prompt that doubled refunds

A support agent ships release 41 with a reworded refund policy in its instruction and a new lookup_warranty tool. Within two hours the refund-rate dashboard, which the team tracks per release, shows refunds per hundred sessions at 9.8 against a baseline of 4.6. The manifest diff shows three changes: prompt version 17 to 18, the new tool, and an ADK library bump.

The on-call engineer turns off the refund tool's auto-approve flag first, so every refund now needs human approval; the rate of executed refunds drops within a minute. Next, because the prompt is versioned separately, they revert only the prompt to version 17, keeping the library bump and the new tool. Refund requests return to the baseline over the next hour of sessions, which confirms the prompt as the cause; the binary never moved. Finally, the containment query returns 212 refunds tagged with release 41 and prompt 18. Finance reviews them; 31 were outside policy and are handled as goodwill or reversed through the payment provider. Long-term memories written by those sessions are quarantined for review because several recorded "customer is eligible for refunds on accessories", which would have leaked the bad policy into future conversations.

Had the prompt shipped inside the image, the only lever would have been a binary revert, which would also have pulled the warranty tool out from under sessions that were using it, producing exactly the Tool not found situation described above.

Rehearsing rollback

A rollback path that is never exercised does not work when it is needed. Make it a test that runs in CI on every release candidate: start a session on release N, roll back, and continue the same session on N-1. With the agent definitions for both releases buildable in one test (or both images pulled in a container test), the core check looks like this:

@Test
void sessionsCreatedOnCandidateSurviveRollback() {
    Session s = runOnRelease(candidateAgent(), "u1", recordedScript("refund_happy_path"));
    Map<String, Object> before = Map.copyOf(s.state());

    // Continue the same persisted session on the previous release.
    Session after = continueOnRelease(previousAgent(), s, "and what is the status now?");

    assertThat(logs()).doesNotContain("Tool not found");
    assertThat(after.state().keySet()).containsAll(before.keySet());   // no lossy write-back
    assertThat(lastTextResponse(after)).isNotEmpty();
}

Run it against recorded model responses so it is deterministic, and add a case to it after every incident. Separately, measure time to revert for each lever in a staging drill once a quarter; if the prompt revert takes 40 minutes because it needs a pull request and a build, it is not really a fast lever.

Failure modes

  • Everything in one image. Every rollback reverts unrelated fixes. Move prompts, model names and flags out of the binary.
  • Rolling back past a state change. The old binary throws on, or silently truncates, state written by the new one. Enforce the N-1 rule with the drill test.
  • Removed tools in old histories. Tool not found warnings and confused turns. Disable tools by flag; keep aliases for one release.
  • Poisoned memory. The bad behaviour returns weeks later through recalled memories. Tag and quarantine by release.
  • Unattributable side effects. No release stamp on writes means no containment list.
  • Rollback to a version that no longer works. A model identifier the old prompt was tuned for may have been retired by the provider; check that every older manifest you might roll back to is still runnable.

Trade-offs

Separating prompts, models and flags from the binary buys fast, narrow rollbacks at the cost of more moving parts and a manifest you must keep honest; a prompt changed outside the manifest is worse than one baked into the image. The N-1 rule slows schema changes to two releases. Release stamps add a field to every write and a dependency on downstream systems accepting metadata. In exchange, rollbacks become minutes of routine work rather than hours of reconstruction, and a team that can roll back cheaply ships more often.

What to do next

  1. Write a release manifest for your agent listing image digest, prompt version, model identifier and flag snapshot.
  2. Move the instruction prompt and model identifier out of the image so each can be reverted alone.
  3. Put a kill-switch flag on every tool with external side effects.
  4. Stamp the release identifier on every external write, memory entry and state change you may need to find later.
  5. Adopt the N-1 rule: expand/contract for state shapes, tolerant readers, no lossy write-back.
  6. Add the rollback drill test to CI and an alert on the Tool not found log line.
  7. Time each lever in a staging drill and write the decision procedure into the on-call runbook.
Key takeaway: An agent release is a binary, a prompt, a model, tool contracts and a state shape, and a traffic rollback reverts only the first. Version the parts separately so you can pull the smallest lever, keep state readable by the previous release, disable tools by flag instead of removing them, and stamp every write with its release so you can contain what a bad version did. Rehearse the whole path in CI, because rollback never reverts data.