An ADK Java agent can fail in ways a normal service cannot. It can stay up and answer every request while calling a refund tool twice per conversation. It can loop on a failing tool and burn a day's token budget in an hour. A prompt injection hidden in a retrieved document can steer it into a tool it should never touch. A model version change can quietly make it worse at following instructions. None of these trips a liveness probe, and rolling back the container does not help when the bad behaviour comes from the model, the data or a downstream API.
This playbook is about the first hour of those incidents, and about the engineering you must do before them so that the first hour goes well. It defines what an agent incident is, builds the containment levers into one ADK plugin using the callback contract of the framework's Plugin interface, lists the signals that detect each class of failure, walks the response phases, reconstructs a run from session events, and works a full incident from alert to recovery. For the general shape of an AI incident runbook, independent of framework, see designing an AI incident response runbook; this article is the ADK Java implementation.
What counts as an agent incident
Start by agreeing on what an incident is. A useful definition for agents: any event where the agent's behaviour, cost or data handling departs from its contract badly enough that a human must act now. Severity is set by impact, not by cause.
| Class | Example | Typical first signal | Default severity |
|---|---|---|---|
| Availability | Model provider returns 429 or 503 for most calls | Run error rate, model error callback counts | SEV2, SEV1 if total |
| Wrong action | Tool called with wrong arguments or twice | Downstream audit, user reports, tool call count per run | SEV1 if money or data moved |
| Runaway loop | Agent retries a failing tool until its call limit | Model calls per run, tokens per run | SEV2 |
| Security | Injected instructions trigger exfiltration through a tool | Guardrail hits, unusual tool and argument pairs | SEV1 |
| Quality regression | New model version ignores a policy instruction | Eval scores, escalation rate, thumbs-down rate | SEV3, SEV2 if harmful |
| Cost | Context growth makes every call 5x larger | Tokens per run, daily spend forecast | SEV3 |
Two rules follow. Wrong actions outrank outages: an agent that is down is safer than one that is moving money incorrectly. And the incident commander's first question is always which lever contains the harm now, never what the root cause is.
Containment levers in one plugin
ADK Java plugins are registered once on the Runner and see every agent, model call and tool call in the tree, including sub-agents. The Plugin interface documents three short-circuit points this playbook relies on. beforeRunCallback returns Maybe<Content>, and a non-empty value halts execution. beforeModelCallback returns Maybe<LlmResponse>, and a non-empty value is used instead of calling the model. beforeToolCallback returns Maybe<Map<String, Object>>, and a non-empty value stops the tool and is returned as its response. Empty means carry on.
public final class IncidentControlPlugin extends BasePlugin {
private final FlagSnapshot flags; // refreshed in the background, never blocks a run
private final IncidentLog log;
public IncidentControlPlugin(FlagSnapshot flags, IncidentLog log) {
super("incident_control");
this.flags = flags;
this.log = log;
}
@Override
public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
IncidentFlags f = flags.current();
if (f.agentDisabled(ctx.agent().name())) {
log.decision(ctx.invocationId(), "kill_switch", f.version());
return Maybe.just(Content.fromParts(Part.fromText(f.userMessage())));
}
return Maybe.empty();
}
@Override
public Maybe<LlmResponse> beforeModelCallback(CallbackContext cb, LlmRequest.Builder req) {
IncidentFlags f = flags.current();
int calls = log.countModelCall(cb.invocationId());
if (calls > f.maxModelCallsPerRun()) {
log.decision(cb.invocationId(), "loop_guard", f.version());
return Maybe.just(LlmResponse.builder()
.content(Content.fromParts(Part.fromText(f.loopMessage())))
.build());
}
return Maybe.empty();
}
@Override
public Maybe<Map<String, Object>> beforeToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext tc) {
IncidentFlags f = flags.current();
if (f.toolQuarantined(tool.name())) {
log.decision(tc.invocationId(), "quarantine:" + tool.name(), f.version());
return Maybe.just(Map.of(
"status", "unavailable",
"reason", "temporarily_disabled",
"instruction", "Do not retry. Tell the user this action is paused and nothing was done."));
}
return Maybe.empty();
}
}Runner runner = Runner.builder()
.agent(rootAgent)
.appName("support-desk")
.sessionService(sessionService)
.plugins(new IncidentControlPlugin(flagSnapshot, incidentLog), new ObservabilityPlugin(meter))
.build();Design points that matter under pressure. The flag snapshot is read from memory and refreshed on a timer; if the flag service is down the plugin keeps the last good snapshot, because an incident is exactly when dependencies fail. Every decision is logged with the flag version so the post-incident review can prove when containment took effect. The quarantine response tells the model explicitly that nothing happened and not to retry; a vague error invites the model to retry or, worse, to tell the user the refund was issued. The patterns for shaping those maps are in wrapping tool errors for the model.
The loop guard counts model calls per invocation. Clean up its counters in afterRunCallback and onRunErrorCallback. Check against the ADK release you run whether a built-in limit on model calls per run exists and is enabled; the plugin guard is a second line that you control with a flag either way.
Detection signals
Containment only helps if you notice. Each incident class needs at least one signal that fires within minutes. Most come from callbacks you can count in an observability plugin, as described in observability as a first-class concern.
| Signal | Source | Alert when |
|---|---|---|
| Run error rate | onRunErrorCallback count over runs | Above 2 percent for 5 minutes |
| Model error rate by type | onModelErrorCallback, error class | Any 429 or 5xx burst above baseline x3 |
| Model calls per run, p99 | beforeModelCallback count per invocation | Above 2x the 7-day p99 |
| Tool calls per run, by tool | afterToolCallback plus onToolErrorCallback | A write tool called more than once per run |
| Tokens per run | usage metadata in afterModelCallback | Above 2x baseline for 15 minutes |
| Guardrail and policy hits | your policy callbacks | Any hit on an exfiltration rule |
| Quality | Offline eval on sampled sessions | Score drops more than an agreed margin |
The write-tool rule deserves emphasis. Read tools can be called many times safely; a tool that changes state should almost never run twice in one invocation. An alert on that single condition catches the most expensive class of agent incident early. Error-handling paths that make it more likely, such as retrying a whole run, are discussed in ADK Java error recovery.
The response, phase by phase
- Detect and declare (0 to 5 minutes). Acknowledge the page, open an incident channel, name an incident commander. Record the agent name, app name, model version and the flag-snapshot version currently live.
- Triage (5 to 15 minutes). Classify with the table above. Ask three questions: is the agent taking wrong actions, is data leaving, is it only failing to answer? Pull five recent failing invocation ids from traces.
- Contain (target under 15 minutes). Pick the narrowest lever that stops the harm: quarantine one tool, then lower the loop guard, then switch the model, then disable the agent. Confirm in the decision log that the lever is firing on every replica.
- Preserve evidence. Snapshot affected sessions, traces and audit records before retention or compaction removes them. Freeze the agent configuration that was live.
- Eradicate. Fix the cause: tool idempotency, prompt, model pin, downstream contract, guardrail. Test the fix by replaying the captured inputs.
- Recover. Ship the fix through a canary, as in ADK Java canary deployment, then remove flags in stages and watch the detection signals for a full traffic cycle.
- Remediate users and review. Reverse wrong actions, notify affected users, and hold a blameless review within five working days.
Forensics from session events
Agent forensics answers one question: for each affected run, what did the agent see, decide and do? ADK gives you three sources. The session holds the ordered list of events for the conversation; each event records its author, its content (including function call and function response parts), and the invocation id it belongs to. Traces carry timing and the parent-child structure of agent, model and tool spans. And an audit log, if you built one as in ADK Java audit logging, holds a tamper-evident copy that survives session deletion.
The working method is to load each affected session through your session service, walk its events in order, and emit one line per model turn and tool call. Group by invocation id and you get a per-run story: user message, model reasoning output, tool call with arguments, tool response, next model turn. Patterns jump out quickly: the same tool and arguments twice in one invocation, a function response that contained an error string the model misread, or instructions inside a retrieved document that preceded an unusual tool call.
Two cautions. Session events are model input and may contain personal data, so run forensics in the same access-controlled environment as production data. And sessions can be compacted or expired by your storage policy, so preserving them is a containment step, not an afterthought.
Worked example: duplicate refunds
Consider a hypothetical incident. A support agent has an issue_refund tool that calls a payments API. At 08:55 the payments team deploys a change that occasionally returns HTTP 500 after the refund has been recorded. The tool reports an error; the model, seeing an error it believes is transient, calls the tool again, which succeeds. The user receives two refunds.
At 09:12 the model-calls-per-run p99 alert fires at four times baseline. At 09:16 triage finds issue_refund called twice in several invocations and declares SEV2. At 09:19 the commander quarantines the tool with one flag; within one 10-second poll the decision log shows quarantine decisions on all six replicas, and users are told refunds are paused. Nothing is lost by pausing: refunds queue for a human. At 09:40 forensics over sessions since 08:55 finds 37 sessions with two successful calls, and the payments team reverses the duplicates.
The eradication has two parts. The tool now sends an idempotency key derived from the invocation id and order id, so a repeated call cannot pay twice, and the error map tells the model the outcome is unknown rather than failed. The fix ships through a canary at 11:05, the quarantine is removed for 10 percent of sessions, then all, and the flag is deleted at 13:30. The review produced three actions: idempotency keys on every write tool, the write-tool-called-twice alert, and a contract test against the payments API.
Failure modes of the playbook itself
- Levers that need a deploy. If quarantining a tool means a code change, containment takes an hour. Build and test the plugin before you need it.
- Flag service as a single point of failure. Plugins that read flags synchronously stall every run when the flag service is slow. Read from memory and fail static.
- Plugin exceptions. An exception inside a callback can fail the run it was meant to protect. Wrap callback bodies and count suppressed errors.
- Ambiguous quarantine messages. The model may claim the action succeeded. State plainly that nothing was done and test the wording.
- Remote agents out of reach. A plugin on your runner does not control an agent you call over A2A. Agree kill switches with the owning team.
- Unrehearsed playbooks. Run a game day each quarter: flip each lever in staging and measure time to effect on every replica.
Trade-offs between levers
| Lever | Blast radius | User impact | Use when |
|---|---|---|---|
| Tool quarantine | One tool | One capability paused | Wrong or dangerous tool actions |
| Loop guard lowered | Runs that loop | Long tasks cut short | Runaway cost or retries |
| Model switch via delegating BaseLlm | Every model call | Possible quality change | Provider outage or bad model version |
| Agent kill switch | Whole agent | No answers | Security incident or unknown wrong actions |
| Rollback deploy | Code and prompts | Minutes of churn | Regression introduced by your own release |
What to do next
- Write the incident class table for your agent and agree severities with product and security owners.
- Build the incident control plugin with kill switch, loop guard and tool quarantine, reading flags from a fail-static in-memory snapshot.
- List every write tool, add an idempotency key to each, and alert when one runs twice in an invocation.
- Add the detection signals to your observability plugin and set baselines from a week of traffic.
- Script session forensics: load a session, group events by invocation id, print tool calls with arguments and responses.
- Rehearse: flip each lever in staging, time it, and put the measured numbers in the runbook.