An ADK Java agent that runs for minutes, calls tools that touch real systems and waits on humans will eventually have its JVM die under it: an out-of-memory kill, a node drain, a deploy that does not wait for in-flight work. When the process comes back, nothing in it remembers what it was doing. The session event log survives if you use a persistent session service, but no component wakes up and says that invocation X was half finished. Recovering is your job, and doing it well depends on knowing exactly what a crash can leave behind.
This article is about that unplanned case. The resumability API itself (ResumabilityConfig, pausing on long-running tools, and which pieces are released versus on main) is covered in Task Resumability in ADK Java, and planned stops are covered in Cancellation and Resumption Semantics. Here we build what sits around them: a run ledger, an orphan sweeper, a side-effect journal and tests that actually kill the process. Those are application code; every ADK method used was checked against the google/adk-java sources on 2026-10-04.
What a crash destroys and what survives
The session object, the RxJava pipeline, any in-flight model request and tool threads all vanish. What remains is whatever was appended through the session service before the process died.
On current main the runner persists each event through sessionService.appendEvent(session, event) inside a concatMap, so events are appended in the order agents emit them, and a persist barrier lets the LLM flow wait until an event is stored before taking its next step. Partial events (streaming fragments) are deliberately not persisted. Check your pinned release's Runner for the barrier; older versions append in order but do not promise the flow waits. Either way, three things are never in the log:
- A model call in progress. The request was sent, tokens were being generated, nothing was stored. Cost: one model call.
- A tool's side effect without its response event. The tool sent the email or created the ticket, then the process died before the function response was appended. The world changed; the log says the call is unanswered.
- Anything held only in fields of your agent or tool objects. Keep durable state in session state or your own store.
The second item is the dangerous one, and most of this article exists to handle it. An in-memory session service loses everything, so the first rule is simple: if you need crash recovery, the session service must be persistent, whether that is the Vertex AI session service or your own BaseSessionService over a database, as in storing the ADK event log in Postgres.
The four crash windows
Name the moments a crash can land, because recovery differs for each. The diagram marks four windows on one invocation.
| Window | Last durable event | What actually happened | Safe action |
|---|---|---|---|
| W1 | none for this invocation | crash before the user message was stored | ask the client to resend, or replay from the ledger's copy |
| W2 | user message | model was thinking; no side effects | re-drive: run the agent again for that message |
| W3 | function call, no response | tool may not have started | check the journal; if no entry, re-drive the tool |
| W4 | function call, no response | tool may have finished its side effect | check the journal; reconcile, then append the result |
W3 and W4 look identical in the event log: the log alone cannot tell you whether an unanswered tool call did anything. Only a record written by the tool itself can. A long-running call that paused cleanly also leaves an unanswered call, but its event carries the call id in longRunningToolIds(), and its invocation ended normally; the ledger tells the two apart.
A run ledger: knowing what was running
Before you can recover an invocation you have to know it existed. Keep a small table, outside ADK, with one row per invocation you start. The row is written before runAsync is called and finished when the last event arrives. A worker that is alive renews a lease on its rows; a row whose lease has expired belongs to a dead worker.
CREATE TABLE agent_run (
run_id uuid PRIMARY KEY,
app_name text NOT NULL,
user_id text NOT NULL,
session_id text NOT NULL,
invocation_id text, -- filled from the first event we see
request jsonb NOT NULL, -- the user message, so W1 can be replayed
status text NOT NULL, -- RUNNING | PAUSED | DONE | FAILED | PARKED
owner text, -- worker id holding the lease
lease_until timestamptz,
attempts int NOT NULL DEFAULT 0,
updated_at timestamptz NOT NULL DEFAULT now()
);
CREATE INDEX agent_run_expired ON agent_run (lease_until) WHERE status = 'RUNNING';public Flowable<Event> run(String userId, String sessionId, Content msg) {
UUID runId = ledger.start(APP, userId, sessionId, msg, workerId, Duration.ofSeconds(60));
Disposable heartbeat = Flowable.interval(20, TimeUnit.SECONDS)
.subscribe(t -> ledger.renew(runId, workerId, Duration.ofSeconds(60)));
AtomicBoolean paused = new AtomicBoolean(false);
return runner.runAsync(userId, sessionId, msg)
.doOnNext(e -> {
ledger.recordInvocation(runId, e.invocationId()); // idempotent: sets once
if (e.longRunningToolIds().map(ids -> !ids.isEmpty()).orElse(false)) paused.set(true);
})
.doOnComplete(() -> ledger.finish(runId, paused.get() ? "PAUSED" : "DONE"))
.doOnError(err -> ledger.finish(runId, "FAILED"))
.doFinally(heartbeat::dispose);
}The lease is renewed by a timer, not by events, because a model call can be silent for a long time. The renew is conditional on owner = workerId, so a worker presumed dead cannot reclaim a reassigned run. A clean pause is recorded as PAUSED, so a waiting approval is never treated as an orphan.
The recovery sweeper and the orphan classifier
On startup, and then every minute, a sweeper claims expired RUNNING rows with UPDATE ... SET owner = me, attempts = attempts + 1, lease_until = now() + interval '60 seconds' WHERE status = 'RUNNING' AND lease_until < now() RETURNING * so that two replicas never recover the same run. For each claimed row it loads the session and classifies the tail of the log for that invocation.
enum Orphan { NOT_STARTED, MODEL_INFLIGHT, TOOL_UNANSWERED, PAUSED, ACTUALLY_DONE }
Orphan classify(RunRow row) {
Session s = sessions.getSession(APP, row.userId(), row.sessionId(), Optional.empty())
.blockingGet(); // null if the session itself is gone
if (s == null) return Orphan.NOT_STARTED;
// runAsync never emits the user event downstream, so a crash during the first model
// call leaves invocation_id null. Recover it from the stored user event.
String inv = row.invocationId() != null ? row.invocationId()
: s.events().stream()
.filter(e -> "user".equals(e.author()) && e.timestamp() >= row.startedAt())
.map(Event::invocationId).findFirst().orElse(null);
if (inv == null) return Orphan.NOT_STARTED;
List<Event> mine = s.events().stream().filter(e -> inv.equals(e.invocationId())).toList();
Event last = mine.get(mine.size() - 1);
if (last.longRunningToolIds().map(ids -> !ids.isEmpty()).orElse(false)) return Orphan.PAUSED;
// finalResponse() is also true for every sub-agent's last text: require the terminal agent.
if (last.finalResponse() && TERMINAL_AGENT.equals(last.author())) return Orphan.ACTUALLY_DONE;
Set<String> answered = mine.stream().flatMap(e -> e.functionResponses().stream())
.map(r -> r.id().orElse("")).collect(Collectors.toSet());
boolean open = mine.stream().flatMap(e -> e.functionCalls().stream())
.anyMatch(fc -> !answered.contains(fc.id().orElse("")));
return open ? Orphan.TOOL_UNANSWERED : Orphan.MODEL_INFLIGHT;
}Two checks in that code came from reading the ADK source rather than guessing. runAsync appends the user event but does not emit it, so the ledger may never have seen the invocation id; the classifier recovers it from the stored user event (startedAt is the ledger row's start, in the same unit as Event.timestamp() in your version). And finalResponse() is true for any event with a pending long-running call and for the last text of every sub-agent, so it only means done when the terminal agent wrote it. A crash between that final event and the ledger update is common, and re-driving would answer the user twice; a paused run is simply marked PAUSED.
For MODEL_INFLIGHT the action is to re-drive. There is no public way to re-enter an arbitrary invocation half way through an LlmAgent's loop, so the practical re-drive is a new invocation in the same session with a short system-authored continuation message, which lets the agent see its own history and carry on. A non-resumable SequentialAgent root, however, starts again from its first sub-agent, so earlier steps must be idempotent or skip themselves when state shows they finished. On main, a resumable app can also call the runAsync overload that takes an explicit invocationId; the overview article explains when that path is available. For TOOL_UNANSWERED, the next section decides.
Side effects across the gap: a journal keyed by business identity
A tool that changes the outside world should write down what it is about to do before doing it, and what happened after. That record is the only evidence that separates W3 from W4. The key must be the business identity of the effect, not the function call id: when the agent is re-driven, the model emits a new call with a new id for the same intent, and a call-id key would let the second email go out. The idempotency article covers key design in general; this is the crash-specific shape.
public static Map<String, Object> sendInvoice(
@Schema(name = "invoiceId") String invoiceId, ToolContext ctx) {
String key = "send-invoice:" + invoiceId; // business identity, not the call id
Journal.Entry prior = journal.find(key);
if (prior != null && prior.isDone()) return prior.result(); // replayed result, no second send
if (prior != null && prior.isIntent()) { // W4 candidate: did the send happen?
Optional<Map<String, Object>> seen = billing.lookupByIdempotencyKey(key);
if (seen.isPresent()) { journal.done(key, seen.get()); return seen.get(); }
}
journal.intent(key, ctx.invocationId(), ctx.functionCallId().orElse(""));
Map<String, Object> result = billing.send(invoiceId, /* idempotencyKey= */ key);
journal.done(key, result);
return result;
}On recovery there are three outcomes. A done entry means the effect happened: return the stored result. An intent only means ask the downstream system whether it saw the key, which requires an API with idempotency keys or a lookup. No entry means the tool never started and is safe to run. If the downstream system offers neither key nor lookup, park the run and ask a human; guessing either way is worse than waiting.
Crash loops and poison invocations
Some crashes are caused by the work itself, such as a tool that allocates a huge buffer. Re-driving such a run kills the next worker too, and a sweeper with no limit turns one bad input into a rolling outage across replicas.
Increment attempts in the claiming statement, as above, and back off exponentially between attempts. After a small fixed number, say three, move the row to PARKED and copy the request, the classification and the last events into a dead-letter record, as in the ADK Java dead-letter queue. Alert on parked runs, not on recoveries: recoveries are expected after every deploy, parked runs need a person. Also cap how many orphans one worker recovers at once, so a fleet restart does not stampede the model provider.
Worked example: killing an invoicing flow four times
Take a three-step flow: a draft agent writes an invoice summary, a billing agent calls sendInvoice, and a notify agent tells the customer. Kill the JVM with kill -9 at each window and watch recovery.
- Kill during the draft model call (W2). The log holds the user message only. The sweeper classifies
MODEL_INFLIGHT, re-drives, and the draft is written once. Cost: one wasted model call. - Kill after the billing call event, before
journal.intent(W3). The log has an unansweredsendInvoicecall; the journal has nothing forsend-invoice:INV-77. Re-drive; the new call sends once. - Kill after
billing.sendreturned, beforejournal.done(W4). The journal holds an intent. On re-drive the tool asks billing for the key, finds the invoice, records done and returns it. Exactly one invoice. - Kill after the notify agent's final event, before
ledger.finish. The sweeper seesfinalResponse()authored by the terminal agent and marks the run done. One notification. Had the kill come after the draft agent's final text instead, the author check would correctly classify it as unfinished.
Key the journal by function call id instead and case three sends a second invoice, because the re-driven call has a new id.
Testing recovery by actually crashing
Crash recovery that has never been exercised does not work. Put named crash points in the code paths above, controlled by a test-only flag, that call Runtime.getRuntime().halt(137): halt skips shutdown hooks, which is what a real kill does, whereas System.exit runs them and hides bugs. Run the agent in a child JVM against a real database, halt it at each point, start a fresh JVM, run the sweeper and assert on the outside world: count invoices in a fake billing service, count notifications, check the ledger reaches a terminal state. Use a scripted fake model so runs are deterministic, and add a staging chaos job that kills random workers under load.
Trade-offs
- Lease length. Short leases recover quickly but may declare a slow worker dead. The owner check and the journal contain the damage; keep leases around three heartbeats.
- Re-drive versus ask. Re-drive cheap, read-only or journaled steps; park anything without an idempotency story.
- At-least-once is the ceiling. Model calls may repeat. What you get is at-most-once external effects for tools that cooperate, which is what users notice.
What to do next
- Confirm your session service is persistent and survives a restart of every replica; an in-memory service makes everything else moot.
- Add the
agent_runledger and wrap everyrunAsynccall with start, heartbeat and finish. - Write the sweeper with a claiming update, the four-way classifier and the
ACTUALLY_DONEcheck first. - List every tool with an external side effect; give each a journal entry keyed by business identity and pass that key downstream.
- Add an attempts cap, backoff and a dead-letter path; alert on parked runs.
- Build the halt-at-crash-point test for each window and run it in CI; add a staging chaos job.
- Read the runtime execution loop to see exactly where events are emitted in your version.