An ADK Java agent that runs for minutes, calls tools that touch real systems and waits on humans will eventually have its JVM die under it: an out-of-memory kill, a node drain, a deploy that does not wait for in-flight work. When the process comes back, nothing in it remembers what it was doing. The session event log survives if you use a persistent session service, but no component wakes up and says that invocation X was half finished. Recovering is your job, and doing it well depends on knowing exactly what a crash can leave behind.

This article is about that unplanned case. The resumability API itself (ResumabilityConfig, pausing on long-running tools, and which pieces are released versus on main) is covered in Task Resumability in ADK Java, and planned stops are covered in Cancellation and Resumption Semantics. Here we build what sits around them: a run ledger, an orphan sweeper, a side-effect journal and tests that actually kill the process. Those are application code; every ADK method used was checked against the google/adk-java sources on 2026-10-04.

What a crash destroys and what survives

The session object, the RxJava pipeline, any in-flight model request and tool threads all vanish. What remains is whatever was appended through the session service before the process died.

On current main the runner persists each event through sessionService.appendEvent(session, event) inside a concatMap, so events are appended in the order agents emit them, and a persist barrier lets the LLM flow wait until an event is stored before taking its next step. Partial events (streaming fragments) are deliberately not persisted. Check your pinned release's Runner for the barrier; older versions append in order but do not promise the flow waits. Either way, three things are never in the log:

  • A model call in progress. The request was sent, tokens were being generated, nothing was stored. Cost: one model call.
  • A tool's side effect without its response event. The tool sent the email or created the ticket, then the process died before the function response was appended. The world changed; the log says the call is unanswered.
  • Anything held only in fields of your agent or tool objects. Keep durable state in session state or your own store.

The second item is the dangerous one, and most of this article exists to handle it. An in-memory session service loses everything, so the first rule is simple: if you need crash recovery, the session service must be persistent, whether that is the Vertex AI session service or your own BaseSessionService over a database, as in storing the ADK event log in Postgres.

The four crash windows

Name the moments a crash can land, because recovery differs for each. The diagram marks four windows on one invocation.

One invocation on a timeline, and where a crash can landledger: RUNNINGlease + heartbeatuser eventappendedmodel callnothing durablecall eventappendedtool side effectoutside ADKresponse eventappendedfinal eventledger: DONEW1W2W3W4Recovery sweeperexpired leases onlyClassify from event logW1..W4 each look differentAct per classre-drive, wait, reconcile, parkSide-effect journalkeyed by business idW4 needs thisA crash at W1 to W3 loses at most a model call; a crash at W4 may leave the world changed and the log silent.
The four crash windows of one invocation. The sweeper reads the event log to tell them apart; only W4 needs the side-effect journal.

WindowLast durable eventWhat actually happenedSafe action
W1none for this invocationcrash before the user message was storedask the client to resend, or replay from the ledger's copy
W2user messagemodel was thinking; no side effectsre-drive: run the agent again for that message
W3function call, no responsetool may not have startedcheck the journal; if no entry, re-drive the tool
W4function call, no responsetool may have finished its side effectcheck the journal; reconcile, then append the result

W3 and W4 look identical in the event log: the log alone cannot tell you whether an unanswered tool call did anything. Only a record written by the tool itself can. A long-running call that paused cleanly also leaves an unanswered call, but its event carries the call id in longRunningToolIds(), and its invocation ended normally; the ledger tells the two apart.

A run ledger: knowing what was running

Before you can recover an invocation you have to know it existed. Keep a small table, outside ADK, with one row per invocation you start. The row is written before runAsync is called and finished when the last event arrives. A worker that is alive renews a lease on its rows; a row whose lease has expired belongs to a dead worker.

CREATE TABLE agent_run (
  run_id        uuid PRIMARY KEY,
  app_name      text NOT NULL,
  user_id       text NOT NULL,
  session_id    text NOT NULL,
  invocation_id text,                 -- filled from the first event we see
  request       jsonb NOT NULL,       -- the user message, so W1 can be replayed
  status        text NOT NULL,        -- RUNNING | PAUSED | DONE | FAILED | PARKED
  owner         text,                 -- worker id holding the lease
  lease_until   timestamptz,
  attempts      int NOT NULL DEFAULT 0,
  updated_at    timestamptz NOT NULL DEFAULT now()
);
CREATE INDEX agent_run_expired ON agent_run (lease_until) WHERE status = 'RUNNING';
public Flowable<Event> run(String userId, String sessionId, Content msg) {
  UUID runId = ledger.start(APP, userId, sessionId, msg, workerId, Duration.ofSeconds(60));
  Disposable heartbeat = Flowable.interval(20, TimeUnit.SECONDS)
      .subscribe(t -> ledger.renew(runId, workerId, Duration.ofSeconds(60)));
  AtomicBoolean paused = new AtomicBoolean(false);
  return runner.runAsync(userId, sessionId, msg)
      .doOnNext(e -> {
        ledger.recordInvocation(runId, e.invocationId());   // idempotent: sets once
        if (e.longRunningToolIds().map(ids -> !ids.isEmpty()).orElse(false)) paused.set(true);
      })
      .doOnComplete(() -> ledger.finish(runId, paused.get() ? "PAUSED" : "DONE"))
      .doOnError(err -> ledger.finish(runId, "FAILED"))
      .doFinally(heartbeat::dispose);
}

The lease is renewed by a timer, not by events, because a model call can be silent for a long time. The renew is conditional on owner = workerId, so a worker presumed dead cannot reclaim a reassigned run. A clean pause is recorded as PAUSED, so a waiting approval is never treated as an orphan.

The recovery sweeper and the orphan classifier

On startup, and then every minute, a sweeper claims expired RUNNING rows with UPDATE ... SET owner = me, attempts = attempts + 1, lease_until = now() + interval '60 seconds' WHERE status = 'RUNNING' AND lease_until < now() RETURNING * so that two replicas never recover the same run. For each claimed row it loads the session and classifies the tail of the log for that invocation.

enum Orphan { NOT_STARTED, MODEL_INFLIGHT, TOOL_UNANSWERED, PAUSED, ACTUALLY_DONE }

Orphan classify(RunRow row) {
  Session s = sessions.getSession(APP, row.userId(), row.sessionId(), Optional.empty())
      .blockingGet();                                   // null if the session itself is gone
  if (s == null) return Orphan.NOT_STARTED;
  // runAsync never emits the user event downstream, so a crash during the first model
  // call leaves invocation_id null. Recover it from the stored user event.
  String inv = row.invocationId() != null ? row.invocationId()
      : s.events().stream()
          .filter(e -> "user".equals(e.author()) && e.timestamp() >= row.startedAt())
          .map(Event::invocationId).findFirst().orElse(null);
  if (inv == null) return Orphan.NOT_STARTED;
  List<Event> mine = s.events().stream().filter(e -> inv.equals(e.invocationId())).toList();
  Event last = mine.get(mine.size() - 1);
  if (last.longRunningToolIds().map(ids -> !ids.isEmpty()).orElse(false)) return Orphan.PAUSED;
  // finalResponse() is also true for every sub-agent's last text: require the terminal agent.
  if (last.finalResponse() && TERMINAL_AGENT.equals(last.author())) return Orphan.ACTUALLY_DONE;
  Set<String> answered = mine.stream().flatMap(e -> e.functionResponses().stream())
      .map(r -> r.id().orElse("")).collect(Collectors.toSet());
  boolean open = mine.stream().flatMap(e -> e.functionCalls().stream())
      .anyMatch(fc -> !answered.contains(fc.id().orElse("")));
  return open ? Orphan.TOOL_UNANSWERED : Orphan.MODEL_INFLIGHT;
}

Two checks in that code came from reading the ADK source rather than guessing. runAsync appends the user event but does not emit it, so the ledger may never have seen the invocation id; the classifier recovers it from the stored user event (startedAt is the ledger row's start, in the same unit as Event.timestamp() in your version). And finalResponse() is true for any event with a pending long-running call and for the last text of every sub-agent, so it only means done when the terminal agent wrote it. A crash between that final event and the ledger update is common, and re-driving would answer the user twice; a paused run is simply marked PAUSED.

For MODEL_INFLIGHT the action is to re-drive. There is no public way to re-enter an arbitrary invocation half way through an LlmAgent's loop, so the practical re-drive is a new invocation in the same session with a short system-authored continuation message, which lets the agent see its own history and carry on. A non-resumable SequentialAgent root, however, starts again from its first sub-agent, so earlier steps must be idempotent or skip themselves when state shows they finished. On main, a resumable app can also call the runAsync overload that takes an explicit invocationId; the overview article explains when that path is available. For TOOL_UNANSWERED, the next section decides.

Side effects across the gap: a journal keyed by business identity

A tool that changes the outside world should write down what it is about to do before doing it, and what happened after. That record is the only evidence that separates W3 from W4. The key must be the business identity of the effect, not the function call id: when the agent is re-driven, the model emits a new call with a new id for the same intent, and a call-id key would let the second email go out. The idempotency article covers key design in general; this is the crash-specific shape.

public static Map<String, Object> sendInvoice(
    @Schema(name = "invoiceId") String invoiceId, ToolContext ctx) {
  String key = "send-invoice:" + invoiceId;                 // business identity, not the call id
  Journal.Entry prior = journal.find(key);
  if (prior != null && prior.isDone()) return prior.result(); // replayed result, no second send
  if (prior != null && prior.isIntent()) {                    // W4 candidate: did the send happen?
    Optional<Map<String, Object>> seen = billing.lookupByIdempotencyKey(key);
    if (seen.isPresent()) { journal.done(key, seen.get()); return seen.get(); }
  }
  journal.intent(key, ctx.invocationId(), ctx.functionCallId().orElse(""));
  Map<String, Object> result = billing.send(invoiceId, /* idempotencyKey= */ key);
  journal.done(key, result);
  return result;
}

On recovery there are three outcomes. A done entry means the effect happened: return the stored result. An intent only means ask the downstream system whether it saw the key, which requires an API with idempotency keys or a lookup. No entry means the tool never started and is safe to run. If the downstream system offers neither key nor lookup, park the run and ask a human; guessing either way is worse than waiting.

Crash loops and poison invocations

Some crashes are caused by the work itself, such as a tool that allocates a huge buffer. Re-driving such a run kills the next worker too, and a sweeper with no limit turns one bad input into a rolling outage across replicas.

Increment attempts in the claiming statement, as above, and back off exponentially between attempts. After a small fixed number, say three, move the row to PARKED and copy the request, the classification and the last events into a dead-letter record, as in the ADK Java dead-letter queue. Alert on parked runs, not on recoveries: recoveries are expected after every deploy, parked runs need a person. Also cap how many orphans one worker recovers at once, so a fleet restart does not stampede the model provider.

Worked example: killing an invoicing flow four times

Take a three-step flow: a draft agent writes an invoice summary, a billing agent calls sendInvoice, and a notify agent tells the customer. Kill the JVM with kill -9 at each window and watch recovery.

  1. Kill during the draft model call (W2). The log holds the user message only. The sweeper classifies MODEL_INFLIGHT, re-drives, and the draft is written once. Cost: one wasted model call.
  2. Kill after the billing call event, before journal.intent (W3). The log has an unanswered sendInvoice call; the journal has nothing for send-invoice:INV-77. Re-drive; the new call sends once.
  3. Kill after billing.send returned, before journal.done (W4). The journal holds an intent. On re-drive the tool asks billing for the key, finds the invoice, records done and returns it. Exactly one invoice.
  4. Kill after the notify agent's final event, before ledger.finish. The sweeper sees finalResponse() authored by the terminal agent and marks the run done. One notification. Had the kill come after the draft agent's final text instead, the author check would correctly classify it as unfinished.

Key the journal by function call id instead and case three sends a second invoice, because the re-driven call has a new id.

Testing recovery by actually crashing

Crash recovery that has never been exercised does not work. Put named crash points in the code paths above, controlled by a test-only flag, that call Runtime.getRuntime().halt(137): halt skips shutdown hooks, which is what a real kill does, whereas System.exit runs them and hides bugs. Run the agent in a child JVM against a real database, halt it at each point, start a fresh JVM, run the sweeper and assert on the outside world: count invoices in a fake billing service, count notifications, check the ledger reaches a terminal state. Use a scripted fake model so runs are deterministic, and add a staging chaos job that kills random workers under load.

Trade-offs

  • Lease length. Short leases recover quickly but may declare a slow worker dead. The owner check and the journal contain the damage; keep leases around three heartbeats.
  • Re-drive versus ask. Re-drive cheap, read-only or journaled steps; park anything without an idempotency story.
  • At-least-once is the ceiling. Model calls may repeat. What you get is at-most-once external effects for tools that cooperate, which is what users notice.

What to do next

  1. Confirm your session service is persistent and survives a restart of every replica; an in-memory service makes everything else moot.
  2. Add the agent_run ledger and wrap every runAsync call with start, heartbeat and finish.
  3. Write the sweeper with a claiming update, the four-way classifier and the ACTUALLY_DONE check first.
  4. List every tool with an external side effect; give each a journal entry keyed by business identity and pass that key downstream.
  5. Add an attempts cap, backoff and a dead-letter path; alert on parked runs.
  6. Build the halt-at-crash-point test for each window and run it in CI; add a staging chaos job.
  7. Read the runtime execution loop to see exactly where events are emitted in your version.
Key takeaway: A crash leaves only the persisted event log, and the log cannot tell an unanswered tool call that did nothing from one that changed the world. Keep a leased run ledger so you know what was running, recover with a sweeper that classifies each orphan from the log and checks for an already-final response, and make every side-effecting tool journal its intent and result under a business key, because a re-driven model call gets a new call id. Cap attempts, park poison runs, and prove it all with tests that halt the JVM at each window.