An agent workflow that waits for a human, a batch job or a payment provider cannot hold a thread open for two days, and it cannot assume the process that started the work is the one that finishes it. Resumability is ADK's answer: the workflow pauses at a well-defined point, the request ends, and a later request picks up from where it stopped instead of re-running everything from the first agent.
ADK Java's support is real but moving quickly, and the released version and the main branch differ in important ways. This article explains the model from first principles, shows working code for the released behaviour, describes what the main branch adds, and covers the operational rules that apply to every version: persistent sessions, idempotent tools and not changing the workflow while it is paused. Every class and method named here was checked against the google/adk-java repository on 2026-10-02 (release v1.10.1 of 2026-09-18 and the main branch). Treat anything marked main-only as subject to change.
The model: invocations, events and a pause point
In ADK, one call to the runner is an invocation. The runner looks up the session, appends your message as an event, runs the root agent, and streams back every event the agents produce: model text, function calls, function responses, state deltas. The session service persists those events. The event log is the durable record of what happened, and it is the raw material for resuming (the runtime execution loop walks through one invocation in detail).
A long-running tool is a function tool whose result is not available when the function returns. In Java it is a LongRunningFunctionTool: the method starts the work (opens an approval ticket, submits a job) and returns something like a pending status. The event carrying that function call lists the call's id in longRunningToolIds. With resumability enabled, that marker is the pause point: the workflow stops instead of continuing to the next agent, and the invocation ends.
Resuming means sending, in a new request, the function response that answers the paused call. The runner must then work out which agent should continue and which agents already finished, so it does not re-run them. How it works that out is exactly where the released version and main differ.
What is released and what is on main
In v1.10.1, ResumabilityConfig exists with a resumable flag, and both the class and App.Builder.resumabilityConfig are annotated @Deprecated with a note that says it is a partial feature: only event-reconstruction pause and resume for SequentialAgent is implemented, and full session resumability (persisted agent state, durable resume, other workflow agents) is not yet available. The deprecation is a warning about completeness, not a removal notice; the note says the same config will drive full resumability.
On main, the annotation is @Experimental instead, and the javadoc describes checkpointing agent state as the workflow runs, with resume being best-effort and at-least-once. The v1.10.0 changelog already lists EventActions.agentState for session-resumability checkpoints, so the field ships in the release even though the released SequentialAgent does not write it.
| Capability | v1.10.1 (released) | main (unreleased, 2026-10-02) |
|---|---|---|
| ResumabilityConfig.resumable | yes, @Deprecated (partial) | yes, @Experimental |
| SequentialAgent pause on long-running call | yes, resume point rebuilt from events | yes, checkpoint before each sub-agent |
| LoopAgent, ParallelAgent resume | LoopAgent stops on a pending call but cannot resume into the paused iteration; ParallelAgent has no resumable path | yes: loop index and count checkpointed, finished parallel branches skipped |
| EventActions.agentState / endOfAgent | fields exist | written as checkpoints and rehydrated on resume |
| runAsync with an explicit invocationId | no | yes, @Experimental |
| plainTextContinuationAutoResume | builder flag present | @Deprecated legacy flow, mutually exclusive with resumable |
The ADK documentation's resume page (adk.dev, runtime/resume), read on the same date, lists Python v1.16.0 and Kotlin v0.1.0 as supported and does not mention Java. Read that as documentation lagging the code, not as Java lacking the feature, and pin your ADK version deliberately.
Enabling it and building a pausable workflow
Resumability is configured once on the App and applies to every agent in it. The example is a three-step publishing flow: draft, ask a human for approval, publish. The approval step uses a long-running tool. The model names are placeholders; use whatever your project already runs.
import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.SequentialAgent;
import com.google.adk.apps.App;
import com.google.adk.apps.ResumabilityConfig;
import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.LongRunningFunctionTool;
import java.util.Map;
public final class PublishFlow {
/** Starts a human approval and returns immediately; the answer arrives later. */
public static Map<String, Object> requestApproval(
@Schema(name = "draftId", description = "id of the draft to approve") String draftId) {
String ticket = ApprovalQueue.open(draftId); // your system; must be idempotent
return Map.of("status", "pending", "ticket", ticket);
}
public static App build() {
LlmAgent draft = LlmAgent.builder()
.name("draft_agent").model("gemini-2.5-flash")
.instruction("Write a release note draft and save it with an id.")
.build();
LlmAgent approval = LlmAgent.builder()
.name("approval_agent").model("gemini-2.5-flash")
.instruction("Ask for approval of the draft using requestApproval, then report the decision.")
.tools(LongRunningFunctionTool.create(PublishFlow.class, "requestApproval"))
.build();
LlmAgent publish = LlmAgent.builder()
.name("publish_agent").model("gemini-2.5-flash")
.instruction("If the draft was approved, publish it; otherwise explain why not.")
.build();
SequentialAgent flow = SequentialAgent.builder()
.name("publish_flow").subAgents(draft, approval, publish).build();
return App.builder()
.name("publish_app")
.rootAgent(flow)
.resumabilityConfig(ResumabilityConfig.builder().resumable(true).build())
.build();
}
}Two things in that code are deliberate. The tool returns immediately with a pending status and a ticket; it does not block waiting for the human. And opening the ticket must itself be idempotent, for reasons the at-least-once section explains. Building with the release, the resumabilityConfig call produces a deprecation warning; suppress it locally with a comment explaining why, so the next reader knows it was a conscious choice.
Pausing and resuming on the released version
Run the flow. draft_agent produces its draft, approval_agent calls requestApproval, and the event carrying that call lists its id in longRunningToolIds. The resumable SequentialAgent sees the pending long-running call and stops; publish_agent does not run. Your code should record the pending call id next to the ticket, because that id is how the answer finds its way back.
import com.google.adk.artifacts.InMemoryArtifactService;
import com.google.adk.events.Event;
import com.google.adk.runner.Runner;
import com.google.genai.types.Content;
import com.google.genai.types.FunctionCall;
import com.google.genai.types.FunctionResponse;
import com.google.genai.types.Part;
import java.util.Map;
Runner runner = Runner.builder()
.app(PublishFlow.build())
.sessionService(persistentSessionService) // NOT in-memory if you must survive restarts
.artifactService(new InMemoryArtifactService())
.build();
// Request 1: run until the long-running call pauses the sequence.
String pendingCallId = null;
for (Event e : runner.runAsync(userId, sessionId,
Content.fromParts(Part.fromText("Prepare the 2.4 release note"))).blockingIterable()) {
for (FunctionCall call : e.functionCalls()) {
if (e.longRunningToolIds().map(ids -> ids.contains(call.id().orElse(""))).orElse(false)) {
pendingCallId = call.id().orElseThrow(); // persist this with the ticket
}
}
}
// Request 2, possibly in another process days later: answer the paused call.
Part answer = Part.builder().functionResponse(
FunctionResponse.builder()
.id(pendingCallId)
.name("requestApproval")
.response(Map.of("status", "approved", "approver", "alice"))
.build())
.build();
runner.runAsync(userId, sessionId, Content.fromParts(answer))
.blockingForEach(e -> log(e)); // approval_agent resumes, then publish_agentWhen the second request arrives, the runner looks back through the session for the function call that the new function response answers, and finds the agent that authored it. With resumability on, v1.10.1 routes the request to that agent's top-most SequentialAgent ancestor rather than straight to the agent, so the sequence can continue past it. The SequentialAgent then rebuilds its resume point from the session events: it finds which direct sub-agent's subtree made the call and starts from that index. draft_agent is not re-run; approval_agent sees the approval and reports it; publish_agent runs next.
Without resumability, the same function response would be routed straight to approval_agent, which would answer and stop, and publish_agent would never run. That difference is the clearest test that your configuration took effect (the sequential chain pattern covers the non-resumable baseline).
What main adds: durable checkpoints
Rebuilding the resume point from events works for a simple sequence, but it cannot capture state that never appears in an event, such as how many times a loop has run. On main, workflow agents write explicit checkpoints. A SequentialAgent writes an event whose agentState records current_sub_agent before each sub-agent runs; a LoopAgent records current_sub_agent and times_looped; a ParallelAgent checkpoints its start and skips branches already marked finished. An agent that completes is marked with endOfAgent, which drops its stored state.
On resume, the invocation context rehydrates those checkpoints from the invocation's events, each workflow agent fast-forwards to its checkpoint, and an invocation whose active agent already finished resolves to a no-op. Main also adds an experimental runAsync overload that takes the invocation id explicitly:
// Unreleased, main branch only, marked @Experimental (checked 2026-10-02).
// Resume a specific invocation; invocationId may be null when it can be inferred
// from a function response carried by the message.
runner.runAsync(
userId,
sessionId,
invocationId, // the paused invocation, from Event.invocationId()
Content.fromParts(answer), // optional: may be null to just continue
RunConfig.builder().build(),
/* stateDelta= */ null)
.blockingForEach(e -> log(e));
// Throws IllegalStateException if the App is not configured with resumable(true).Two caveats come from the main-branch javadoc itself. Resume is best-effort and at-least-once. And in-memory state is lost on resumption: anything an agent keeps in a Java field rather than in session state or a checkpoint is gone after the pause. The older plainTextContinuationAutoResume flag, which made any plain user message resume the last unfinished invocation, is deprecated on main and cannot be combined with resumable; the builder rejects that combination.
At-least-once: why every tool must be idempotent
A crash can land between a tool doing its work and the event recording that work being persisted. On resume, the framework cannot know the work happened, so it runs the step again. The adk.dev resume page states the contract plainly: tools run at least once and may run more than once. For read-only tools that is harmless. For a tool that charges a card, sends an email or opens a ticket it is a duplicate side effect.
The fix is an idempotency key derived from business data, checked against a durable store before acting, and passed to downstream APIs that accept one:
public static Map<String, Object> chargeCard(
@Schema(name = "orderId") String orderId,
@Schema(name = "amountCents") long amountCents,
ToolContext ctx) {
// The key must be stable across retries and resumes: derive it from business data,
// never from a random UUID generated inside the tool.
String key = "charge:" + orderId;
Optional<Receipt> done = receipts.find(key); // durable store, not a field
if (done.isPresent()) {
return Map.of("status", "ok", "receipt", done.get().id(), "replayed", true);
}
Receipt r = payments.charge(orderId, amountCents, /* idempotencyKey= */ key);
receipts.save(key, r);
return Map.of("status", "ok", "receipt", r.id(), "replayed", false);
}The key must be the same on every retry, which is why it comes from the order id and not from a UUID generated inside the tool. The receipt store must be durable and outside the agent process. Idempotency in ADK Java covers the pattern in more depth, including keys for tools whose arguments are generated by the model.
Operating resumable workflows
- Use a persistent session service. The in-memory session service loses the event log when the process exits, and with it everything resumption needs. Persist events in a database; storing ADK events in Postgres shows one layout.
- Persist the pending call id with your external ticket. The webhook or UI that receives the human's answer needs the session id, user id and function call id to build the function response.
- Do not change the workflow while instances are paused. The adk.dev page lists this as a limitation. Renaming an agent or reordering sub-agents breaks the mapping from recorded authors and checkpoint names to code. Version workflows and drain paused instances before deploying structural changes.
- Keep agent memory in session state. Values in Java fields do not survive the pause. Use state deltas, which are events and therefore durable.
- Custom agents need work. Only the built-in workflow agents implement the resume logic. A custom BaseAgent subclass has to read and write its own checkpoints to resume correctly.
- Expire stale pauses. Decide how long an approval may stay pending, and send a function response with a timeout or rejection status when it expires, so the workflow ends deliberately rather than never.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Function response is answered but the next agent never runs | resumability not enabled, or the runner built from an agent instead of the App | build the Runner with app(...) and resumable(true) |
| Earlier agents run again after resume | workflow agent without resume support, or a custom agent | use SequentialAgent on the release; add checkpoints to custom agents |
| Resume finds nothing to continue | session lost (in-memory service) or wrong session id | persistent session service; store ids with the ticket |
| Duplicate charges, emails or tickets | at-least-once re-execution after a crash | idempotency keys and a durable receipt store |
| Resume errors after a deploy | agents renamed or reordered while paused | version workflows; drain before structural changes |
| Builder throws on build() (main) | resumable and plainTextContinuationAutoResume both set | set only resumable |
What to do next
- Pin your ADK Java version and read
ResumabilityConfigandRunnerin that exact version's source; the behaviour changes between releases. - Build the three-agent flow above with a persistent session service and confirm that publish_agent runs after the function response, and does not without resumable(true).
- Kill the process between the two requests and resume from a fresh one to prove the pause really is durable.
- Audit every tool reachable from a resumable workflow for side effects, and add idempotency keys to each one that has them.
- Write down your policy for stale pauses and structural deploys before the first real paused instance exists.