Agents are asked to do things that take longer than a request should: generate a report that runs for twenty minutes, wait for a manager to approve a refund, provision infrastructure, run a batch evaluation. ADK Java's answer is the long-running function tool: the tool starts the work, returns a pending status at once, and the application resumes the conversation later by sending the tool's real result as a function response with the same call id.
The mechanics fit in a paragraph and are covered in advanced function calling. This article is about everything around them that a production system needs: where the pending call is recorded, how completions arrive and are deduplicated, how progress is reported, what happens when the work never finishes, and how to keep a resume from colliding with the user's next message. Code targets ADK Java's public APIs; where behaviour differs between releases, it says so.
What the framework gives you
Three facts from the framework shape every pattern here. First, LongRunningFunctionTool.create(...) wraps an ordinary method, with the same overloads as FunctionTool (a class and method name, an instance and method name, or a Method). The method runs like any function tool and its return value is sent to the model as the initial function response. The ADK documentation is explicit that these tools start and manage long work; they do not perform it.
Second, the event carrying the model's function call lists the ids of long-running calls in longRunningToolIds(), which in Java returns Optional<Set<String>>. That set is how your application knows a call is still open.
Third, to deliver an update or the final result, the client sends a FunctionResponse with the same id and function name as new content into the same session. The model treats it as the tool's result and continues. Intermediate responses with the same id are allowed, which is how progress works.
If your app enables resumability, read task resumability first: what a resume needs, including whether an explicit invocation id can be passed, depends on your ADK Java release.
The architecture: one table joins three facts
The tool knows the job id it created. The event stream knows the session id and the function call id. The completion, arriving minutes later from a webhook, knows only the job id. The design problem is joining those three facts durably, because the process that started the job may be gone when the result arrives.
A single table does it. Each row is one long-running call:
CREATE TABLE agent_job (
job_id TEXT PRIMARY KEY,
tool_name TEXT NOT NULL,
user_id TEXT,
session_id TEXT, -- filled by the event observer
call_id TEXT, -- filled by the event observer
state TEXT NOT NULL, -- STARTED, SUCCEEDED, FAILED, EXPIRED, RESUMED
result_json JSONB,
deadline_at TIMESTAMPTZ NOT NULL,
resumed_at TIMESTAMPTZ,
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX agent_job_ready ON agent_job (state, resumed_at)
WHERE resumed_at IS NULL;The row is ready to resume when it has a terminal state and both ids, in whichever order those arrived. That order independence matters: a fast external job can finish before the observer has recorded the call id, and the design must not lose that completion.
The tool: start, record, return
The tool does three things: validate input, create the job durably, and return a small pending map. It must be idempotent, because a model may repeat a call and a run may be retried, so derive the job id from something stable when you can.
import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.LongRunningFunctionTool;
import java.time.Duration;
import java.util.Map;
public final class ReportTools {
private final JobStore jobs; // your DAO over agent_job
private final ReportQueue queue; // Pub/Sub, SQS, a batch API...
public ReportTools(JobStore jobs, ReportQueue queue) { this.jobs = jobs; this.queue = queue; }
/** Starts a sales report build. Returns at once; the result arrives later. */
public Map<String, Object> startSalesReport(
@Schema(name = "quarter", description = "quarter such as 2026-Q3") String quarter,
@Schema(name = "region", description = "sales region code") String region) {
if (!quarter.matches("\\d{4}-Q[1-4]")) {
return Map.of("status", "error", "message", "quarter must look like 2026-Q3");
}
String jobId = jobs.createIfAbsent("report:" + quarter + ":" + region,
"startSalesReport", Duration.ofMinutes(45)); // deadline for the sweeper
queue.submit(jobId, quarter, region); // must tolerate duplicates
return Map.of("status", "pending", "jobId", jobId,
"message", "Report build started; typical time is 15-25 minutes.");
}
public LongRunningFunctionTool asTool() {
return LongRunningFunctionTool.create(this, "startSalesReport");
}
}The returned message gives the model something true to tell the user. Keep the map small: it is stored in the session and re-read by the model on every later turn. If the tool needs its own call id, ToolContext exposes functionCallId(); the observer approach below works without it and also captures the session.
Recording the pending call
Wherever your application consumes the runner's event stream, watch for the function response that answers a long-running call, and attach the session and call id to the job:
Content userMsg = Content.fromParts(Part.fromText(text));
runner.runAsync(userId, sessionId, userMsg)
.doOnNext(event -> {
Set<String> pending = event.longRunningToolIds().orElse(Set.of());
if (!pending.isEmpty()) {
openCalls.putAll(sessionId, pending); // call ids seen as long-running
}
for (FunctionResponse fr : event.functionResponses()) {
String callId = fr.id().orElse(null);
if (callId != null && openCalls.containsEntry(sessionId, callId)) {
Object jobId = fr.response().map(r -> r.get("jobId")).orElse(null);
if (jobId != null) {
jobs.attach(jobId.toString(), userId, sessionId, callId);
}
}
}
})
.blockingForEach(event -> ui.render(event));The openCalls multimap can live in memory because it is only needed between the function call event and its response in the same run; the durable record is the job row. attach should be an update that only fills empty columns, so a replayed stream cannot overwrite an earlier mapping.
Ingesting completions
Completions arrive by webhook, queue message or a poller. Treat every one as possibly duplicated or out of order. A single conditional update gives you exactly-once state transition on top of at-least-once delivery:
-- Only the first terminal result wins; retries and late duplicates change nothing.
UPDATE agent_job
SET state = :state, result_json = :result, updated_at = now()
WHERE job_id = :job_id AND state = 'STARTED';Return success to the sender whether or not a row changed, or it will retry forever. Then wake the resumer for that job.
Resuming the conversation
The resumer turns a ready row into a function response and runs the agent again. Two rules keep it safe. Claim the row before running, so two resumer instances never answer the same call. And serialise per session, because a resume that runs while the user is mid-turn interleaves two runs into one conversation history.
void resume(Job job) {
if (!jobs.claimForResume(job.jobId())) return; // sets resumed_at if null
FunctionResponse done = FunctionResponse.builder()
.id(job.callId()) // the original call id
.name(job.toolName()) // the same function name
.response(job.resultMap()) // e.g. status, url, rowCount
.build();
Content content = Content.fromParts(Part.builder().functionResponse(done).build());
sessionQueue.submit(job.sessionId(), () -> // one run per session at a time
runner.runAsync(job.userId(), job.sessionId(), content)
.blockingForEach(e -> notifier.push(job.userId(), e)));
}A per-session single-threaded executor, or a lease row in the database when you run several pods, is enough. The user sees the agent's follow-up through whatever channel you use for push: a websocket, an email, or a chat message. Cancellation and resumption explains what is persisted when a run is cut short, which matters if the resumer crashes mid-run.
Progress updates and parallel calls
Because intermediate responses with the same id are allowed, the worker can report progress: send {status: running, percent: 40} and the model can tell the user. Use this sparingly. Every update is a model call and a session event, so a job that reports every second will cost more in tokens than the report itself. A good rule is to report state changes that a user would act on, such as queued to running, or a step that needs their input, and otherwise stay silent until the end.
If the model issues two long-running calls in one turn, the event lists both ids. Each gets its own job row. When both finish close together, you can put both function responses as two parts of one Content so the model reasons about them once; otherwise resume each separately.
Deadlines and expiry
Some jobs never finish: the approver goes on holiday, the webhook is lost, the external system drops the job. Without a deadline the call stays open forever and the model keeps saying it is waiting. A sweeper running every minute marks expired rows and resumes them with an explicit status:
UPDATE agent_job SET state = 'EXPIRED',
result_json = '{"status":"expired","message":"No result within 45 minutes"}'
WHERE state = 'STARTED' AND deadline_at < now()
RETURNING job_id;The agent's instruction should say what to do with an expired result: apologise and offer to retry, or escalate. Ask the external system to cancel the work too, or a late result will arrive for a call that has already been closed; the conditional update above makes that late result a harmless no-op.
Worked example: a 19-minute report
Follow one request through the system. At 10:00:00 a user asks for the Q3 EMEA sales report. The model calls startSalesReport; the tool creates job report:2026-Q3:EMEA with a 10:45 deadline, submits it and returns pending within 40 milliseconds. The observer attaches the session and call id at 10:00:01, and the model replies that the build has started.
At 10:05 the user asks an unrelated question; that turn runs normally, and the open call sits in the history. At 10:19 the batch system posts a webhook; the conditional update moves the row to SUCCEEDED. The resumer claims it and queues a run behind the session's current turn. At 10:19:03 the model receives the response with the download URL and row count and tells the user the report is ready. At 10:21 the batch system retries the same webhook; the update matches no row and nothing happens.
Failure modes
- Blocking inside the tool. A tool that waits for the job holds a thread and the run for minutes and hits tool timeouts. Return pending.
- Wrong id or name on resume. The model cannot match the response to its call. Persist both exactly as they appeared.
- Lost completion. The result arrived before the ids were recorded and the code dropped it. Store results by job id first and resume when both halves exist.
- Double resume. Two resumers or a retried webhook answer one call twice. Claim rows atomically.
- Interleaved turns. A resume races with a user message in the same session. Serialise per session.
- Chatty progress. Frequent updates inflate context and cost. Throttle to meaningful state changes.
- Forever pending. No deadline. Add the sweeper.
Trade-offs
| Pattern | Use when | Cost |
|---|---|---|
| Blocking tool with timeout | Work reliably finishes in seconds | Holds the run; fails on slow days |
| Async tool (non-blocking, same run) | Under a minute, user waits | Run stays open; see async tools |
| Long-running tool + resume | Minutes to days, or humans involved | Job table, observer, resumer |
| Start tool + status-check tool | No push channel; user asks for updates | Model polls; more tokens |
For work under a minute, asynchronous tools are simpler. Choose the long-running pattern when the work outlives a request or depends on people.
What to do next
- List your tools whose 99th-percentile latency exceeds your run timeout; those are candidates.
- Create the job table with a deadline column and a conditional terminal update.
- Wrap one tool with
LongRunningFunctionTool.createand make it return a pending map with a job id and a user-facing message. - Add the event observer that attaches session and call ids from
longRunningToolIds(). - Build the resumer with an atomic claim and a per-session queue, and the expiry sweeper.
- Test duplicate webhooks, completion-before-attach, resume during a user turn, and expiry.