"Workflow" means two different things to an ADK Java developer. Inside one process, ADK has workflow agents, SequentialAgent, ParallelAgent and LoopAgent, which run sub-agents in a fixed pattern within a single invocation. Outside the process, Google Cloud Workflows is a managed orchestrator that runs a YAML-defined sequence of HTTP calls, with retries, parallel branches and waits, for up to a year per execution. This page is about the second: using Cloud Workflows to drive ADK Java agents deployed on Cloud Run, so a multi-step agent job survives restarts, waits days for a human, and never pays twice for a step it has already finished.
You will see the architecture, the limits that shape it, a step endpoint in Java that makes agent calls safe to retry, a workflow definition with parallel research, a custom retry predicate and a human-approval callback, a traced execution, and the failure modes. The generic theory of durable workflows is in the workflow orchestration architecture article; here every detail is specific to Cloud Workflows and ADK Java.
Two kinds of workflow, and when to use each
Use ADK's workflow agents when the whole job finishes in one request: a pipeline of three agents that takes forty seconds, where losing it to an instance restart just means the caller retries. Use Cloud Workflows when the job is longer than a request can live, waits on people or external systems, mixes agent steps with ordinary API calls, or must not repeat an expensive step after a crash. The two compose: a single Cloud Workflows step can call an endpoint that runs a SequentialAgent internally.
The division of labour is strict. Cloud Workflows owns control flow, retries, timers, waits and the record of which steps finished. The ADK Java service owns reasoning: it runs one agent invocation per step and keeps the conversation in a durable session. Neither should try to do the other's job; an agent that loops forever waiting for approval inside a Cloud Run request is the classic mistake.
Architecture and the limits that shape it
A handful of documented limits drive the design. Check them against the current Workflows quotas page before you build, since quotas change:
| Limit | Value | Design consequence |
|---|---|---|
| Execution duration | 1 year | Multi-day approvals are fine inside one execution |
| HTTP request timeout | 300 s default, 1800 s max | Agent steps longer than 30 minutes need callbacks |
| HTTP response size | 2 MB | Return the final answer, not the event history |
| Variables, arguments and events | 512 KB per execution | Store large outputs in GCS or session state; pass references |
| Steps per execution | 100,000 | Loops over thousands of items need batching |
| Parallel branches | 10 per step; 20 concurrent branches and iterations | Use concurrency_limit to protect model quotas |
| Callback wait | 43200 s default timeout | Set an explicit timeout and handle TimeoutError |
An idempotent step endpoint in Java
Cloud Workflows retries a step by sending the same HTTP request again, and it cannot know whether the first attempt reached the agent. So the step endpoint has to make repeats harmless. The design below derives the ADK session id from the workflow execution id plus the step name, and uses the agent's outputKey so the final answer lands in session state. A repeat of a finished step finds the stored answer and returns it without calling the model. A repeat that arrives while the first delivery is still running gets 409, which the workflow retries after a backoff.
@RestController
public class StepController {
private static final String APP = "research_app";
private static final String USER = "workflow";
private final Runner runner;
public StepController(BaseSessionService durableSessions) { // not InMemorySessionService
LlmAgent researcher = LlmAgent.builder()
.name("researcher")
.model("gemini-2.5-flash")
.instruction("Research the topic in the user message. Answer in under 200 words.")
.outputKey("step_result") // final text is saved to session state
.build();
this.runner = Runner.builder()
.agent(researcher)
.appName(APP)
.sessionService(durableSessions)
.artifactService(new InMemoryArtifactService())
.build();
}
public record StepRequest(String executionId, String step, String input) {}
public record StepResult(String text, boolean replayed) {}
@PostMapping("/steps/research")
public StepResult research(@RequestBody StepRequest req) {
String sessionId = req.executionId() + "-" + req.step();
BaseSessionService sessions = runner.sessionService();
Session s = sessions.getSession(APP, USER, sessionId, Optional.empty()).blockingGet();
if (s != null && s.state().containsKey("step_result")) {
return new StepResult(String.valueOf(s.state().get("step_result")), true); // replay
}
if (s != null) { // a delivery of this step is still running
throw new ResponseStatusException(HttpStatus.CONFLICT, "step in progress");
}
sessions.createSession(APP, USER, new ConcurrentHashMap<>(), sessionId).blockingGet();
Content msg = Content.fromParts(Part.fromText(req.input()));
StringBuilder text = new StringBuilder();
runner.runAsync(USER, sessionId, msg)
.filter(Event::finalResponse)
.blockingForEach(e -> text.append(e.stringifyContent()));
return new StepResult(text.toString(), false);
}
}The class names come from the google/adk-java core module: Runner.builder() requires an artifact service, getSession returns an RxJava Maybe (so blockingGet() is null when the session does not exist), and Event.finalResponse() marks the answer. Inject a durable BaseSessionService: Cloud Run scales instances away, and an in-memory store forgets every finished step with them. Two refinements belong in production. Record a start time in the initial state so a session left behind by a crashed instance can be detected as stale and restarted, and map transient model errors to 503 so the retry policy can tell them from bugs.
The workflow definition
The workflow below researches several topics in parallel, at most three at a time to protect the model quota, then asks the agent service to draft a report and hand it to a reviewer, and waits up to two days for the reviewer's decision. The auth: type: OIDC block makes Workflows attach an identity token for its service account, so the Cloud Run service can require authentication. The audience defaults to the full request URL, path included; setting it to the service base URL is the safe choice for Cloud Run. Raise the Cloud Run service's request timeout to at least the workflow's timeout: value: Cloud Run defaults to 300 seconds (maximum 3,600), and a service that cuts the agent off at 300 seconds leaves a session with no result, so every retry would get 409.
main:
params: [args]
steps:
- init:
assign:
- exec_id: ${sys.get_env("GOOGLE_CLOUD_WORKFLOW_EXECUTION_ID")}
- base: ${args.agentUrl}
- notes: {}
- research:
parallel:
shared: [notes]
concurrency_limit: 3
for:
value: topic
in: ${args.topics}
steps:
- call_agent:
try:
call: http.post
args:
url: ${base + "/steps/research"}
auth:
type: OIDC
audience: ${base}
timeout: 900
body:
executionId: ${exec_id}
step: ${"research-" + topic}
input: ${topic}
result: resp
retry:
predicate: ${agent_retryable}
max_retries: 6
backoff:
initial_delay: 2
max_delay: 60
multiplier: 2
- keep:
assign:
- notes[topic]: ${resp.body.text}
- make_callback:
call: events.create_callback_endpoint
args:
http_callback_method: "POST"
result: cb
- request_review:
call: http.post
args:
url: ${base + "/steps/draft"}
auth:
type: OIDC
audience: ${base}
body:
executionId: ${exec_id}
step: "draft"
input: ${json.encode_to_string(notes)}
callbackUrl: ${cb.url}
- wait_for_review:
try:
call: events.await_callback
args:
callback: ${cb}
timeout: 172800
result: review
except:
as: e
steps:
- escalate:
raise: ${"no review within 48 hours: " + json.encode_to_string(e)}
- decide:
switch:
- condition: ${review.http_request.body.approved == true}
return: "approved"
next: rejected
- rejected:
return: "rejected"
agent_retryable:
params: [e]
steps:
- check:
switch:
- condition: ${"code" in e and e.code in [409, 429, 502, 503, 504]}
return: true
- condition: ${"TimeoutError" in e.tags or "ConnectionError" in e.tags or "ConnectionFailedError" in e.tags}
return: true
- otherwise:
return: falseThe retry predicate is deliberate. The built-in http.default_retry retries 429, 502, 503 and 504 plus connection and timeout errors, five times, with a backoff starting at 1 second, multiplying by 1.25 and capped at 60 seconds. http.default_retry_non_idempotent retries only 429, 503 and connection failures. Neither retries 409, which this design uses for "still running", and neither retries 500, which is correct: a 500 from an agent is usually a bug or a refusal that will repeat. Because the step endpoint is replay-safe, retrying a timeout is safe here; against an endpoint without the session check, it would run the agent twice.
The draft endpoint runs the drafting agent, stores the draft, passes the callback URL to the review queue and only then returns, so no work happens after the response, when Cloud Run may throttle the CPU. Whoever approves, a review UI or a worker, POSTs a JSON body to that URL. The caller's identity needs the workflows.callbacks.send permission, which the Workflows Invoker role includes. Requests that arrive before the workflow reaches await_callback are handled as the documentation describes: the first is kept, later ones are rejected.
A traced execution
Here is how one execution behaves, as an illustrative trace rather than a captured log. The workflow starts with three topics. The research branch for "pricing" hits a cold Cloud Run instance and gets 503; the predicate accepts it and the branch retries 2 seconds later and succeeds. The "competitors" branch runs long: the model answers at 920 seconds, but the HTTP call timed out at 900, so Workflows raises TimeoutError and retries. Deliveries at about 902, 906 and 914 seconds find the session without a result and get 409, which the predicate retries. The delivery at about 930 seconds finds step_result in state and returns it with replayed: true, so the model is not called again and the bill is not doubled.
The draft step runs, the reviewer receives the callback URL and approves the next morning. During the night, the Cloud Run revision was redeployed twice; nothing was lost because the open wait lives in Workflows, not in a container. The execution returns "approved" after about fifteen hours, with five retries and one replay in its history.
Why not call the dev server's /run directly
ADK Java's dev module ships AdkWebServer, whose POST /run endpoint takes appName, userId, sessionId, newMessage and optional stateDelta. It is tempting to point Workflows at it directly, and for a prototype that works, but read the source first. /run blocks until the agent finishes and returns the entire event list, which on a tool-heavy run can approach the 2 MB response limit. It maps every agent failure to 500. Re-sending the same request appends another user turn to the session, so it is not idempotent. And POST /apps/{app}/users/{user}/sessions/{id} returns 400 "Session already exists" when the id is taken, so a retried create step keyed by execution id fails unless you treat that 400 as success or fall back to a GET. A thin endpoint of your own, as above, avoids all four problems. Deployment of that service is covered in the Cloud Run deployment article.
Failure modes
- Double execution. A timed-out call is retried while the first is still running. Without the session check and 409, the agent runs twice and any side-effecting tool fires twice.
- Variable bloat. Accumulating full agent outputs in workflow variables hits the 512 KB limit mid-run. Keep summaries in variables and the full text in GCS or session state.
- Lost state on scale-in. An in-memory session service passes every test and loses finished steps in production.
- Callback never arrives. Without an explicit timeout and a handler for TimeoutError, the execution waits out the default and fails without a reason. Catch it and route to escalation.
- Quota stampede. A parallel step over fifty topics with no concurrency limit fires fifty model calls at once and earns a wall of 429s; see the dead-letter article for parking items that keep failing.
Operational guidance
Give the workflow its own service account with the Cloud Run Invoker role on the agent service and nothing broader. Log the execution id in every agent request and attach it to traces, so one id joins the Workflows execution history to the agent's spans and model costs. Watch three numbers per step: retries, replays and 409s. Rising replays mean timeouts are too short; rising 409s mean steps overlap. The long-running task patterns article covers alternatives when a step must outlive even a callback-based design.
Trade-offs
| Option | Durability | Best for | Cost of choosing it |
|---|---|---|---|
| ADK workflow agents only | One request | Short in-process pipelines | Lost on restart; no long waits |
| Cloud Workflows + ADK on Cloud Run | Managed, up to a year | GCP-native multi-step jobs with waits | YAML, HTTP limits, idempotent endpoints |
| Cloud Tasks queue per step | Per task | Independent fan-out jobs | You build the control flow yourself |
| Code-first durable engine | Event-sourced replay | Complex logic, many versions | Run and operate the engine |
What to do next
- Write down which agent steps can exceed 30 minutes or wait on people; those need callbacks.
- Build the step endpoint with session ids derived from execution id plus step name, a durable session service, replay of finished steps and 409 for in-progress ones.
- Deploy it to Cloud Run with authentication required, and grant the workflow's service account the invoker role.
- Write the workflow with OIDC auth, an explicit timeout per call, a custom retry predicate and a concurrency limit on parallel steps.
- Kill an instance mid-step and confirm the retry replays instead of re-running.
- Read the Cloud Workflows overview for the rest of the syntax before adding more steps.