"Workflow" means two different things to an ADK Java developer. Inside one process, ADK has workflow agents, SequentialAgent, ParallelAgent and LoopAgent, which run sub-agents in a fixed pattern within a single invocation. Outside the process, Google Cloud Workflows is a managed orchestrator that runs a YAML-defined sequence of HTTP calls, with retries, parallel branches and waits, for up to a year per execution. This page is about the second: using Cloud Workflows to drive ADK Java agents deployed on Cloud Run, so a multi-step agent job survives restarts, waits days for a human, and never pays twice for a step it has already finished.

You will see the architecture, the limits that shape it, a step endpoint in Java that makes agent calls safe to retry, a workflow definition with parallel research, a custom retry predicate and a human-approval callback, a traced execution, and the failure modes. The generic theory of durable workflows is in the workflow orchestration architecture article; here every detail is specific to Cloud Workflows and ADK Java.

Two kinds of workflow, and when to use each

Use ADK's workflow agents when the whole job finishes in one request: a pipeline of three agents that takes forty seconds, where losing it to an instance restart just means the caller retries. Use Cloud Workflows when the job is longer than a request can live, waits on people or external systems, mixes agent steps with ordinary API calls, or must not repeat an expensive step after a crash. The two compose: a single Cloud Workflows step can call an endpoint that runs a SequentialAgent internally.

The division of labour is strict. Cloud Workflows owns control flow, retries, timers, waits and the record of which steps finished. The ADK Java service owns reasoning: it runs one agent invocation per step and keeps the conversation in a durable session. Neither should try to do the other's job; an agent that loops forever waiting for approval inside a Cloud Run request is the classic mistake.

Architecture and the limits that shape it

Cloud Workflows owns the run; ADK Java on Cloud Run owns each stepCloud Workflowsexecution (up to 1 year)init: exec idresearch (parallel)draftawait callbackpublishCloud Run serviceADK Java step endpointsPOST /steps/researchPOST /steps/draftGeminimodel callsSession storedurable, keyed by exec idReviewer or async workerPOSTs to callback URLOIDC POST, retrycallback URLcallbackEach step is an idempotent HTTP call keyed by execution id plus step name; long waits use callbacks, not open HTTP requests.Limits that shape the design: HTTP timeout 1800 s max, 2 MB response, 512 KB of variables per execution.
Cloud Workflows holds the control flow and the waits; each step is an authenticated, retry-safe call to an ADK Java endpoint on Cloud Run backed by a durable session store.

A handful of documented limits drive the design. Check them against the current Workflows quotas page before you build, since quotas change:

LimitValueDesign consequence
Execution duration1 yearMulti-day approvals are fine inside one execution
HTTP request timeout300 s default, 1800 s maxAgent steps longer than 30 minutes need callbacks
HTTP response size2 MBReturn the final answer, not the event history
Variables, arguments and events512 KB per executionStore large outputs in GCS or session state; pass references
Steps per execution100,000Loops over thousands of items need batching
Parallel branches10 per step; 20 concurrent branches and iterationsUse concurrency_limit to protect model quotas
Callback wait43200 s default timeoutSet an explicit timeout and handle TimeoutError

An idempotent step endpoint in Java

Cloud Workflows retries a step by sending the same HTTP request again, and it cannot know whether the first attempt reached the agent. So the step endpoint has to make repeats harmless. The design below derives the ADK session id from the workflow execution id plus the step name, and uses the agent's outputKey so the final answer lands in session state. A repeat of a finished step finds the stored answer and returns it without calling the model. A repeat that arrives while the first delivery is still running gets 409, which the workflow retries after a backoff.

@RestController
public class StepController {
  private static final String APP = "research_app";
  private static final String USER = "workflow";
  private final Runner runner;

  public StepController(BaseSessionService durableSessions) {  // not InMemorySessionService
    LlmAgent researcher = LlmAgent.builder()
        .name("researcher")
        .model("gemini-2.5-flash")
        .instruction("Research the topic in the user message. Answer in under 200 words.")
        .outputKey("step_result")             // final text is saved to session state
        .build();
    this.runner = Runner.builder()
        .agent(researcher)
        .appName(APP)
        .sessionService(durableSessions)
        .artifactService(new InMemoryArtifactService())
        .build();
  }

  public record StepRequest(String executionId, String step, String input) {}
  public record StepResult(String text, boolean replayed) {}

  @PostMapping("/steps/research")
  public StepResult research(@RequestBody StepRequest req) {
    String sessionId = req.executionId() + "-" + req.step();
    BaseSessionService sessions = runner.sessionService();
    Session s = sessions.getSession(APP, USER, sessionId, Optional.empty()).blockingGet();
    if (s != null && s.state().containsKey("step_result")) {
      return new StepResult(String.valueOf(s.state().get("step_result")), true);  // replay
    }
    if (s != null) {                          // a delivery of this step is still running
      throw new ResponseStatusException(HttpStatus.CONFLICT, "step in progress");
    }
    sessions.createSession(APP, USER, new ConcurrentHashMap<>(), sessionId).blockingGet();
    Content msg = Content.fromParts(Part.fromText(req.input()));
    StringBuilder text = new StringBuilder();
    runner.runAsync(USER, sessionId, msg)
        .filter(Event::finalResponse)
        .blockingForEach(e -> text.append(e.stringifyContent()));
    return new StepResult(text.toString(), false);
  }
}

The class names come from the google/adk-java core module: Runner.builder() requires an artifact service, getSession returns an RxJava Maybe (so blockingGet() is null when the session does not exist), and Event.finalResponse() marks the answer. Inject a durable BaseSessionService: Cloud Run scales instances away, and an in-memory store forgets every finished step with them. Two refinements belong in production. Record a start time in the initial state so a session left behind by a crashed instance can be detected as stale and restarted, and map transient model errors to 503 so the retry policy can tell them from bugs.

The workflow definition

The workflow below researches several topics in parallel, at most three at a time to protect the model quota, then asks the agent service to draft a report and hand it to a reviewer, and waits up to two days for the reviewer's decision. The auth: type: OIDC block makes Workflows attach an identity token for its service account, so the Cloud Run service can require authentication. The audience defaults to the full request URL, path included; setting it to the service base URL is the safe choice for Cloud Run. Raise the Cloud Run service's request timeout to at least the workflow's timeout: value: Cloud Run defaults to 300 seconds (maximum 3,600), and a service that cuts the agent off at 300 seconds leaves a session with no result, so every retry would get 409.

main:
  params: [args]
  steps:
    - init:
        assign:
          - exec_id: ${sys.get_env("GOOGLE_CLOUD_WORKFLOW_EXECUTION_ID")}
          - base: ${args.agentUrl}
          - notes: {}
    - research:
        parallel:
          shared: [notes]
          concurrency_limit: 3
          for:
            value: topic
            in: ${args.topics}
            steps:
              - call_agent:
                  try:
                    call: http.post
                    args:
                      url: ${base + "/steps/research"}
                      auth:
                        type: OIDC
                        audience: ${base}
                      timeout: 900
                      body:
                        executionId: ${exec_id}
                        step: ${"research-" + topic}
                        input: ${topic}
                    result: resp
                  retry:
                    predicate: ${agent_retryable}
                    max_retries: 6
                    backoff:
                      initial_delay: 2
                      max_delay: 60
                      multiplier: 2
              - keep:
                  assign:
                    - notes[topic]: ${resp.body.text}
    - make_callback:
        call: events.create_callback_endpoint
        args:
          http_callback_method: "POST"
        result: cb
    - request_review:
        call: http.post
        args:
          url: ${base + "/steps/draft"}
          auth:
            type: OIDC
            audience: ${base}
          body:
            executionId: ${exec_id}
            step: "draft"
            input: ${json.encode_to_string(notes)}
            callbackUrl: ${cb.url}
    - wait_for_review:
        try:
          call: events.await_callback
          args:
            callback: ${cb}
            timeout: 172800
          result: review
        except:
          as: e
          steps:
            - escalate:
                raise: ${"no review within 48 hours: " + json.encode_to_string(e)}
    - decide:
        switch:
          - condition: ${review.http_request.body.approved == true}
            return: "approved"
        next: rejected
    - rejected:
        return: "rejected"

agent_retryable:
  params: [e]
  steps:
    - check:
        switch:
          - condition: ${"code" in e and e.code in [409, 429, 502, 503, 504]}
            return: true
          - condition: ${"TimeoutError" in e.tags or "ConnectionError" in e.tags or "ConnectionFailedError" in e.tags}
            return: true
    - otherwise:
        return: false

The retry predicate is deliberate. The built-in http.default_retry retries 429, 502, 503 and 504 plus connection and timeout errors, five times, with a backoff starting at 1 second, multiplying by 1.25 and capped at 60 seconds. http.default_retry_non_idempotent retries only 429, 503 and connection failures. Neither retries 409, which this design uses for "still running", and neither retries 500, which is correct: a 500 from an agent is usually a bug or a refusal that will repeat. Because the step endpoint is replay-safe, retrying a timeout is safe here; against an endpoint without the session check, it would run the agent twice.

The draft endpoint runs the drafting agent, stores the draft, passes the callback URL to the review queue and only then returns, so no work happens after the response, when Cloud Run may throttle the CPU. Whoever approves, a review UI or a worker, POSTs a JSON body to that URL. The caller's identity needs the workflows.callbacks.send permission, which the Workflows Invoker role includes. Requests that arrive before the workflow reaches await_callback are handled as the documentation describes: the first is kept, later ones are rejected.

A traced execution

Here is how one execution behaves, as an illustrative trace rather than a captured log. The workflow starts with three topics. The research branch for "pricing" hits a cold Cloud Run instance and gets 503; the predicate accepts it and the branch retries 2 seconds later and succeeds. The "competitors" branch runs long: the model answers at 920 seconds, but the HTTP call timed out at 900, so Workflows raises TimeoutError and retries. Deliveries at about 902, 906 and 914 seconds find the session without a result and get 409, which the predicate retries. The delivery at about 930 seconds finds step_result in state and returns it with replayed: true, so the model is not called again and the bill is not doubled.

The draft step runs, the reviewer receives the callback URL and approves the next morning. During the night, the Cloud Run revision was redeployed twice; nothing was lost because the open wait lives in Workflows, not in a container. The execution returns "approved" after about fifteen hours, with five retries and one replay in its history.

Why not call the dev server&#x27;s /run directly

ADK Java's dev module ships AdkWebServer, whose POST /run endpoint takes appName, userId, sessionId, newMessage and optional stateDelta. It is tempting to point Workflows at it directly, and for a prototype that works, but read the source first. /run blocks until the agent finishes and returns the entire event list, which on a tool-heavy run can approach the 2 MB response limit. It maps every agent failure to 500. Re-sending the same request appends another user turn to the session, so it is not idempotent. And POST /apps/{app}/users/{user}/sessions/{id} returns 400 "Session already exists" when the id is taken, so a retried create step keyed by execution id fails unless you treat that 400 as success or fall back to a GET. A thin endpoint of your own, as above, avoids all four problems. Deployment of that service is covered in the Cloud Run deployment article.

Failure modes

  • Double execution. A timed-out call is retried while the first is still running. Without the session check and 409, the agent runs twice and any side-effecting tool fires twice.
  • Variable bloat. Accumulating full agent outputs in workflow variables hits the 512 KB limit mid-run. Keep summaries in variables and the full text in GCS or session state.
  • Lost state on scale-in. An in-memory session service passes every test and loses finished steps in production.
  • Callback never arrives. Without an explicit timeout and a handler for TimeoutError, the execution waits out the default and fails without a reason. Catch it and route to escalation.
  • Quota stampede. A parallel step over fifty topics with no concurrency limit fires fifty model calls at once and earns a wall of 429s; see the dead-letter article for parking items that keep failing.

Operational guidance

Give the workflow its own service account with the Cloud Run Invoker role on the agent service and nothing broader. Log the execution id in every agent request and attach it to traces, so one id joins the Workflows execution history to the agent's spans and model costs. Watch three numbers per step: retries, replays and 409s. Rising replays mean timeouts are too short; rising 409s mean steps overlap. The long-running task patterns article covers alternatives when a step must outlive even a callback-based design.

Trade-offs

OptionDurabilityBest forCost of choosing it
ADK workflow agents onlyOne requestShort in-process pipelinesLost on restart; no long waits
Cloud Workflows + ADK on Cloud RunManaged, up to a yearGCP-native multi-step jobs with waitsYAML, HTTP limits, idempotent endpoints
Cloud Tasks queue per stepPer taskIndependent fan-out jobsYou build the control flow yourself
Code-first durable engineEvent-sourced replayComplex logic, many versionsRun and operate the engine

What to do next

  1. Write down which agent steps can exceed 30 minutes or wait on people; those need callbacks.
  2. Build the step endpoint with session ids derived from execution id plus step name, a durable session service, replay of finished steps and 409 for in-progress ones.
  3. Deploy it to Cloud Run with authentication required, and grant the workflow's service account the invoker role.
  4. Write the workflow with OIDC auth, an explicit timeout per call, a custom retry predicate and a concurrency limit on parallel steps.
  5. Kill an instance mid-step and confirm the retry replays instead of re-running.
  6. Read the Cloud Workflows overview for the rest of the syntax before adding more steps.
Key takeaway: Let Cloud Workflows own control flow, retries and waits, and let ADK Java own reasoning inside short, authenticated, idempotent HTTP steps. Key each step's session by execution id and step name, replay finished steps from session state, return 409 while a step is running, and retry only the codes that mean try again later. Use callbacks for anything that outlives a 30-minute request, keep large outputs out of workflow variables, and keep ADK workflow agents for pipelines that finish within one request.