A sequential chain is the simplest multi-agent shape in ADK Java: a SequentialAgent runs a fixed list of sub-agents in order, each one building on what the previous one left in session state. It is also the shape most teams reach for first and then struggle with, because the hard problems are not in the framework. They are in deciding where one stage ends and the next begins, what each stage promises the next, how a chain stops when a stage produces garbage, and how to keep latency and token cost from growing with every stage you add.
This page is about those design decisions. The mechanics of moving data between steps, meaning outputKey, {key} placeholders and includeContents, are covered in passing context between sequential steps and are assumed here. We start from what the framework actually guarantees, read from its source, then build a four-stage invoice-approval chain with a deterministic validation gate, a custom sequence that can halt, a parallel enrichment stage, and a test strategy for each stage.
When a chain is the right shape
ADK gives you two ways to compose agents. Workflow agents (SequentialAgent, ParallelAgent, LoopAgent) decide the order in code. An LlmAgent with sub-agents lets the model decide, by transferring control to whichever sub-agent it picks. A chain is right when the order of work is known at design time and does not depend on the content: extract, then validate, then decide. It is wrong when the next step genuinely depends on what the input turns out to be; in that case, a router or model-driven delegation is honest about the branching instead of hiding it in prompts.
| Shape | Who decides the order | Good for | Cost of getting it wrong |
|---|---|---|---|
| SequentialAgent | Your code, fixed | Pipelines: extract, check, decide, format | Stages run that should have been skipped |
| ParallelAgent | Your code, all at once | Independent lookups on the same input | Race on shared state keys |
| LoopAgent | Your code, until escalate or maxIterations | Draft, critique, revise | Runaway iterations |
| LlmAgent with sub-agents | The model, per turn | Open-ended requests that need routing | Unpredictable paths, hard to test |
Chains also compose: a stage can itself be a parallel or loop agent, each with one testable job.
What SequentialAgent guarantees, from the source
The whole of the non-resumable path in SequentialAgent.runAsyncImpl in the google/adk-java repository is Flowable.fromIterable(subAgents).concatMap(subAgent -> subAgent.runAsync(invocationContext)). Three properties follow. Stages run strictly one after another, because concatMap subscribes to the next stage only when the previous stream completes. Every stage receives the same InvocationContext, so they share one session and one invocation id. And the sequential agent itself never calls a model; it adds no tokens.
When the invocation is resumable, the same method fast-forwards to the stage being resumed, so completed stages are not re-run, and it stops starting new stages once an event reports a pending long-running tool call. That is the only built-in way the loop stops early. At the time of writing the file contains no reference to escalate: a stage that sets escalate on its event actions ends a LoopAgent, which checks event.actions().escalate(), but a SequentialAgent simply starts the next stage. Verify this against the version you depend on, but design as if a chain always runs to the end unless you build a stop into it. The same is true of errors: an exception in any stage terminates the Flowable and the stages after it never run, with no retry and no compensation.
Designing the stages
Treat each stage as a function with a declared input and output, even though ADK does not enforce one. Write the contract down before writing prompts:
| Stage | Reads | Writes | Kind | Model |
|---|---|---|---|---|
| extract | user message | invoice (JSON) | LlmAgent | small, fast |
| validate | invoice | validation: ok or a reason | custom BaseAgent | none |
| enrich | invoice | vendor_risk, po_match | ParallelAgent of two LlmAgents with tools | small |
| decide | all of the above | decision | LlmAgent | stronger |
Four rules come out of that table. First, one stage, one output key; two stages writing the same key is a bug that surfaces only when the order changes. Second, anything that can be checked without a model should be a stage without a model: JSON parsing, required fields, arithmetic, allowlists. Third, choose the model per stage; extraction and lookups rarely need the model that makes the final judgement, and the chain's cost is dominated by whichever stages use the big model. Fourth, keep the chain short. Every stage adds latency and another place for format drift, so merge two stages whenever neither ever needs to be tested or replaced alone.
The LLM stages
import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.ParallelAgent;
static final String FAST = "gemini-2.5-flash"; // use models your project can call
static final String STRONG = "gemini-2.5-pro";
static final LlmAgent EXTRACT = LlmAgent.builder()
.name("extract")
.model(FAST)
.instruction("Extract the invoice from the user's message. Reply with JSON only: "
+ "{\"vendor\": string, \"po_number\": string, \"currency\": string, "
+ "\"lines\": [{\"sku\": string, \"qty\": number, \"unit_price\": number}], "
+ "\"total\": number}")
.outputKey("invoice")
.build();
static final LlmAgent VENDOR_RISK = LlmAgent.builder()
.name("vendor_risk").model(FAST)
.instruction("Invoice: {invoice}. Use the vendor lookup tool and reply LOW, MEDIUM or HIGH with one reason.")
.outputKey("vendor_risk")
.build(); // tools omitted for brevity
static final LlmAgent PO_MATCH = LlmAgent.builder()
.name("po_match").model(FAST)
.instruction("Invoice: {invoice}. Use the purchase-order tool and report MATCH or MISMATCH with the differing lines.")
.outputKey("po_match")
.build();
static final ParallelAgent ENRICH = ParallelAgent.builder()
.name("enrich")
.subAgents(VENDOR_RISK, PO_MATCH)
.build();
static final LlmAgent DECIDE = LlmAgent.builder()
.name("decide").model(STRONG)
.instruction("Invoice: {invoice}\nVendor risk: {vendor_risk}\nPO check: {po_match}\n"
+ "Reply APPROVE, HOLD or REJECT, then one paragraph of reasons.")
.outputKey("decision")
.build();The two enrichment agents write different keys, so running them in parallel cannot race on state. See the parallel fan-out pattern for what happens when branches share keys or one branch fails.
A deterministic gate stage
The validation stage is a plain Java class. It reads the extracted JSON from session state, checks it, and emits one event that writes validation, either ok or the reason it failed, through the event's state delta, which is how every state change in ADK is recorded. It never calls a model, so it is fast, free and unit-testable.
public final class ValidateStage extends BaseAgent {
private static final ObjectMapper JSON =
new ObjectMapper().enable(DeserializationFeature.USE_BIG_DECIMAL_FOR_FLOATS);
public ValidateStage() {
super("validate", "Checks the extracted invoice before any further model calls.",
List.of(), List.of(), List.of());
}
/** Pure function: easy to unit test without ADK. Empty means valid. */
static Optional<String> problem(Object raw) {
try {
JsonNode inv = JSON.readTree(String.valueOf(raw));
if (!inv.hasNonNull("vendor") || !inv.hasNonNull("po_number") || !inv.hasNonNull("total")) return Optional.of("missing vendor, po_number or total");
if (!inv.path("lines").isArray() || inv.path("lines").isEmpty()) return Optional.of("no line items");
BigDecimal sum = BigDecimal.ZERO;
for (JsonNode l : inv.path("lines")) {
sum = sum.add(l.path("unit_price").decimalValue().multiply(l.path("qty").decimalValue()));
}
if (sum.compareTo(inv.path("total").decimalValue()) != 0) return Optional.of("line items do not add up to total");
return Optional.empty();
} catch (Exception e) {
return Optional.of("invoice is not valid JSON");
}
}
@Override
protected Flowable<Event> runAsyncImpl(InvocationContext ctx) {
return Flowable.defer(() -> {
Map<String, Object> delta = new ConcurrentHashMap<>();
delta.put(GatedSequence.GATE_KEY, problem(ctx.session().state().get("invoice")).orElse("ok"));
return Flowable.just(Event.builder()
.id(Event.generateEventId())
.invocationId(ctx.invocationId())
.author(name())
.branch(ctx.branch().orElse(null))
.actions(EventActions.builder().stateDelta(new ConcurrentHashMap<>(delta)).build())
.timestamp(System.currentTimeMillis())
.build());
});
}
@Override
protected Flowable<Event> runLiveImpl(InvocationContext ctx) {
return Flowable.error(new UnsupportedOperationException("live mode not supported"));
}
}Parsing numbers as BigDecimal and comparing with compareTo rather than using doubles is deliberate: a gate that rejects valid invoices because 0.1 plus 0.2 is not 0.3 teaches people to ignore the gate.
Halting the chain: a GatedSequence
The gate reports a failure, but a stock SequentialAgent would still run enrichment and the decision stage, spending tokens on an invoice already known to be broken, and possibly approving it. The fix mirrors how the framework's own resumable path stops: watch each event as it passes, and stop starting new stages once a stop signal appears.
public final class GatedSequence extends BaseAgent {
static final String GATE_KEY = "validation";
public GatedSequence(String name, List<? extends BaseAgent> stages) {
super(name, "Runs stages in order; stops after a stage writes " + GATE_KEY + " other than ok.",
stages, List.of(), List.of());
}
@Override
protected Flowable<Event> runAsyncImpl(InvocationContext ctx) {
return Flowable.defer(() -> {
AtomicBoolean halted = new AtomicBoolean(false); // per run, not per agent
return Flowable.fromIterable(subAgents()).concatMap(stage ->
halted.get()
? Flowable.<Event>empty()
: stage.runAsync(ctx).doOnNext(e -> {
Object v = e.actions().stateDelta().get(GATE_KEY);
if (v != null && !"ok".equals(v)) halted.set(true);
}));
});
}
@Override
protected Flowable<Event> runLiveImpl(InvocationContext ctx) {
return Flowable.error(new UnsupportedOperationException("live mode not supported"));
}
}
static final GatedSequence ROOT =
new GatedSequence("invoice_chain", List.of(EXTRACT, new ValidateStage(), ENRICH, DECIDE));Two details are easy to get wrong. The flag is created inside Flowable.defer, so each run gets a fresh one; a field on the agent would leak a halt from one user's invocation into the next. And concatMap evaluates the lambda only when the previous stage has completed, so the check sees the gate's event in time. This version drops the resumable fast-forward that the stock agent has; if you use resumable invocations, keep the stock agent and put a before-agent callback on each later stage that returns content when validation is not ok. Writing the key on every run matters: it has no temp: prefix, so it stays in session state after the turn, and a key written only on failure would still be there, stale, when the next clean invoice arrives in the same session. The patterns for writing custom agents like these are covered in extending the runtime with custom executors.
Worked example: two invoices through the chain
Invoice A: three lines, 12 units at 4.50, 2 at 30.00 and 1 at 9.00, total 123.00. Extract writes the JSON. Validate computes 54.00 plus 60.00 plus 9.00, which is 123.00, emits an event with an empty state delta, and the chain continues. Decide reads all three keys and writes APPROVE. Four stages ran, four LLM agents did work (two of them inside enrich), and only one of them used the strong model.
Invoice B: the email says total 132.00 for the same lines, a transposition typo or an attempted overbilling. Validate computes 123.00, writes validation = line items do not add up to total, and GatedSequence never starts enrichment or the decision stage. One cheap model call, zero strong-model calls, and a precise machine-readable reason your application can route to a human queue. Without the gate the decision model would have seen a plausible-looking invoice and might have approved it.
Latency, tokens and model choice
A chain's latency is the sum of its stages, and a parallel stage contributes its slowest branch. Its token cost grows faster than the number of stages, because with the default history setting each LLM stage can see the replies of the stages before it. Measure a single run per stage before tuning: if extraction takes 1.2 seconds, each lookup 0.8 and 1.5, and the decision 3.0, the chain costs about 5.7 seconds end to end, and putting the lookups in parallel saved 0.8.
Testing a chain
- Deterministic stages as pure functions.
ValidateStage.problemtakes an object and returns an optional reason. Test it with ordinary JUnit cases: valid, missing fields, wrong total, non-JSON, empty lines. - LLM stages against fixtures. Keep 20 to 50 real inputs per stage with expected outputs, run each stage alone through an in-memory runner, and assert structure (parses, required fields present), not exact wording.
- The whole chain on golden paths. A handful of end-to-end cases, including at least one that must halt, asserting which stages produced events by their
author.
Failure modes
- Assuming escalate stops the chain. It stops a loop, not a sequence. Build a gate.
- Silent format drift. An output key stores whatever text the model produced, and even with an output schema a failed validation still stores the raw string, so garbage flows downstream. Put a deterministic check after every LLM stage whose output is parsed.
- Shared keys in parallel stages. Two branches write the same key and the last writer wins. One key per branch.
- Side effects mid-chain. A stage that sends email or books a payment before a later stage rejects the request leaves you needing compensation. Put irreversible actions last, behind the final decision.
- Stale outcome keys. Session state outlives the turn; overwrite outcome keys every run or read the outcome from this run's events.
- Per-agent state in custom agents. Mutable fields on an agent instance are shared by every concurrent invocation.
Operating a chain
Every event carries its stage's name as author, so per-stage latency, error rate and halt rate fall out of your event log directly. Chart halts by reason: a rising rate of one reason after a prompt or model change is the earliest signal that a stage's contract has drifted. See observability for ADK Java for exporting those spans and metrics, and workflow orchestration for when a chain needs to outlive a single request.
What to do next
- Write the stage contract table for your chain: name, keys read, key written, model, and whether it can be deterministic.
- Replace every check a model is doing that code could do with a
BaseAgentstage, and unit-test its pure function. - Decide whether any stage must be able to stop the chain; if so, adopt a gate like
GatedSequenceor per-stage before-agent callbacks, and add a test that proves the later stages never run. - Move independent lookups into a
ParallelAgentstage with one output key per branch. - Measure latency and tokens per stage, then pick the smallest model that passes each stage's fixtures.
- Move irreversible side effects to the end of the chain and alert on halt rate by reason.