Many agent tasks are several independent questions wearing one coat. Review a pull request for security, performance and style. Look up a customer in billing, support and product usage. Draft an answer from three retrieval sources. Running those sub-tasks one after another makes the user wait for the sum of their latencies. Running them at the same time makes the user wait for the slowest one. That is the fan-out pattern, and in the Agent Development Kit for Java it is expressed with ParallelAgent.
The API is small: a builder, a list of sub-agents, and optionally a scheduler. The design around it is not. You have to decide what each branch may see, where it writes its result, how the next step gathers those results, how many model calls you can afford to run at once, and what the user sees while the branches run. This article covers those decisions with a worked example, a code-review pipeline, and points to the companion article on handling failures in parallel agents for what happens when a branch throws. API details are from the adk-java source and documentation at the time of writing; confirm them against the version you depend on.
What ParallelAgent does, precisely
ParallelAgent is a workflow agent: it makes no model calls of its own. When it runs, it takes its sub-agents, derives a child invocation context whose branch includes its own name, and starts each sub-agent's runAsync stream subscribed on a scheduler. The scheduler defaults to RxJava's Schedulers.io() and can be replaced through the builder's scheduler(...) method. The per-branch event streams are combined with Flowable.merge, and the merged stream ends early if an event from a direct sub-agent sets escalate. Live bidirectional streaming through runLive is not supported for this agent.
Three consequences follow. First, the events you receive are interleaved: a token from the style reviewer can arrive between two tokens of the security reviewer, and the documentation is explicit that ordering is not deterministic. Second, the parallel agent completes only when every branch completes, so its latency is that of the slowest branch. Third, because the underlying stream is a merge, one branch that errors takes the whole fan-out down unless you guard it, which is the subject of the companion article.
When fan-out pays, and when it does not
Fan-out is worth it when the sub-tasks are genuinely independent: no branch needs another's output, and the combined result is better than any one. It buys latency, not cost. Three model calls cost the same tokens whether they run in series or in parallel, and the synthesis step adds a fourth.
It does not pay when branches depend on each other (use a SequentialAgent and pass state forward, as in passing context between sequential agents), when one branch dominates the latency so the others save little, when your model quota cannot absorb N concurrent calls per user request, or when the answer only needs one perspective and the extra branches are there for show. A useful rule: if you cannot say what the synthesizer does with each branch's output, remove the branch.
It is also worth separating fan-out across agents from parallel tool calls within one agent. When a single LlmAgent decides it needs three lookups, the model can request several function calls in one turn. That keeps one reasoning thread and one prompt. A ParallelAgent gives each branch its own instruction, model, tools and output, which is what you want when the branches are different jobs, not different lookups.
Branch isolation and shared state
Each sub-agent runs in its own branch, and the branch name is the dotted path of agent names, for example review_pipeline.review_fanout.security_reviewer. Conversation history is filtered by branch, so an LLM agent in one branch does not see the model turns of its siblings. That isolation is a feature: three reviewers primed by each other's opinions converge and stop being independent.
Session state, however, is one map shared by the whole invocation. The documentation describes branches as having no automatic sharing of history or state during execution, which you should read as: do not design branches that read each other's state while running, because whether a sibling's write has landed yet is a race. What branches can safely do is read state written before the fan-out started, and write to keys that no other branch writes. Give every branch a distinct outputKey. Two branches writing the same key both succeed, and the value left behind is whichever finished last, which changes from run to run.
Worked example: a three-reviewer pipeline
The pipeline reviews a code diff that the caller has placed in state under diff. Three reviewers run in parallel, each with a narrow instruction and its own output key, and a synthesizer merges them into one prioritized review. Instructions reference state with the {key} placeholder syntax, which ADK fills from session state before the model call.
import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.ParallelAgent;
import com.google.adk.agents.SequentialAgent;
public final class ReviewPipeline {
static final String MODEL = "gemini-2.5-flash"; // any model your project can call
static LlmAgent reviewer(String name, String focus, String key) {
return LlmAgent.builder()
.name(name)
.model(MODEL)
.description("Reviews a diff for " + focus)
.instruction("""
You review code changes for %s only. Ignore every other concern.
Diff:
{diff}
Return at most five findings as lines of the form
SEVERITY | file:line | finding | suggested fix
Return NONE if you find nothing.""".formatted(focus))
.outputKey(key)
.build();
}
public static SequentialAgent build() {
ParallelAgent fanout = ParallelAgent.builder()
.name("review_fanout")
.description("Independent security, performance and style reviews")
.subAgents(
reviewer("security_reviewer", "security", "review_security"),
reviewer("performance_reviewer", "performance", "review_perf"),
reviewer("style_reviewer", "readability and style", "review_style"))
.build();
LlmAgent synthesizer = LlmAgent.builder()
.name("synthesizer")
.model(MODEL)
.instruction("""
Merge three independent reviews into one list ordered by severity.
Remove duplicates, keep file:line references, do not add new findings.
Security: {review_security}
Performance: {review_perf}
Style: {review_style}""")
.outputKey("review_final")
.build();
return SequentialAgent.builder()
.name("review_pipeline")
.subAgents(fanout, synthesizer)
.build();
}
}Running it uses the ordinary runner. The event stream carries events from all three branches interleaved, then the synthesizer's events. Group or filter by event.author() or by branch rather than assuming an order.
InMemoryRunner runner = new InMemoryRunner(ReviewPipeline.build(), "code-review");
Session session = runner.sessionService()
.createSession("code-review", userId, new ConcurrentHashMap<>(Map.of("diff", diffText)), null)
.blockingGet();
Content msg = Content.fromParts(Part.fromText("Review this change."));
runner.runAsync(userId, session.id(), msg)
.filter(Event::finalResponse)
.blockingForEach(e -> log.info("{} finished", e.author()));The four-argument createSession that seeds initial state exists in current adk-java releases; if your version lacks it, create the session and have a first agent or a callback copy the diff into state. Keep the diff out of the user message if possible: every branch receives the same user content, and a large diff in the message is paid for in every branch's prompt anyway, but in state you can choose which agents template it in.
The instructions do two jobs beyond describing the task. They narrow each reviewer to one concern, which is what makes the outputs complementary rather than three copies of the same review. And they fix an output format, which lets the synthesizer, or plain Java code, parse and de-duplicate findings. The synthesizer is told not to invent findings, because a merge step that adds content is a fourth reviewer nobody asked for.
Dynamic fan-out: when N is not known at build time
A ParallelAgent's sub-agents are fixed when it is built. Many real fan-outs are data-driven: one branch per changed file, per retrieved document or per ticket. Agents are ordinary immutable Java objects and cheap to construct, so the simplest correct approach is to build the pipeline per request from the data, with branch names and output keys derived from a stable index.
static SequentialAgent perFileReview(List<String> files) {
List<BaseAgent> branches = new ArrayList<>();
for (int i = 0; i < files.size(); i++) {
branches.add(LlmAgent.builder()
.name("file_reviewer_" + i) // names must be valid identifiers
.model(MODEL)
.instruction("Review only this file:\n{file_" + i + "}")
.outputKey("review_file_" + i)
.build());
}
ParallelAgent fanout = ParallelAgent.builder()
.name("per_file_fanout").subAgents(branches).build();
return SequentialAgent.builder()
.name("per_file_pipeline").subAgents(fanout, buildSynthesizer(files.size())).build();
}Build fresh agent instances for every pipeline rather than reusing one instance under two parents; the parent link is set when an agent is attached, and sharing instances across trees is a source of confusing branch names. Cap N before you build: a pull request touching 400 files should not create 400 concurrent model calls. Batch files into a fixed number of branches, or review the top files by change size and summarize the rest.
Concurrency, quotas and the scheduler
Every branch is a model call, and model endpoints enforce requests-per-minute and tokens-per-minute quotas per project. A fan-out of three on every user request triples your request rate and multiplies the chance that a burst of traffic produces 429 responses in the middle of a fan-out. Plan capacity in branch-calls, not user requests, and put a shared limiter in front of the model client so that concurrency is bounded across all requests, not only within one; rate limiting in ADK Java covers the patterns.
The default Schedulers.io() is an unbounded, caching thread pool shared with anything else in the JVM that uses it. Passing a dedicated scheduler through ParallelAgent.builder().scheduler(...) isolates fan-out work from other I/O and makes its thread usage visible, a bulkhead in the classic sense. A scheduler backed by a virtual-thread executor is a natural fit for branches that spend their time waiting on network calls; virtual threads in ADK Java discusses the trade-offs. A bounded pool also limits how many branches start at once, but treat it as isolation rather than as your rate limiter: it bounds threads in this process, not calls to the provider.
Set deadlines. Without one, the fan-out waits for the slowest branch indefinitely, and a single hung connection holds the user's request open. A per-branch timeout, turned into a recorded 'timed out' result rather than an error, lets the synthesizer proceed with partial input; the guard-wrapper pattern for that lives in the failures article.
Streaming and user experience
Because events interleave, streaming partial text from all branches straight to a chat window produces an unreadable mix. Two patterns work. Show progress, not content: display one status line per branch that flips from running to done as each branch's final response arrives, then stream only the synthesizer's answer. Or render each branch into its own panel keyed by author. Either way, the consumer must route by author or branch, never by arrival order.
Time to first useful token is dominated by the slowest branch plus the synthesizer's first token. If that is too slow, consider whether the synthesizer is necessary at all: for some fan-outs, a deterministic Java merge that sorts findings by severity is faster, cheaper and more predictable than another model call.
Testing and observability
Test the pipeline shape without a model. Stub each branch with a small custom agent that emits a fixed final response and writes its output key, then assert on the final state: all keys present, synthesizer ran after the fan-out, and nothing depended on arrival order. Run the test many times with randomized delays in the stubs; ordering bugs are intermittent by nature and only show up when timing changes.
In production, trace one span per branch under the parallel agent's span and record the branch name, model, latency, token counts and outcome. The metrics that matter are the latency spread between the fastest and slowest branch (a large spread means one branch is wasting the others' parallelism), the per-branch error and timeout rate, and cost per request. Observability for ADK Java shows how to wire callbacks into tracing.
Failure modes and trade-offs
| Failure | Cause | Mitigation |
|---|---|---|
| Whole request fails on one 429 | merge propagates the first error | Guard each branch; see the failures article |
| Synthesizer sees a missing or stale key | Branch failed or wrote a different key | Distinct keys, status keys, quorum check before synthesis |
| Results differ between identical runs | Two branches share an outputKey | One key per branch |
| Requests hang | No deadline on a slow branch | Per-branch timeout recorded as a result |
| Quota exhaustion under load | Capacity planned per request, not per branch call | Shared limiter; cap N |
| Branches agree on everything | Instructions not differentiated | Narrow each branch to one concern |
The core trade-off is latency against cost and complexity: fan-out makes the user wait less and pays the same or more, and it introduces concurrency into a system that was otherwise a straight line. For an overview of how parallel, sequential and loop agents compose, see ADK workflow agents.
What to do next
- List the sub-tasks in your agent and mark which are independent; only those are fan-out candidates.
- Wrap the fan-out in a SequentialAgent with an explicit gather step and give every branch a distinct outputKey.
- Narrow each branch's instruction to one concern and fix its output format so the merge can parse it.
- Cap the number of branches and add a shared limiter sized in branch-calls per minute.
- Pass a dedicated scheduler and set a per-branch deadline that produces a recorded result.
- Guard branches against errors using the companion failures article before going to production.
- Add per-branch spans and alert on latency spread and per-branch error rate.