An ADK Java agent runs a turn by loading a session, calling a model one or more times, running tools, and appending events back to the session. Each step is I/O. That makes agent services easy to scale out in one sense, because a replica needs little CPU per turn, and hard in another, because every replica shares the same session store, the same model quota and the same downstream tools.
This article is about the second part. It builds a capacity model for an ADK Java service, shows where the real ceilings are, explains what happens when two turns for the same session land on different replicas, and gives code for admission control and per-session leases. For failover and regional design see the high-availability article; for container sizing and autoscaling signals see the Kubernetes one.
The precondition: nothing lives in a replica
Scaling out starts with a precondition: a replica must not hold anything between turns that another replica would need. In ADK Java a Runner is built around an agent and three services, a BaseSessionService, a BaseArtifactService and optionally a BaseMemoryService. Its entry point, runAsync(userId, sessionId, newMessage), returns a Flowable<Event> and loads everything the turn needs from those services.
The in-memory implementations break the precondition. InMemorySessionService keeps sessions in the JVM's heap, so a second replica cannot see them and a restart loses them. For more than one replica you need a shared implementation: the ADK Java javadoc lists VertexAiSessionService (the managed Vertex AI session API) and FirestoreSessionService, and you can implement BaseSessionService over your own database, as the Postgres schema article on this site does. Artifacts need the same treatment, for example a Cloud Storage-backed artifact service instead of the in-memory one.
Everything else in the process should be a cache that can be lost: compiled prompts, tool clients, connection pools. If a value cannot be rebuilt from the shared stores, it is state, and it has to move out.
A capacity model for agent turns
Capacity for I/O-bound services follows Little's law: the average number of requests in flight equals the arrival rate times the average time each spends in the system. For agents the time is long. A turn that makes three model calls of about 1.5 seconds and two tool calls of 300 milliseconds takes about 5 seconds, so a peak of 40 turns per second means about 200 turns in flight at once.
In-flight turns cost memory, not CPU. Each holds the loaded session (every event, unless you filter them), the prompt being assembled, the streamed response and any tool payloads. Measure this in a load test; a few megabytes per turn is common for long sessions. If a replica can safely hold 60 turns, 200 in flight needs four replicas, and N plus one or two for failure and deploys gives five or six.
That per-replica limit should be enforced, not assumed. Without an admission limit a traffic spike lets every replica accept more turns than its heap can hold, garbage collection stalls everything, latency rises, and Little's law turns rising latency into even more in-flight turns: a feedback loop that ends in out-of-memory restarts. A semaphore that rejects work over the limit, with a retryable error, keeps the replica in its linear range and lets the load balancer or client retry elsewhere.
The real ceiling is upstream
Adding replicas raises capacity only until a shared dependency saturates, and for agents the first one is almost always the model endpoint. Model quotas are set per project and per model, in requests per minute and tokens per minute, and every replica draws on the same pool.
Turn the capacity model into model load. At 40 turns per second with three calls each, the service makes 120 model requests per second, 7,200 per minute. If each prompt carries 8,000 input tokens of instructions, history and tool results, that is about 58 million input tokens per minute. Compare both numbers with the quota you actually have; if either exceeds it, more replicas only produce more quota errors.
Agent structure multiplies this. A ParallelAgent that fans out to four sub-agents, each making two calls, turns one user turn into eight model calls in a burst; a LoopAgent multiplies calls by its iteration count. Count calls per turn from traces, not from the design diagram, and treat the 95th percentile, because a few runaway loops can consume a large share of quota.
Tools are the second ceiling. An agent that calls an internal search service on most turns sends that service the full scaled-out load. Give each downstream dependency its own concurrency limit per replica, so one slow tool cannot absorb every in-flight slot.
Load on the session store
The session store sees two kinds of load. Writes: every turn appends several events, the user message, each model response, each function call and response, so 40 turns per second with eight events each is 320 appends per second, plus state updates. Reads: getSession loads the session at the start of each turn, and by default that means its event history. Read volume therefore grows with session length, and long-lived sessions quietly become the dominant cost.
Control both. GetSessionConfig lets a caller filter which events are loaded; check which options your ADK version supports. The Runner also accepts an events compaction configuration in recent releases, which summarises older history; verify its fields against your version's javadoc before relying on it. Set a retention policy for idle sessions, and partition or shard the store by session id so hot users do not concentrate on one partition. For a database-backed implementation, the session row is updated on every turn, so it behaves like a hot row; the Postgres schema article covers how to tune for that.
Two turns, one session
The subtle problem in scale-out is two turns for the same session running at once. It happens more often than expected: a user double-clicks send, a client retries a request that is still running, or a webhook and a user message arrive together. With one replica and an in-memory store the turns at least interleave in one process. With many replicas they can run on different machines, each loading the same session, each appending its own events.
The ADK Java javadoc for BaseSessionService.appendEvent documents no concurrency, staleness or locking guarantee, so do not assume one. Depending on the backing store, the outcome ranges from interleaved events, which confuse the next model call, to a lost state update, where one turn's state changes overwrite the other's. Either way the conversation is corrupted silently.
There are three defences, from weakest to strongest. Client serialisation: the UI disables send until the previous turn completes. Necessary, but not enough, because retries and server-side triggers bypass it. Routing affinity: hash the session id at the load balancer so a session's turns usually reach one replica, where a local lock serialises them. Good for cache locality, but not a correctness guarantee, because a scale event or deploy reshuffles the hash ring mid-conversation. A per-session lease: before running a turn, acquire a short lease on the session id in a shared store, and reject or queue a second turn while it is held. This is the one that holds across replicas.
Admission control and leases in code
The executor below combines admission control and a lease around runAsync. SessionLeases is your interface over a shared store: in Redis, SET key token NX PX ttl to acquire and a compare-and-delete script to release; in a database, a row with an owner and expiry.
public final class TurnExecutor {
private final Runner runner;
private final SessionLeases leases; // shared store: Redis, Firestore, a DB table
private final Semaphore slots; // per-replica admission, sized by load test
public TurnExecutor(Runner runner, SessionLeases leases, int maxInFlight) {
this.runner = runner;
this.leases = leases;
this.slots = new Semaphore(maxInFlight);
}
public Flowable<Event> turn(String userId, String sessionId, String text) {
Content message = Content.fromParts(Part.fromText(text));
// defer: each subscription (including a retry) takes exactly one permit
// and gives back exactly one.
return Flowable.defer(() -> {
if (!slots.tryAcquire()) {
// Map to HTTP 429 with Retry-After; the client or LB tries another replica.
return Flowable.<Event>error(new OverloadedException("replica at capacity"));
}
return Flowable.using(
// Throws SessionBusyException (map to 409) if another turn holds it.
() -> leases.acquire(sessionId, Duration.ofMinutes(2)),
lease -> runner.runAsync(userId, sessionId, message),
Lease::release)
.subscribeOn(Schedulers.io()) // lease I/O off the caller's thread
.doFinally(slots::release);
});
}
}Two details matter. The lease TTL must exceed the longest legitimate turn, or a slow turn loses its lease and a second turn starts beside it; renew it from a timer for long tool calls. And a lease can still expire under a paused JVM, so for strict safety have the session service reject an append whose turn id is not the current lease holder's, the fencing-token pattern. A SessionBusyException should be visible to clients as "a turn is already running", not as a generic failure.
Worked example
A support agent serves 40 turns per second at peak. Traces show 3.2 model calls and 1.4 tool calls per turn at the median, 9 calls at the 95th percentile, and a 5.5 second mean turn. Little's law gives 220 turns in flight. A load test shows heap per turn reaching 3 MB on long sessions, so a replica with a 1 GB budget for turns is capped at 60 in flight after headroom. Four replicas carry 240; six are deployed so the service survives a zone loss and a rolling deploy.
Model load is 128 requests per second, about 7,700 per minute, and at about 6,500 input tokens per call that is roughly 50 million input tokens per minute. The project's token quota is the binding limit, so the team requests a quota increase, moves static instructions into a context cache where the model supports it, and caps history loading. Session-store load is about 350 appends per second; the store is sharded by session id. Leases fix a recurring bug in which mobile clients retried slow turns and produced duplicated tool calls.
| Resource | Demand at peak | Limit | Binding? |
|---|---|---|---|
| Replica heap | 220 turns in flight | 60 per replica x 6 = 360 | no; 300 after a zone loss |
| Model requests | ~7,700 per minute | project quota | check |
| Model input tokens | ~50M per minute | project quota | yes, first |
| Session appends | ~350 per second | store throughput | no, sharded |
Failure modes and trade-offs
- Quota storms. Autoscaling on CPU or latency adds replicas while the model returns quota errors, which raises error rate without raising throughput. Scale on in-flight turns and alert on quota errors separately.
- Retry amplification. Client, gateway and model client each retry, multiplying load at the worst moment. Retry in one layer, with jittered backoff and a budget.
- Duplicate side effects. A retried turn reruns tools. Make tool calls idempotent with a key derived from the session and turn.
- Unbounded history. Sessions that never end make every turn slower and every read larger. Filter, compact and expire.
- Hot sessions. A shared support session or a bot account concentrates writes. Detect it in store metrics and rate-limit per session.
- Lost leases. TTLs shorter than real turns cause the very concurrency they prevent. Set TTL from the latency tail, renew, and fence.
The trade-off is between simplicity and guarantees. Affinity alone is simple and fast but fails during rebalancing; leases add a store round trip per turn and a new dependency, but make same-session concurrency impossible. For most agents the round trip is negligible beside a multi-second model call.
What to do next
- Replace every in-memory service with a shared one and prove it by killing a replica mid-conversation.
- Measure calls per turn, turn latency and heap per turn from traces and a load test, and compute in-flight demand with Little's law.
- Set a per-replica admission limit and return a retryable 429 above it.
- Compare model request and token demand with your quota and fix the binding limit before adding replicas.
- Add per-session leases with fencing, and route by session id for locality.
- Bound session history and give each downstream tool its own concurrency limit.
Keep learning: ADK Java high availability, Deploying ADK Java on Kubernetes, Sessions, state and context in ADK Java and A Postgres session table for ADK.