A multi-tenant agent platform usually starts with a handful of customers and one deployment. Adding the hundredth tenant feels the same as adding the tenth. Somewhere between a thousand and ten thousand tenants, three things change at once: the per-tenant objects you cache start to dominate memory, one shared model quota has to be divided among tenants whose demand differs by four orders of magnitude, and a single deployment becomes a single blast radius for everyone.
This article is about scaling the number and size of tenants in an ADK Java platform. Keeping tenant data apart is covered in Agent Tenant Data Separation, and per-tenant admission limits are a separate topic; here we assume both exist and ask how the system grows. We will bound the runner registry, introduce cells, do the capacity arithmetic that shows which resource actually runs out, split a provider quota fairly, place tenants with a simple bin-packing planner, and move a hot tenant without breaking its conversations.
Two axes of tenant scale
Tenant scale has two independent axes. The first is tenant count: how many distinct customers the platform holds, most of them idle at any moment. The second is tenant size: how much traffic the largest tenant sends. They break different things, so plan for each separately.
| Stage | What breaks | Fix |
|---|---|---|
| Tens of tenants | Nothing; one deployment, one quota | Measure per-tenant tokens and turns from day one |
| Hundreds | Unbounded per-tenant caches grow; a busy tenant crowds others | Bounded runner cache, per-tenant limits |
| Thousands | One quota cannot be shared by static limits; one bad deploy hits all | Fair-share allocator, cells |
| One giant tenant | A single tenant needs more quota than a shared cell holds | Dedicated cell with its own quota |
The common mistake is to treat this as a pod-autoscaling problem. Agent turns are I/O bound: a pod spends most of a turn waiting on the model and on tools. Pods are cheap to add. Model quota, cold-start cost and blast radius are not, and those are what the design below manages.
The architecture: cells
The architecture has one new idea: the cell. A cell is a complete copy of the serving stack (pods, runner cache, session store, quota allocator) with its own slice of model quota. Every tenant is assigned to exactly one cell by a placement table. The gateway authenticates the request, resolves the tenant, and the cell router forwards it to that tenant's cell.
Cells give you three things. A bad deploy rolls out cell by cell, so it reaches a fraction of tenants first. Quota is held per cell, so a tenant that exhausts its cell's slice cannot starve tenants elsewhere. And capacity grows by adding cells rather than by making one deployment ever larger, which keeps each cell's behaviour within the range you have already tested. The placement table carries a version number, so the router can tell a stale cached entry from a current one during a move.
A bounded runner cache
The simplest tenant registry is a ConcurrentHashMap<String, Runner> filled with computeIfAbsent. It is correct and it never shrinks. At ten thousand tenants each with an agent tree, tool objects, HTTP clients and per-tenant configuration, it holds every tenant ever seen since the pod started. Replace it with a bounded cache that evicts tenants nobody has used recently:
public final class TenantRunnerCache {
private final LoadingCache<String, Runner> cache;
public TenantRunnerCache(Function<String, Runner> factory, MeterRegistry meters) {
this.cache = Caffeine.newBuilder()
.maximumSize(2_000) // sized from a heap measurement
.expireAfterAccess(Duration.ofMinutes(30))
.removalListener((String tenant, Runner r, RemovalCause cause) ->
meters.counter("tenant_runner_evicted", "cause", cause.name()).increment())
.build(tenant -> {
Timer.Sample t = Timer.start(meters);
Runner runner = factory.apply(tenant); // loads config, builds agent and tools
t.stop(meters.timer("tenant_runner_build"));
return runner;
});
}
public Runner forTenant(String tenantId) { return cache.get(tenantId); }
}Two properties make eviction safe. First, the runner must hold no conversation state of its own. Sessions live in the session service, so the factory must hand every runner the cell's durable session service. If each runner owned an in-memory session service, evicting the runner would silently delete every open conversation of that tenant. Second, an evicted runner may still be serving a turn that started before eviction. That is fine as long as eviction only drops the cache reference and does not close shared clients; let the turn finish and the garbage collector reclaim the rest.
Size the cache from measurement, not guesswork. Build a few hundred runners in a test, take a heap histogram, and divide. Then watch tenant_runner_build: if its p99 is two seconds because the factory fetches remote configuration, a tenant returning after eviction pays two seconds on its first turn, and a cache that is too small turns that into a constant tax. The eviction counter by cause tells you which limit is biting.
Capacity arithmetic: find the binding resource
Before choosing cell sizes, find the binding resource. Little's law says the number of turns in flight equals arrival rate times duration. Take a platform with these assumed figures, chosen for illustration:
- 2,000 small tenants at 0.4 turns per minute each at peak: 800 turns/min.
- 40 medium tenants at 8 turns per minute: 320 turns/min.
- 1 large tenant at 600 turns per minute.
- An average turn uses 8,000 tokens across its model calls and lasts 5 seconds.
- A provider quota of 5,000,000 tokens per minute per project, planned to 70 percent (3,500,000), leaving headroom for retries and bursts.
Total peak demand is 1,720 turns per minute, which is 13.76 million tokens per minute. One project's planned budget covers 3,500,000 / 8,000 = 437 turns per minute. Turns in flight for a full cell are 437 / 60 × 5 = about 36. A single JVM handles 36 concurrent I/O-bound turns without strain, so pods are not the constraint: even with three replicas for availability, each pod is mostly idle. Quota is the constraint.
The large tenant alone needs 600 × 8,000 = 4.8 million tokens per minute, more than a shared cell's planned budget. It gets a dedicated cell with its own project and a quota increase negotiated for it. The remaining 1,120 turns per minute need 1,120 / 437 = 2.6 cells, so three shared cells. That is the arithmetic behind the diagram above, and it is worth redoing every quarter with measured numbers.
Sharing one quota fairly
Inside a cell, hundreds of tenants share one quota. Static per-tenant limits waste it: most tenants are idle most of the time, so their reserved share sits unused while a busy tenant is throttled. The standard answer is max-min fairness computed by water-filling. Every few seconds, collect each tenant's recent demand and a weight from its plan, then give every tenant the smaller of its demand and an equal weighted share, and redistribute what the satisfied tenants did not use:
/** Max-min fair split of a cell's tokens-per-minute across tenants by weight. */
static Map<String, Double> allocate(double capacity, Map<String, Double> demand,
Map<String, Double> weight) {
Map<String, Double> grant = new HashMap<>();
Set<String> open = new HashSet<>(demand.keySet());
double left = capacity;
while (!open.isEmpty() && left > 1e-9) {
double w = open.stream().mapToDouble(weight::get).sum();
double perWeight = left / w;
List<String> satisfied = open.stream()
.filter(t -> demand.get(t) - grant.getOrDefault(t, 0.0) <= perWeight * weight.get(t))
.toList();
if (satisfied.isEmpty()) { // nobody is satisfied: split the rest by weight
for (String t : open) grant.merge(t, perWeight * weight.get(t), Double::sum);
return grant;
}
for (String t : satisfied) { // give satisfied tenants exactly what they asked
double need = demand.get(t) - grant.getOrDefault(t, 0.0);
grant.merge(t, need, Double::sum);
left -= need;
open.remove(t);
}
}
return grant;
}The grants become the refill rates of per-tenant token buckets that the admission layer and a model callback consult before each call. Recompute every 5 to 10 seconds from a sliding window of demand, including requests that were refused, or a throttled tenant's demand looks artificially small and it never earns more. A worked case: capacity 3.5 million, weights equal, demands of 2.5 million, 0.6 million and 0.2 million plus an idle tenant. The idle and smaller tenants are satisfied at 0.2 and 0.6, and the busy tenant needs 2.5 million, which fits inside the 2.7 million left, so it is fully served and 0.2 million stays spare. When a second busy tenant arrives with 2.5 million, the two split the 2.7 million left after the small ones, 1.35 million each, and both are throttled equally instead of first-come-first-served.
Placing tenants into cells
Placement decides which cell each tenant lives in. It is a bin-packing problem with peak tokens per minute as the item size and each cell's planned budget as the bin. First-fit decreasing is simple, explainable and within a small factor of optimal:
record Tenant(String id, double peakTpm) {}
record Cell(String id, double budgetTpm, List<Tenant> tenants) {
double used() { return tenants.stream().mapToDouble(Tenant::peakTpm).sum(); }
}
static List<Cell> place(List<Tenant> tenants, List<Cell> cells, double dedicatedOver) {
tenants.stream()
.sorted(Comparator.comparingDouble(Tenant::peakTpm).reversed())
.forEach(t -> {
if (t.peakTpm() > dedicatedOver) { throw new IllegalStateException(t.id() + " needs a dedicated cell"); }
cells.stream()
.filter(c -> c.used() + t.peakTpm() <= c.budgetTpm())
.findFirst()
.orElseThrow(() -> new IllegalStateException("add a cell"))
.tenants().add(t);
});
return cells;
}Run the planner offline against measured peaks, review its diff, and apply it as a new placement table version. Do not re-run it on every request or every hour: each reassignment is a tenant move with a cost, so a placement that is stable and slightly unbalanced beats one that is optimal and churning. Use the 95th percentile of daily peaks rather than the mean, and keep a reserve of a few percent per cell for tenants that grow between planning runs.
Moving a hot tenant
Tenants grow. When one tenant's share of a cell's quota crosses a threshold, say 30 percent for a week, it becomes a candidate to move to a less loaded cell or to its own. A move has to preserve the tenant's open conversations, which live in the source cell's session store. The sequence that does so:
- Copy the tenant's sessions to the target cell's store while the source keeps serving.
- Publish placement version N+1 with the tenant marked moving: the router sends new sessions to the target and existing sessions to the source.
- Wait for the source's open sessions for that tenant to go idle, then copy the deltas.
- Publish version N+2 with the tenant fully on the target, and keep the source copy read-only for a rollback window before deleting it.
The router must compare the version on any cached placement entry with the current one, so a gateway instance that missed an update cannot send a turn to the cell that no longer owns the tenant. Rehearse a move on an internal tenant before you need one in an incident.
Failure modes
- Eviction deletes conversations. Runners built with their own in-memory session service lose sessions on eviction. Detection: users report forgotten context after quiet periods. Fix: one durable session service per cell, shared by all runners.
- Cold-start storms. After a deploy every runner is rebuilt at once, and a slow factory multiplies first-turn latency across thousands of tenants. Pre-warm the most active tenants from the previous pod's access log before taking traffic.
- Starved demand signal. Fair-share computed only from admitted requests never grows a throttled tenant's grant. Count refused requests as demand.
- Placement churn. A planner run on noisy data moves tenants back and forth. Apply a hysteresis band and a minimum dwell time per tenant.
- Split-brain during a move. Two cells serve one session when the router trusts a stale placement entry. Version every entry and refuse turns in the source once the tenant's version moves past it.
- Quota assumed, not held. A new cell's project quota defaults low. Raise and verify it before placing tenants there.
Trade-offs
Cells cost duplicated infrastructure and a placement system to operate, so a platform with a few hundred tenants and one quota is often better served by one deployment with a fair-share allocator. Dedicated cells for large tenants simplify their quota and compliance stories but leave capacity idle off-peak. Smaller runner caches save memory and cost latency on returning tenants. Max-min fairness protects small tenants at the price of capping a busy tenant that has paid for more; add weights by plan rather than abandoning the fairness. For the cluster-level side of capacity, the signal to autoscale pods on is covered in Running ADK Java Agents on Kubernetes, and placing cells in several regions is covered in Multi-Region Deployment for ADK Java. The end-to-end build and release path for those cells is in ADK Java deployment: packaging, sessions and JVM tuning.
What to do next
- Emit tokens, turns and refusals per tenant per minute, and chart the top 20 tenants' share of total quota.
- Replace any unbounded tenant map with a bounded cache, and check that every runner shares one durable session service.
- Measure runner build time and per-runner heap, then size the cache from those numbers.
- Redo the capacity arithmetic with your measured tokens per turn and quota, and write down which resource runs out first.
- Add a max-min fair allocator in front of model calls and count refused requests as demand.
- Introduce a versioned placement table, even with a single cell, so the second cell is a configuration change.
- Rehearse moving one internal tenant between cells with open sessions.