A multi-tenant agent platform runs one codebase for many customers. Keeping their data apart is the first requirement, and it is covered in Agent Tenant Data Separation: how ADK Java keys sessions, state, artifacts and memory, and how to guard those keys. This article is about everything else tenants share: threads, model quota, tool credentials, outbound connections, caches and failure domains. A platform can have perfect data separation and still let one tenant's batch job starve everyone of model capacity, or let a prompt injection in one tenant's document call a tool holding another tenant's API key.
We will define what isolation means at runtime, choose between pooled, bridge and silo deployment, then build the enforcement points in ADK Java: an admission layer, a per-tenant runtime registry, a policy plugin on each tenant's Runner, and tools bound to tenant credentials when they are constructed. A worked noisy-neighbour example puts numbers on why the admission layer exists, and the article closes with failure modes and a checklist.
What runtime isolation means
Runtime isolation means three properties. Performance isolation: one tenant's load cannot push another tenant's latency or error rate past its SLO. Authority isolation: a run on behalf of tenant A can only use credentials, tools and model accounts provisioned for tenant A, whatever the model is tricked into asking for. Fault isolation: a crash, loop or poisoned dependency in one tenant's run is contained to that tenant. Data isolation is the fourth leg and already has its own page.
| Shared resource | How it leaks between tenants | Control |
|---|---|---|
| JVM threads and event loops | one tenant's slow tools hold all workers | per-tenant bulkhead |
| Model provider quota | a burst exhausts the project's requests or tokens per minute | per-tenant rate and token budget |
| Tool credentials | a shared client authenticates as the platform for every tenant | tools built per tenant with scoped secrets |
| Caches and pools | cached tool results or HTTP connections keyed without tenant | tenant in every cache key, or per-tenant caches |
| Agent loops | a runaway run burns budget and threads for minutes | per-run model call and token caps |
Pooled, bridge and silo
The first decision is how much infrastructure tenants share. In the pooled model all tenants run in the same deployment and isolation is enforced in code; it is cheapest and the default for small tenants. In the silo model each tenant gets its own deployment, model project and secrets, so isolation comes from infrastructure; it is what regulated or very large customers ask for, and it costs a deployment per tenant. The bridge model shares the agent code and gateway but gives a tenant dedicated worker pools and a dedicated model account. Most platforms end up with all three, and the trick is to make the agent code identical in each so that moving a tenant between tiers is a configuration change. That is why everything below keys off a TenantScope object rather than off global configuration.
The architecture: one runtime per tenant
The tenant is resolved exactly once, at the gateway, from the authenticated identity, never from a request field or a tool argument the model could fill in. From there the request carries a TenantScope and every enforcement point reads it. The runtime registry holds one fully built Runner per tenant: the same LlmAgent definition, but its own plugin instance, its own tool instances bound to that tenant's clients and its own appName. Building per tenant costs a few objects and removes a whole class of bugs, because no code path inside a run needs to look up the tenant again.
public record TenantScope(String tenantId, String appName, Tier tier,
int maxConcurrentRuns, double runsPerSecond,
long dailyTokenBudget, int maxModelCallsPerRun,
Set<String> allowedTools) {}
public final class TenantRuntimes {
private final Map<String, Runner> runners = new ConcurrentHashMap<>();
private final SecretStore secrets;
private final BaseSessionService sessions;
private final UsageLedger usage;
public Runner runnerFor(TenantScope t) {
return runners.computeIfAbsent(t.tenantId(), id -> build(t));
}
private Runner build(TenantScope t) {
CrmClient crm = new CrmClient(secrets.get(t.tenantId(), "crm")); // tenant's own key
CrmTools tools = new CrmTools(crm);
LlmAgent agent = LlmAgent.builder()
.name("support_agent")
.model("gemini-2.5-flash")
.instruction(SupportPrompts.INSTRUCTION)
.tools(FunctionTool.create(tools, "lookupAccount"),
FunctionTool.create(tools, "openTicket"))
.build();
return Runner.builder()
.agent(agent)
.appName(t.appName())
.sessionService(sessions)
.plugins(new TenantPolicyPlugin(t, usage))
.build();
}
}Note what the tools do not take: a tenant id parameter. lookupAccount searches through a client that can only see the tenant's CRM. If the model is tricked into asking for another customer's account, the request fails at the CRM with the tenant's own permissions, which is the property you want. Remove the runner from the map when a tenant's secrets rotate or the tenant is offboarded.
Admission control before the Runner
Admission control runs before ADK sees the request. It holds two limits per tenant: a bulkhead capping concurrent runs, and a rate limiter capping new runs per second. A bulkhead is just a semaphore whose permit is held for the life of the run, which in ADK Java is the life of the Flowable<Event> returned by runAsync. The general pattern is in ADK Java bulkhead isolation; here the bulkhead is keyed by tenant.
public Flowable<Event> submit(TenantScope t, String userId, String sessionId, Content msg) {
Semaphore slots = bulkheads.computeIfAbsent(t.tenantId(),
id -> new Semaphore(t.maxConcurrentRuns()));
RateLimiter rate = limiters.computeIfAbsent(t.tenantId(),
id -> RateLimiter.create(t.runsPerSecond()));
return Flowable.defer(() -> { // acquire per subscription, not per call
if (!rate.tryAcquire()) return Flowable.<Event>error(new TenantThrottled(t.tenantId(), "rate"));
if (!slots.tryAcquire()) return Flowable.<Event>error(new TenantThrottled(t.tenantId(), "concurrency"));
return runtimes.runnerFor(t)
.runAsync(userId, sessionId, msg)
.doFinally(slots::release); // complete, error or cancel
}).subscribeOn(Schedulers.from(executors.forTier(t.tier()))); // bridge tenants: own pool
}Fail fast with a typed error the gateway maps to HTTP 429 and a retry hint; queuing inside the JVM just moves the starvation somewhere you cannot see it. Acquire inside Flowable.defer so a stream that is never subscribed holds nothing, and release in doFinally so cancelled streams return their permit. The rate limiter shown is Guava's; any token bucket works, and in a multi-replica deployment the limits must either be divided across replicas or held in a shared store.
A tenant policy plugin
Admission bounds how many runs a tenant has; the policy plugin bounds what each run may do. ADK Java plugins registered on a Runner see every agent, model call and tool call in the tree, and a non-empty return short-circuits the step. Because each tenant has its own Runner, the plugin holds its TenantScope as a field and never has to trust anything in the request.
public final class TenantPolicyPlugin extends BasePlugin {
private final TenantScope t;
private final UsageLedger usage;
private final Map<String, AtomicInteger> callsPerRun = new ConcurrentHashMap<>();
public TenantPolicyPlugin(TenantScope t, UsageLedger usage) {
super("tenant_policy_" + t.tenantId());
this.t = t;
this.usage = usage;
}
@Override
public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
if (!t.appName().equals(ctx.session().appName())) // wiring bug: fail closed
return Maybe.just(Content.fromParts(Part.fromText("Request rejected.")));
if (usage.tokensToday(t.tenantId()) >= t.dailyTokenBudget())
return Maybe.just(Content.fromParts(Part.fromText(
"Your organisation has reached today's usage limit.")));
return Maybe.empty();
}
@Override
public Maybe<LlmResponse> beforeModelCallback(CallbackContext cb, LlmRequest.Builder req) {
int n = callsPerRun.computeIfAbsent(cb.invocationId(), k -> new AtomicInteger())
.incrementAndGet();
if (n > t.maxModelCallsPerRun())
return Maybe.just(LlmResponse.builder()
.content(Content.fromParts(Part.fromText("Stopping: step limit reached.")))
.build());
return Maybe.empty();
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext cb, LlmResponse resp) {
resp.usageMetadata().ifPresent(u ->
usage.add(t.tenantId(), u.totalTokenCount().orElse(0)));
return Maybe.empty();
}
@Override
public Maybe<Map<String, Object>> beforeToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext tc) {
if (!t.allowedTools().contains(tool.name()))
return Maybe.just(Map.of("status", "unavailable",
"instruction", "This action is not enabled for this organisation. Do not retry."));
return Maybe.empty();
}
}Clean up callsPerRun when a run ends, for example in afterRunCallback, or the map grows without bound. The daily budget check is deliberately approximate: it runs before the run, so one run can overshoot by its own size. Per-run call caps bound that overshoot. Feed the same ledger to billing, covered in Agent Tenant Billing, so the number that throttles a tenant is the number you invoice.
Worked example: the nightly replay job
A pooled deployment serves 40 tenants from one model project with a quota of 1,000 requests per minute. A typical support run makes 4 model calls and lasts 12 seconds. Normal load is 150 runs per minute across all tenants, 600 model requests per minute, comfortably inside quota.
Tenant Acme starts a nightly job that replays 3,000 old tickets through the agent as fast as its client can submit. Without admission control Acme alone generates hundreds of runs per minute; at 4 calls each, Acme wants several thousand requests per minute. The provider starts returning 429s to the whole project within seconds, and the other 39 tenants see failures although they changed nothing.
With a per-tenant limit of 30 runs per minute and 10 concurrent runs, Acme's job takes 100 minutes instead of ten, uses at most 120 requests per minute, and the project peaks at 720, under quota. Set the sum of tenant limits above quota to allow for idle tenants, but keep any single tenant's limit well below it; a useful rule is that the largest tenant limit plus normal load should stay under 90 percent of quota. Tenants who legitimately need more get the bridge tier with their own model account.
Testing isolation
Isolation claims need tests that try to break them. Run these in staging on every release: a flood test, in which one synthetic tenant submits at ten times its limit while a second tenant's p95 latency is measured and must stay within its SLO; a budget test that drives a tenant past its daily tokens and checks the next run is refused before any model call; a loop test with a tool that always asks to be called again, which must stop at the per-run cap; and an authority test that plants an instruction in tenant A's document telling the agent to open a ticket for tenant B, then checks B's CRM saw nothing. Prompt-injection defences in general are in ADK Java guardrails; the point here is that authority isolation holds even when those defences fail.
Failure modes
- Tenant from the request body. A field or tool argument names the tenant; one forged value crosses the boundary. Resolve it from authenticated claims only.
- Shared platform credential. Tools call SaaS APIs with a platform key and filter by tenant in code; one missing filter exposes every customer.
- Leaked permits. Bulkhead permits released only on success; after a few errors the tenant is permanently throttled. Release in
doFinally. - Per-replica limits. Limits enforced per pod multiply with replica count during a scale-out, exactly when the provider quota is tightest.
- Cache without tenant key. A tool result cache keyed by query text serves one tenant's answer to another.
- Stale runners. A rotated secret keeps working inside a cached Runner until restart, or a revoked one keeps failing. Evict on rotation and offboarding.
Trade-offs
Pooled deployment is cheap and simple to operate but makes performance isolation a code property you must test continuously. Silo deployment makes isolation easy to prove to an auditor and impossible to break with a code bug, at the cost of per-tenant infrastructure and slower rollouts. Per-tenant Runners trade a little memory and warm-up time for removing tenant lookups from hot paths. Strict limits protect neighbours but frustrate the tenant who hit them, so pair every limit with a visible error and an upgrade path. Token budgets checked before a run are approximate; exact enforcement needs a check before every model call, which adds latency to each step.
What to do next
- List every resource your agents share across tenants, using the table above as a template, and name the control for each.
- Move tenant resolution to the gateway and delete any tool parameter or request field that names a tenant.
- Build tools per tenant with
FunctionTool.create(instance, method)over clients that hold that tenant's credentials. - Add a per-tenant bulkhead and rate limit in front of
runAsync, releasing permits indoFinally, and return 429 with a retry hint. - Register a tenant policy plugin with a per-run model call cap, a daily token budget and a tool allowlist, and log every refusal with the tenant id.
- Write the flood, budget, loop and authority tests and run them on every release.
- Define the criteria that move a tenant from pooled to bridge or silo, and rehearse one move.