A multi-tenant agent platform serves many customers from one code base, and customers do not want the same agent. One pays for a stronger model, one has the refund tool and one does not, one needs answers in German and a stricter tone, one has a lower spend ceiling. Without a deliberate place for those differences, they end up as if (tenantId.equals("acme")) branches scattered through tools and prompts, and every new customer becomes a code change and a deployment.
Tenant config is that deliberate place: a typed, versioned record per tenant that the runtime reads to decide how the agent behaves for that tenant's requests. This article designs the record, shows where its values are applied in ADK for Java (some when an agent is built, some on every model and tool call), and covers rollout, rollback, auditing and the failure modes that turn configuration into a cross-tenant incident. It assumes the isolation model from tenant isolation; how layers of configuration override each other is in config layering and precedence.
What belongs in tenant config
The first design decision is what does not belong. Secrets do not: a tenant's API credentials belong in a secret manager, and the config record holds only a reference. Code does not: a tenant config that can name an arbitrary class to load is a plugin system with no review. Free-form prompts are a grey area, covered below. What remains is a small set of fields, each with a clear point of application.
| Field | Example | Applied at | Changes |
|---|---|---|---|
| Model tier | standard or premium | Agent build | Rarely, with the contract |
| Tool entitlements | [lookupOrder, issueRefund] | Agent build and every tool call | On upsell or incident |
| Tool limits | max refund 200.00 EUR | Every tool call | Occasionally |
| Instruction overlay | Reply in German, formal register | Every model call | Occasionally |
| Sampling and output caps | temperature 0.2, 1,024 tokens | Every model call | Rarely |
| Budgets | 40 tool calls per invocation | Every tool call | Rarely |
| Feature flags | new_triage_flow: true | Agent build | Often, during rollouts |
| Data region | eu-west | Routing, before the Runner | Almost never |
The architecture
A typed, versioned record
Make the record a typed value, validated in its constructor, so an invalid config cannot exist in memory. Store it as JSON with an explicit schema version and a monotonically increasing config version per tenant.
public record TenantConfig(
String tenantId,
long version, // increases on every change, never reused
ModelTier modelTier,
Set<String> allowedTools,
BigDecimal maxRefund,
String instructionOverlay, // validated text, not a template engine
float temperature,
int maxOutputTokens,
int maxToolCallsPerInvocation,
Map<String, Boolean> flags) {
public TenantConfig {
Objects.requireNonNull(tenantId);
if (version < 1) throw new IllegalArgumentException("version");
if (temperature < 0f || temperature > 1f) throw new IllegalArgumentException("temperature");
if (maxOutputTokens < 64 || maxOutputTokens > 8192) throw new IllegalArgumentException("maxOutputTokens");
if (instructionOverlay.length() > 2000) throw new IllegalArgumentException("overlay too long");
allowedTools = Set.copyOf(allowedTools);
flags = Map.copyOf(flags);
}
public boolean flag(String name) { return flags.getOrDefault(name, false); }
}{ "schema": 3, "tenantId": "acme", "version": 7, "modelTier": "premium",
"allowedTools": ["lookupOrder", "issueRefund"], "maxRefund": "200.00",
"instructionOverlay": "Answer in German using the formal register.",
"temperature": 0.2, "maxOutputTokens": 1024, "maxToolCallsPerInvocation": 40,
"flags": { "new_triage_flow": true } }The validator in front of the store does more than the constructor can: it checks that every allowed tool exists in the platform's tool registry, that the tenant's contract entitles them to the model tier and tools they asked for, and that the overlay passes a content check. Startup-time validation patterns are in config validation at startup.
Where the tenant id comes from
Everything depends on knowing which tenant a request belongs to, and that answer must come from authentication, never from the conversation. The edge service verifies the caller's token, reads the tenant claim and selects the tenant's Runner. If you also want the tenant id visible to tools and callbacks, pass it as server-side state when the run starts, using the stateDelta argument of runAsync, and treat it as read-only.
Do not let the model or a tool set it. Session state is writable from inside the run (tools can write it and agents can store output in it), so a value read from state is only as trustworthy as the code paths that can write that key. The robust pattern is to fix the tenant by construction: each Runner is built for exactly one tenant, and its plugin holds that tenant's config, so no lookup inside the run can be steered to another tenant.
// Edge handler: the tenant comes from the verified token, never from the request body.
String tenantId = jwt.getClaim("tenant");
Runner runner = tenantRunners.forTenant(tenantId);
Flowable<Event> events = runner.runAsync(
userId, sessionId, Content.fromParts(Part.fromText(message)),
RunConfig.builder().build(),
Map.of("tenant_id", tenantId)); // informational copy for tools; not the source of truthSessions deserve the same care. If all tenants share one session service and one app name, a session id guessed or replayed from another tenant must not resolve. Prefix user ids with the tenant, or check the session's recorded tenant against the token before running, so that a valid session id is never enough on its own.
Applying config: build time and call time
Values split cleanly into build-time and call-time settings. Build-time values shape the agent tree: which BaseLlm it uses, which tools are declared, the base instruction and any flag-gated sub-agents. Changing the model by rewriting the request's model name inside a callback is the wrong tool, because the flow calls the BaseLlm the agent was built with. Build a tree per tenant config version instead, and cache the resulting Runner.
public final class TenantRunners {
// Sketch: use a bounded cache with idle eviction in production (see failure modes).
private final ConcurrentHashMap<String, Runner> runners = new ConcurrentHashMap<>();
private final TenantConfigCache configs;
private final BaseSessionService sessions; // shared, keyed by user and session
private final Map<ModelTier, BaseLlm> models;
private final ToolCatalog catalog;
public Runner forTenant(String tenantId) {
TenantConfig cfg = configs.current(tenantId); // last known good on store failure
String key = tenantId + "@" + cfg.version();
return runners.computeIfAbsent(key, k -> build(cfg));
}
private Runner build(TenantConfig cfg) {
LlmAgent agent = LlmAgent.builder()
.name("support")
.model(models.get(cfg.modelTier()))
.instruction(BASE_INSTRUCTION)
.tools(catalog.resolve(cfg.allowedTools())) // only entitled tools are declared
.build();
return Runner.builder()
.agent(agent)
.appName("support")
.sessionService(sessions)
.plugins(new TenantConfigPlugin(cfg))
.build();
}
}Call-time values are applied by a plugin that holds the same snapshot. Before each model call it appends the overlay and tightens sampling; before each tool call it enforces entitlements, argument limits and the call budget, returning a refusal map instead of running the tool.
public final class TenantConfigPlugin extends BasePlugin {
private final TenantConfig cfg;
private final ConcurrentHashMap<String, AtomicInteger> toolCalls = new ConcurrentHashMap<>();
public TenantConfigPlugin(TenantConfig cfg) { super("tenant-config"); this.cfg = cfg; }
@Override
public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder req) {
GenerateContentConfig base = req.config().orElse(GenerateContentConfig.builder().build());
req.config(base.toBuilder()
.temperature(cfg.temperature())
.maxOutputTokens(cfg.maxOutputTokens())
.build());
if (!cfg.instructionOverlay().isBlank()) req.appendInstructions(List.of(cfg.instructionOverlay()));
return Maybe.empty(); // mutate and continue
}
@Override
public Maybe<Map<String, Object>> beforeToolCallback(BaseTool tool, Map<String, Object> args, ToolContext ctx) {
if (!cfg.allowedTools().contains(tool.name()))
return Maybe.just(Map.of("error", "tool_not_enabled", "tool", tool.name()));
int n = toolCalls.computeIfAbsent(ctx.invocationId(), k -> new AtomicInteger()).incrementAndGet();
if (n > cfg.maxToolCallsPerInvocation())
return Maybe.just(Map.of("error", "tool_budget_exhausted"));
if (tool.name().equals("issueRefund")) {
BigDecimal amount;
try { amount = new BigDecimal(String.valueOf(args.get("amount"))); }
catch (NumberFormatException e) { // a throwing hook fails the invocation
return Maybe.just(Map.of("error", "invalid_refund_amount"));
}
if (amount.compareTo(cfg.maxRefund()) > 0)
return Maybe.just(Map.of("error", "refund_above_tenant_limit", "limit", cfg.maxRefund().toString()));
}
return Maybe.empty();
}
@Override
public Completable afterRunCallback(InvocationContext ctx) {
toolCalls.remove(ctx.invocationId());
return Completable.complete();
}
@Override
public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
toolCalls.remove(ctx.invocationId()); // the error path must clean up too
return Completable.complete();
}
}Declaring only entitled tools and still checking in beforeToolCallback is deliberate defence in depth: the declaration keeps the model from trying, the check stops a call that arrives anyway, for example through a shared sub-agent built for another purpose. Because this plugin returns values from beforeToolCallback, register it after observers so metering still sees refused calls.
Pinning, rollout and rollback
Because the Runner is keyed by tenant and version, an invocation that starts on version 7 finishes on version 7 even if version 8 lands mid-run: the next request picks up the new Runner. That pinning is what makes a config change safe to reason about, and it is the same snapshot discipline described in dynamic config reload.
Rollout works on the version pointer. The store is append-only; each tenant has a current pointer, and changing a field writes version N+1 and moves the pointer. Roll out risky changes to a few canary tenants first and watch their error and refusal rates. Rollback moves the pointer back to N, which the cache picks up through the change feed within seconds, with no deployment. Every pointer move records who, when, why and the diff, because "what was acme's config at 14:02 on Tuesday?" is the first question in any tenant incident.
Make the version observable. Tag every trace span, metric and log line from a run with the tenant id and config version, and emit a counter for refusals by reason. When a tenant reports that the agent "stopped issuing refunds this morning", the dashboard should show a jump in refund_above_tenant_limit refusals that starts exactly at the version change, which turns a vague complaint into a one-line diff.
Worked example: the refund add-on
Acme buys the refund add-on. An operator submits a change adding issueRefund with a 200 EUR limit. The validator confirms the contract entitles Acme to the tool and the store writes version 7; the change feed invalidates pod caches. Acme's next request builds a new Runner keyed acme@7 whose agent declares the refund tool.
A customer asks for a 350 EUR refund. The model calls issueRefund with amount 350. The tenant plugin returns {"error": "refund_above_tenant_limit", "limit": "200.00"} instead of running the tool, the model receives that as the tool result and tells the customer the request needs a human. Another tenant on version 3 of its own config, without the add-on, never sees the tool declared at all. An hour later the operator learns the limit should be 150; version 8 goes out, the acme@7 Runner drains and is evicted once idle.
Failure modes
- Config store outage. Falling back to a global default silently downgrades or, worse, upgrades every tenant. Serve the last known good snapshot and alert; refuse new tenants with no cached config.
- Cache keyed by tenant only. Old Runners keep serving after a change. Key by tenant and version.
- Unbounded Runner cache. Thousands of tenants times versions exhaust memory. Bound the cache and evict idle entries; agent trees are cheap to rebuild.
- Prompt injection through overlays. A tenant admin writes "ignore previous rules" into the overlay. Treat overlays as untrusted input: length limits, content checks, review for new tenants, and keep safety rules in the base instruction and in plugins, not in the prompt alone.
- Flag sprawl. Hundreds of flags nobody removes. Give each an owner and an expiry date, and keep the flag backend behind your own interface so it can be replaced.
- Cross-tenant leakage via statics. A static cache inside a tool keyed only by user id serves one tenant's data to another. Key everything by tenant.
Trade-offs
| Design | Strength | Cost |
|---|---|---|
| Runner per tenant and version | Tenant fixed by construction; clean pinning | Memory per active tenant; warm-up per version |
| Shared Runner, lookup per call | One object graph; trivial memory | Every hook must find the tenant correctly; model and tools cannot vary |
| Code branches per tenant | Fast for the first two customers | Deploys per customer change; untestable combinations |
Per-tenant service levels built on the same per-tenant plumbing are covered in tenant SLAs.
What to do next
- Grep your agents, tools and prompts for tenant-specific branches and list each as a candidate config field.
- Define the typed record with constructor validation and a schema version.
- Put a validator and an append-only versioned store with a current pointer in front of it, with an audit entry per change.
- Key Runners by tenant and version and bound the cache.
- Write the tenant plugin for per-call limits and refusals, and test it with a fake model.
- Rehearse a rollback by moving a canary tenant's pointer back and timing propagation.