A multi-tenant agent platform serves many customers from one code base, and customers do not want the same agent. One pays for a stronger model, one has the refund tool and one does not, one needs answers in German and a stricter tone, one has a lower spend ceiling. Without a deliberate place for those differences, they end up as if (tenantId.equals("acme")) branches scattered through tools and prompts, and every new customer becomes a code change and a deployment.

Tenant config is that deliberate place: a typed, versioned record per tenant that the runtime reads to decide how the agent behaves for that tenant's requests. This article designs the record, shows where its values are applied in ADK for Java (some when an agent is built, some on every model and tool call), and covers rollout, rollback, auditing and the failure modes that turn configuration into a cross-tenant incident. It assumes the isolation model from tenant isolation; how layers of configuration override each other is in config layering and precedence.

What belongs in tenant config

The first design decision is what does not belong. Secrets do not: a tenant's API credentials belong in a secret manager, and the config record holds only a reference. Code does not: a tenant config that can name an arbitrary class to load is a plugin system with no review. Free-form prompts are a grey area, covered below. What remains is a small set of fields, each with a clear point of application.

FieldExampleApplied atChanges
Model tierstandard or premiumAgent buildRarely, with the contract
Tool entitlements[lookupOrder, issueRefund]Agent build and every tool callOn upsell or incident
Tool limitsmax refund 200.00 EUREvery tool callOccasionally
Instruction overlayReply in German, formal registerEvery model callOccasionally
Sampling and output capstemperature 0.2, 1,024 tokensEvery model callRarely
Budgets40 tool calls per invocationEvery tool callRarely
Feature flagsnew_triage_flow: trueAgent buildOften, during rollouts
Data regioneu-westRouting, before the RunnerAlmost never

The architecture

Tenant config: written rarely, validated once, read on every requestAdmin APItenant or operatorValidatorschema, entitlementsConfig storeversioned, append-onlyChange feedtenant, versionPod cachelast known goodEdge authtenant id from tokenRunner cachekey: tenant + versionAgent treemodel, tools, promptConfig pluginper-call limitsModel and toolsinvalidatesnapshottenant id
Writes pass a validator into an append-only versioned store; a change feed refreshes pod caches. Requests select a Runner by tenant and config version, and a plugin applies per-call limits.

A typed, versioned record

Make the record a typed value, validated in its constructor, so an invalid config cannot exist in memory. Store it as JSON with an explicit schema version and a monotonically increasing config version per tenant.

public record TenantConfig(
    String tenantId,
    long version,                       // increases on every change, never reused
    ModelTier modelTier,
    Set<String> allowedTools,
    BigDecimal maxRefund,
    String instructionOverlay,          // validated text, not a template engine
    float temperature,
    int maxOutputTokens,
    int maxToolCallsPerInvocation,
    Map<String, Boolean> flags) {

  public TenantConfig {
    Objects.requireNonNull(tenantId);
    if (version < 1) throw new IllegalArgumentException("version");
    if (temperature < 0f || temperature > 1f) throw new IllegalArgumentException("temperature");
    if (maxOutputTokens < 64 || maxOutputTokens > 8192) throw new IllegalArgumentException("maxOutputTokens");
    if (instructionOverlay.length() > 2000) throw new IllegalArgumentException("overlay too long");
    allowedTools = Set.copyOf(allowedTools);
    flags = Map.copyOf(flags);
  }

  public boolean flag(String name) { return flags.getOrDefault(name, false); }
}
{ "schema": 3, "tenantId": "acme", "version": 7, "modelTier": "premium",
  "allowedTools": ["lookupOrder", "issueRefund"], "maxRefund": "200.00",
  "instructionOverlay": "Answer in German using the formal register.",
  "temperature": 0.2, "maxOutputTokens": 1024, "maxToolCallsPerInvocation": 40,
  "flags": { "new_triage_flow": true } }

The validator in front of the store does more than the constructor can: it checks that every allowed tool exists in the platform's tool registry, that the tenant's contract entitles them to the model tier and tools they asked for, and that the overlay passes a content check. Startup-time validation patterns are in config validation at startup.

Where the tenant id comes from

Everything depends on knowing which tenant a request belongs to, and that answer must come from authentication, never from the conversation. The edge service verifies the caller's token, reads the tenant claim and selects the tenant's Runner. If you also want the tenant id visible to tools and callbacks, pass it as server-side state when the run starts, using the stateDelta argument of runAsync, and treat it as read-only.

Do not let the model or a tool set it. Session state is writable from inside the run (tools can write it and agents can store output in it), so a value read from state is only as trustworthy as the code paths that can write that key. The robust pattern is to fix the tenant by construction: each Runner is built for exactly one tenant, and its plugin holds that tenant's config, so no lookup inside the run can be steered to another tenant.

// Edge handler: the tenant comes from the verified token, never from the request body.
String tenantId = jwt.getClaim("tenant");
Runner runner = tenantRunners.forTenant(tenantId);
Flowable<Event> events = runner.runAsync(
    userId, sessionId, Content.fromParts(Part.fromText(message)),
    RunConfig.builder().build(),
    Map.of("tenant_id", tenantId));          // informational copy for tools; not the source of truth

Sessions deserve the same care. If all tenants share one session service and one app name, a session id guessed or replayed from another tenant must not resolve. Prefix user ids with the tenant, or check the session's recorded tenant against the token before running, so that a valid session id is never enough on its own.

Applying config: build time and call time

Values split cleanly into build-time and call-time settings. Build-time values shape the agent tree: which BaseLlm it uses, which tools are declared, the base instruction and any flag-gated sub-agents. Changing the model by rewriting the request's model name inside a callback is the wrong tool, because the flow calls the BaseLlm the agent was built with. Build a tree per tenant config version instead, and cache the resulting Runner.

public final class TenantRunners {
  // Sketch: use a bounded cache with idle eviction in production (see failure modes).
  private final ConcurrentHashMap<String, Runner> runners = new ConcurrentHashMap<>();
  private final TenantConfigCache configs;
  private final BaseSessionService sessions;            // shared, keyed by user and session
  private final Map<ModelTier, BaseLlm> models;
  private final ToolCatalog catalog;

  public Runner forTenant(String tenantId) {
    TenantConfig cfg = configs.current(tenantId);       // last known good on store failure
    String key = tenantId + "@" + cfg.version();
    return runners.computeIfAbsent(key, k -> build(cfg));
  }

  private Runner build(TenantConfig cfg) {
    LlmAgent agent = LlmAgent.builder()
        .name("support")
        .model(models.get(cfg.modelTier()))
        .instruction(BASE_INSTRUCTION)
        .tools(catalog.resolve(cfg.allowedTools()))     // only entitled tools are declared
        .build();
    return Runner.builder()
        .agent(agent)
        .appName("support")
        .sessionService(sessions)
        .plugins(new TenantConfigPlugin(cfg))
        .build();
  }
}

Call-time values are applied by a plugin that holds the same snapshot. Before each model call it appends the overlay and tightens sampling; before each tool call it enforces entitlements, argument limits and the call budget, returning a refusal map instead of running the tool.

public final class TenantConfigPlugin extends BasePlugin {
  private final TenantConfig cfg;
  private final ConcurrentHashMap<String, AtomicInteger> toolCalls = new ConcurrentHashMap<>();

  public TenantConfigPlugin(TenantConfig cfg) { super("tenant-config"); this.cfg = cfg; }

  @Override
  public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder req) {
    GenerateContentConfig base = req.config().orElse(GenerateContentConfig.builder().build());
    req.config(base.toBuilder()
        .temperature(cfg.temperature())
        .maxOutputTokens(cfg.maxOutputTokens())
        .build());
    if (!cfg.instructionOverlay().isBlank()) req.appendInstructions(List.of(cfg.instructionOverlay()));
    return Maybe.empty();                                 // mutate and continue
  }

  @Override
  public Maybe<Map<String, Object>> beforeToolCallback(BaseTool tool, Map<String, Object> args, ToolContext ctx) {
    if (!cfg.allowedTools().contains(tool.name()))
      return Maybe.just(Map.of("error", "tool_not_enabled", "tool", tool.name()));
    int n = toolCalls.computeIfAbsent(ctx.invocationId(), k -> new AtomicInteger()).incrementAndGet();
    if (n > cfg.maxToolCallsPerInvocation())
      return Maybe.just(Map.of("error", "tool_budget_exhausted"));
    if (tool.name().equals("issueRefund")) {
      BigDecimal amount;
      try { amount = new BigDecimal(String.valueOf(args.get("amount"))); }
      catch (NumberFormatException e) {                 // a throwing hook fails the invocation
        return Maybe.just(Map.of("error", "invalid_refund_amount"));
      }
      if (amount.compareTo(cfg.maxRefund()) > 0)
        return Maybe.just(Map.of("error", "refund_above_tenant_limit", "limit", cfg.maxRefund().toString()));
    }
    return Maybe.empty();
  }

  @Override
  public Completable afterRunCallback(InvocationContext ctx) {
    toolCalls.remove(ctx.invocationId());
    return Completable.complete();
  }

  @Override
  public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
    toolCalls.remove(ctx.invocationId());              // the error path must clean up too
    return Completable.complete();
  }
}

Declaring only entitled tools and still checking in beforeToolCallback is deliberate defence in depth: the declaration keeps the model from trying, the check stops a call that arrives anyway, for example through a shared sub-agent built for another purpose. Because this plugin returns values from beforeToolCallback, register it after observers so metering still sees refused calls.

Pinning, rollout and rollback

Because the Runner is keyed by tenant and version, an invocation that starts on version 7 finishes on version 7 even if version 8 lands mid-run: the next request picks up the new Runner. That pinning is what makes a config change safe to reason about, and it is the same snapshot discipline described in dynamic config reload.

Rollout works on the version pointer. The store is append-only; each tenant has a current pointer, and changing a field writes version N+1 and moves the pointer. Roll out risky changes to a few canary tenants first and watch their error and refusal rates. Rollback moves the pointer back to N, which the cache picks up through the change feed within seconds, with no deployment. Every pointer move records who, when, why and the diff, because "what was acme's config at 14:02 on Tuesday?" is the first question in any tenant incident.

Make the version observable. Tag every trace span, metric and log line from a run with the tenant id and config version, and emit a counter for refusals by reason. When a tenant reports that the agent "stopped issuing refunds this morning", the dashboard should show a jump in refund_above_tenant_limit refusals that starts exactly at the version change, which turns a vague complaint into a one-line diff.

Worked example: the refund add-on

Acme buys the refund add-on. An operator submits a change adding issueRefund with a 200 EUR limit. The validator confirms the contract entitles Acme to the tool and the store writes version 7; the change feed invalidates pod caches. Acme's next request builds a new Runner keyed acme@7 whose agent declares the refund tool.

A customer asks for a 350 EUR refund. The model calls issueRefund with amount 350. The tenant plugin returns {"error": "refund_above_tenant_limit", "limit": "200.00"} instead of running the tool, the model receives that as the tool result and tells the customer the request needs a human. Another tenant on version 3 of its own config, without the add-on, never sees the tool declared at all. An hour later the operator learns the limit should be 150; version 8 goes out, the acme@7 Runner drains and is evicted once idle.

Failure modes

  • Config store outage. Falling back to a global default silently downgrades or, worse, upgrades every tenant. Serve the last known good snapshot and alert; refuse new tenants with no cached config.
  • Cache keyed by tenant only. Old Runners keep serving after a change. Key by tenant and version.
  • Unbounded Runner cache. Thousands of tenants times versions exhaust memory. Bound the cache and evict idle entries; agent trees are cheap to rebuild.
  • Prompt injection through overlays. A tenant admin writes "ignore previous rules" into the overlay. Treat overlays as untrusted input: length limits, content checks, review for new tenants, and keep safety rules in the base instruction and in plugins, not in the prompt alone.
  • Flag sprawl. Hundreds of flags nobody removes. Give each an owner and an expiry date, and keep the flag backend behind your own interface so it can be replaced.
  • Cross-tenant leakage via statics. A static cache inside a tool keyed only by user id serves one tenant's data to another. Key everything by tenant.

Trade-offs

DesignStrengthCost
Runner per tenant and versionTenant fixed by construction; clean pinningMemory per active tenant; warm-up per version
Shared Runner, lookup per callOne object graph; trivial memoryEvery hook must find the tenant correctly; model and tools cannot vary
Code branches per tenantFast for the first two customersDeploys per customer change; untestable combinations

Per-tenant service levels built on the same per-tenant plumbing are covered in tenant SLAs.

What to do next

  1. Grep your agents, tools and prompts for tenant-specific branches and list each as a candidate config field.
  2. Define the typed record with constructor validation and a schema version.
  3. Put a validator and an append-only versioned store with a current pointer in front of it, with an audit entry per change.
  4. Key Runners by tenant and version and bound the cache.
  5. Write the tenant plugin for per-call limits and refusals, and test it with a fake model.
  6. Rehearse a rollback by moving a canary tenant's pointer back and timing propagation.
Key takeaway: Tenant config turns per-customer differences into data: a typed, validated, versioned record that the runtime applies at two points. Build-time fields shape a Runner cached per tenant and version, and a plugin applies per-call limits and refusals. Take the tenant from authentication, pin a version per invocation, roll back by moving a pointer, and serve the last known good config when the store fails.