A new customer signs a contract and expects the agent to work for their people tomorrow. Behind that expectation sits a surprising amount of setup: an isolation boundary in ADK Java, a data key, a quota plan, a tool allowlist, configuration, secrets, a runner, and some proof that all of it works together. Teams that treat onboarding as a signup handler which writes one row usually find out months later that a tenant was half-created: it has a directory entry but no key, or a runner that points at the default application name and therefore shares data with everyone else.

This article treats tenant onboarding as a lifecycle with explicit states, the mirror image of tenant offboarding. You will see what a tenant needs before its first run, an idempotent provisioner you can crash and rerun, how to wire a per-tenant Runner, a smoke run that gates activation, and the failure modes that show up in production. The code is Java against ADK Java's runner and plugin APIs; the directory, key service and quota ledger are interfaces you back with your own stores.

Why onboarding is not a signup handler

Three facts drive the design. First, the only isolation boundary ADK Java gives you is the application name. Session, artifact and memory services are addressed by (appName, userId, sessionId), and app: state is shared by every user of an app, as tenant data separation shows in detail. If onboarding forgets to give the tenant its own app name, nothing downstream will notice. Second, onboarding touches several systems that do not share a transaction: a database for the directory, a key management service, a quota ledger, a secrets store and an in-memory runner registry. Any of them can fail after the others succeeded. Third, the last step that makes a tenant usable must be the only step that opens admission, so a partial tenant can never take traffic.

Those facts rule out the usual one-shot handler. What works is a small saga: a durable state field per tenant, one idempotent step per transition, and a resume loop that reads the state and runs the next step. Because each step can be repeated safely, a crash, a timeout or an operator rerun converges on the same end state instead of creating a second key or a second application name.

The lifecycle as a state machine

Tenant onboarding as a resumable state machineRequestedrecord writtenReservedappName + idsKeyeddata key in KMSConfiguredplan, tools, configWiredrunner registeredVerifiedsmoke run passedActiveadmission opensapproveFailed(step)retry resumes hereany step errorAbandonedhand to offboardinggive upEvery step is idempotent and keyed by tenantId, so a crash or retry resumesat the last completed state. Admission refuses every run until the state is Active.
Onboarding states. Each arrow is one idempotent step; failures keep the last completed state and the resume loop retries the next step. Only an explicit approval opens admission.

Each state records one durable fact, and the transition into it is what creates that fact. Requested means the contract system wrote a tenant record with a plan and a region. Reserved means the tenant id and application name are allocated and unique. Keyed means a data encryption key exists for this tenant. Configured means the plan's quota, the tool allowlist and the validated configuration are stored. Wired means a runner exists in the registry, still behind closed admission. Verified means a smoke run passed. Active is set by a separate approval, by a person or by a policy that says verified tenants on a given plan activate automatically.

Failure does not get its own linear state. Instead the record keeps the last completed state plus a lastError field, and the resume loop simply retries the next step. A tenant that cannot be completed, because the customer cancelled or a region is unavailable, moves to Abandoned and is handed to the offboarding lifecycle, which already knows how to delete keys, app names and anything partially written. That reuse matters: you write the deletion logic once.

What a tenant needs before its first run

Before writing code, list what the tenant must own. Each row below is created by exactly one step, and each has a check you can run later to prove it exists.

AssetCreated inWhy it is neededVerification
Tenant id, application nameReservedthe ADK isolation boundaryunique index; name matches pattern
Data encryption keyKeyedencrypts stored sessions, artifacts, exports; shredded at offboardingKMS describe returns enabled
Quota planConfiguredadmission and per-run budgetsledger row with plan limits
Tool allowlistConfiguredwhich tools the tenant may callevery name exists in the tool registry
Tenant configConfiguredmodel choice, instructions overlay, localepasses the startup validator
Secrets for tenant integrationsConfiguredthe tenant's own API credentialssecret reference resolves
RunnerWiredruns the agent tree under the tenant's app nameregistry lookup returns it

Create the data key on day one even if you do not encrypt everything with it yet. Offboarding by crypto-shredding only works if every tenant record, export and backup was written under a tenant key from the start; adding the key later leaves older data that the shred does not cover.

An idempotent provisioner

The provisioner below is the whole saga. The directory stores the state with an optimistic version, so two workers resuming the same tenant cannot both advance it. Every step is written to be repeatable: allocation uses the tenant id as the natural key, the key service is asked to create a key with an alias it can look up first, and configuration writes are upserts.

public final class TenantProvisioner {
  private final TenantDirectory dir;      // durable: state, version, lastError
  private final KeyService kms;           // create-or-get by alias
  private final QuotaLedger quotas;
  private final ConfigValidator validator;
  private final RunnerRegistry runners;
  private final SmokeTester smoke;

  // ctor omitted

  public TenantState resume(String tenantId) {
    while (true) {
      TenantRecord r = dir.load(tenantId);
      TenantState next = switch (r.state()) {
        case REQUESTED  -> reserve(r);
        case RESERVED   -> key(r);
        case KEYED      -> configure(r);
        case CONFIGURED -> wire(r);
        case WIRED      -> verify(r);
        case VERIFIED, ACTIVE, ABANDONED -> null;   // nothing automatic left
      };
      if (next == null) return r.state();
      // compare-and-set on version: a concurrent worker makes this fail, and we reload
      dir.advance(tenantId, r.version(), next);
    }
  }

  private TenantState reserve(TenantRecord r) {
    String app = "t_" + r.tenantId().replaceAll("[^a-z0-9]", "_");
    dir.reserveAppName(r.tenantId(), app);           // unique index; no-op if already ours
    return TenantState.RESERVED;
  }

  private TenantState key(TenantRecord r) {
    String alias = "tenant/" + r.tenantId();
    kms.findByAlias(alias).orElseGet(() -> kms.create(alias, r.region()));
    return TenantState.KEYED;
  }

  private TenantState configure(TenantRecord r) {
    TenantConfig cfg = TenantConfig.fromPlan(r.plan(), r.overrides());
    List<String> errors = validator.validate(cfg);   // collect-all, not first error
    if (!errors.isEmpty()) throw new ProvisioningException("config", errors);
    quotas.upsertPlan(r.tenantId(), cfg.quota());
    dir.saveConfig(r.tenantId(), cfg);
    return TenantState.CONFIGURED;
  }

  private TenantState wire(TenantRecord r) {
    runners.register(r.tenantId(), dir.loadConfig(r.tenantId())); // idempotent put
    return TenantState.WIRED;
  }

  private TenantState verify(TenantRecord r) {
    SmokeResult res = smoke.run(r.tenantId());
    if (!res.passed()) throw new ProvisioningException("smoke", res.failures());
    return TenantState.VERIFIED;
  }
}

A thrown ProvisioningException is caught by the job that called resume, which records lastError and schedules a retry with backoff. Notice that ACTIVE is not reachable from this loop. Activation is a separate call that checks the state is VERIFIED and records who approved it.

Wiring the tenant&#x27;s runner

Wiring builds one Runner per tenant with the reserved application name and the tenant's plugins. Runner.builder() accepts the agent, app name, session, artifact and memory services, and plugins. The shared services are stateless clients over durable stores; the app name is what keeps one tenant's rows apart from another's.

Runner buildRunner(TenantRecord r, TenantConfig cfg) {
  BaseAgent root = agentFactory.create(cfg);           // fresh tree per tenant: no shared mutable tools
  return Runner.builder()
      .agent(root)
      .appName(r.appName())                            // the isolation boundary
      .sessionService(sessions)
      .artifactService(artifacts)
      .memoryService(memory)
      .plugins(List.of(
          new TenantPolicyPlugin(TenantScope.of(r, cfg), usage),   // allowlist + budgets
          new AdmissionGatePlugin(r.tenantId(), dir)))             // refuses unless ACTIVE
      .build();
}

The admission gate is a beforeRunCallback that reads the tenant state from a short-lived cache and returns a refusal message unless the state is ACTIVE. Keeping the gate inside the runner, not only in the HTTP layer, means a queue consumer or a scheduled job that reaches the runner by another path is refused too. The policy plugin enforcing allowlists and budgets is the one from tenant isolation; onboarding only has to construct it with the tenant's scope.

The smoke run that gates activation

The smoke run is the step that turns configuration into evidence. It runs through the real runner, with admission bypassed by an explicit internal flag, and checks four things:

  1. One scripted turn completes. Send a fixed prompt that should call one harmless allowlisted tool and produce a final response. Assert the final event exists and that the tool's name appears in the event stream.
  2. Data lands under the right key. After the turn, list sessions for (appName, smokeUserId) and assert exactly one exists; list the same user under the default app name and assert none do.
  3. A disallowed tool is refused. A second prompt asks for a tool outside the allowlist; assert the policy plugin's refusal appears instead of a call.
  4. Usage is recorded. The quota ledger shows tokens charged to this tenant id, which is the same number billing will invoice.

Afterwards, delete the smoke session so it does not appear in the tenant's history or its usage report. Keep the smoke run cheap: one or two model calls on the tenant's configured model. If your model provider is billed per tenant project, this is also the first proof that the tenant's credentials work.

Worked example: one tenant, two retries

Walk one tenant through it. A contract for acme-eu on the Team plan in an EU region creates a record in REQUESTED. The worker reserves t_acme_eu and advances to RESERVED. It creates a key with alias tenant/acme-eu in the EU key ring, then crashes before writing KEYED.

On restart the resume loop reads RESERVED and runs the key step again. The alias lookup finds the key created before the crash and returns it, so there is still exactly one key. Configuration fails validation because the contract enabled a calendar tool that this region's tool registry does not offer; the record shows lastError=config: tool calendar_book not in registry. Support removes the tool from the overrides and reruns. Configuration passes, wiring registers the runner, and the smoke run catches one more problem: the tenant's instruction overlay told the agent to always ask a clarifying question, so the scripted turn never called its tool. The overlay is fixed, the smoke run passes, the state is VERIFIED, and an account manager activates it. Every one of those problems would otherwise have been found by the customer.

Failure modes

  • Default application name. A factory falls back to a shared app name when the tenant's is missing. Make the builder throw on a blank or default name, and keep the smoke check that lists the default app.
  • Duplicate keys or app names after retries. Steps that create instead of create-or-get. Give every created asset a deterministic name derived from the tenant id and look it up first.
  • Lost updates between workers. Two resumes advance the same tenant. The versioned compare-and-set on the state turns this into a harmless reload.
  • Runner registry drift. Runners live in memory per replica; a replica that started after onboarding has none. Build runners lazily from the directory on first use, and treat Wired as "config stored", not "object exists".
  • Activation without verification. An admin API sets ACTIVE directly. Enforce the allowed transitions in the directory itself, not in callers.
  • Smoke data leaking into reports. Use a reserved smoke user id, delete the session, and exclude that user from usage exports.

Trade-offs

A state machine costs more code than a handler, and it is worth it once more than a few tenants a month arrive or more than two systems are involved. Fully automatic activation shortens time-to-value but removes the last human look at a contract's special terms; many teams auto-activate self-serve plans and require approval for enterprise ones. Building runners eagerly at onboarding makes the first request fast but costs memory for tenants that never show up; lazy building, with the per-tenant runner cache bounded as described in tenant scaling, is usually the better default. Finally, a per-tenant data key adds a KMS dependency to every write path, which is a latency and availability cost you accept in exchange for a clean, provable offboarding later.

What to do next

  1. Write down your tenant states and the single durable fact each one records.
  2. Make every provisioning step create-or-get with a name derived from the tenant id, and test it by running each step twice.
  3. Put an admission gate plugin on every tenant runner that refuses anything not ACTIVE.
  4. Create a per-tenant data key at onboarding, before any tenant data is written.
  5. Build the four-check smoke run and make it the only path to VERIFIED.
  6. Route abandoned tenants into the offboarding lifecycle rather than writing a second cleanup path.
  7. Validate configuration with the collect-all approach in configuration validation at startup, and wire plan limits through quota management.
Key takeaway: Onboard tenants through a durable state machine whose steps are all create-or-get, so crashes and reruns converge instead of duplicating. Reserve a unique application name first, because it is the only boundary ADK Java enforces, and create the tenant's data key before any data exists. Wire a per-tenant runner behind an admission gate, prove it with a smoke run, and open traffic only through an explicit activation.