A new customer signs a contract and expects the agent to work for their people tomorrow. Behind that expectation sits a surprising amount of setup: an isolation boundary in ADK Java, a data key, a quota plan, a tool allowlist, configuration, secrets, a runner, and some proof that all of it works together. Teams that treat onboarding as a signup handler which writes one row usually find out months later that a tenant was half-created: it has a directory entry but no key, or a runner that points at the default application name and therefore shares data with everyone else.
This article treats tenant onboarding as a lifecycle with explicit states, the mirror image of tenant offboarding. You will see what a tenant needs before its first run, an idempotent provisioner you can crash and rerun, how to wire a per-tenant Runner, a smoke run that gates activation, and the failure modes that show up in production. The code is Java against ADK Java's runner and plugin APIs; the directory, key service and quota ledger are interfaces you back with your own stores.
Why onboarding is not a signup handler
Three facts drive the design. First, the only isolation boundary ADK Java gives you is the application name. Session, artifact and memory services are addressed by (appName, userId, sessionId), and app: state is shared by every user of an app, as tenant data separation shows in detail. If onboarding forgets to give the tenant its own app name, nothing downstream will notice. Second, onboarding touches several systems that do not share a transaction: a database for the directory, a key management service, a quota ledger, a secrets store and an in-memory runner registry. Any of them can fail after the others succeeded. Third, the last step that makes a tenant usable must be the only step that opens admission, so a partial tenant can never take traffic.
Those facts rule out the usual one-shot handler. What works is a small saga: a durable state field per tenant, one idempotent step per transition, and a resume loop that reads the state and runs the next step. Because each step can be repeated safely, a crash, a timeout or an operator rerun converges on the same end state instead of creating a second key or a second application name.
The lifecycle as a state machine
Each state records one durable fact, and the transition into it is what creates that fact. Requested means the contract system wrote a tenant record with a plan and a region. Reserved means the tenant id and application name are allocated and unique. Keyed means a data encryption key exists for this tenant. Configured means the plan's quota, the tool allowlist and the validated configuration are stored. Wired means a runner exists in the registry, still behind closed admission. Verified means a smoke run passed. Active is set by a separate approval, by a person or by a policy that says verified tenants on a given plan activate automatically.
Failure does not get its own linear state. Instead the record keeps the last completed state plus a lastError field, and the resume loop simply retries the next step. A tenant that cannot be completed, because the customer cancelled or a region is unavailable, moves to Abandoned and is handed to the offboarding lifecycle, which already knows how to delete keys, app names and anything partially written. That reuse matters: you write the deletion logic once.
What a tenant needs before its first run
Before writing code, list what the tenant must own. Each row below is created by exactly one step, and each has a check you can run later to prove it exists.
| Asset | Created in | Why it is needed | Verification |
|---|---|---|---|
| Tenant id, application name | Reserved | the ADK isolation boundary | unique index; name matches pattern |
| Data encryption key | Keyed | encrypts stored sessions, artifacts, exports; shredded at offboarding | KMS describe returns enabled |
| Quota plan | Configured | admission and per-run budgets | ledger row with plan limits |
| Tool allowlist | Configured | which tools the tenant may call | every name exists in the tool registry |
| Tenant config | Configured | model choice, instructions overlay, locale | passes the startup validator |
| Secrets for tenant integrations | Configured | the tenant's own API credentials | secret reference resolves |
| Runner | Wired | runs the agent tree under the tenant's app name | registry lookup returns it |
Create the data key on day one even if you do not encrypt everything with it yet. Offboarding by crypto-shredding only works if every tenant record, export and backup was written under a tenant key from the start; adding the key later leaves older data that the shred does not cover.
An idempotent provisioner
The provisioner below is the whole saga. The directory stores the state with an optimistic version, so two workers resuming the same tenant cannot both advance it. Every step is written to be repeatable: allocation uses the tenant id as the natural key, the key service is asked to create a key with an alias it can look up first, and configuration writes are upserts.
public final class TenantProvisioner {
private final TenantDirectory dir; // durable: state, version, lastError
private final KeyService kms; // create-or-get by alias
private final QuotaLedger quotas;
private final ConfigValidator validator;
private final RunnerRegistry runners;
private final SmokeTester smoke;
// ctor omitted
public TenantState resume(String tenantId) {
while (true) {
TenantRecord r = dir.load(tenantId);
TenantState next = switch (r.state()) {
case REQUESTED -> reserve(r);
case RESERVED -> key(r);
case KEYED -> configure(r);
case CONFIGURED -> wire(r);
case WIRED -> verify(r);
case VERIFIED, ACTIVE, ABANDONED -> null; // nothing automatic left
};
if (next == null) return r.state();
// compare-and-set on version: a concurrent worker makes this fail, and we reload
dir.advance(tenantId, r.version(), next);
}
}
private TenantState reserve(TenantRecord r) {
String app = "t_" + r.tenantId().replaceAll("[^a-z0-9]", "_");
dir.reserveAppName(r.tenantId(), app); // unique index; no-op if already ours
return TenantState.RESERVED;
}
private TenantState key(TenantRecord r) {
String alias = "tenant/" + r.tenantId();
kms.findByAlias(alias).orElseGet(() -> kms.create(alias, r.region()));
return TenantState.KEYED;
}
private TenantState configure(TenantRecord r) {
TenantConfig cfg = TenantConfig.fromPlan(r.plan(), r.overrides());
List<String> errors = validator.validate(cfg); // collect-all, not first error
if (!errors.isEmpty()) throw new ProvisioningException("config", errors);
quotas.upsertPlan(r.tenantId(), cfg.quota());
dir.saveConfig(r.tenantId(), cfg);
return TenantState.CONFIGURED;
}
private TenantState wire(TenantRecord r) {
runners.register(r.tenantId(), dir.loadConfig(r.tenantId())); // idempotent put
return TenantState.WIRED;
}
private TenantState verify(TenantRecord r) {
SmokeResult res = smoke.run(r.tenantId());
if (!res.passed()) throw new ProvisioningException("smoke", res.failures());
return TenantState.VERIFIED;
}
}A thrown ProvisioningException is caught by the job that called resume, which records lastError and schedules a retry with backoff. Notice that ACTIVE is not reachable from this loop. Activation is a separate call that checks the state is VERIFIED and records who approved it.
Wiring the tenant's runner
Wiring builds one Runner per tenant with the reserved application name and the tenant's plugins. Runner.builder() accepts the agent, app name, session, artifact and memory services, and plugins. The shared services are stateless clients over durable stores; the app name is what keeps one tenant's rows apart from another's.
Runner buildRunner(TenantRecord r, TenantConfig cfg) {
BaseAgent root = agentFactory.create(cfg); // fresh tree per tenant: no shared mutable tools
return Runner.builder()
.agent(root)
.appName(r.appName()) // the isolation boundary
.sessionService(sessions)
.artifactService(artifacts)
.memoryService(memory)
.plugins(List.of(
new TenantPolicyPlugin(TenantScope.of(r, cfg), usage), // allowlist + budgets
new AdmissionGatePlugin(r.tenantId(), dir))) // refuses unless ACTIVE
.build();
}The admission gate is a beforeRunCallback that reads the tenant state from a short-lived cache and returns a refusal message unless the state is ACTIVE. Keeping the gate inside the runner, not only in the HTTP layer, means a queue consumer or a scheduled job that reaches the runner by another path is refused too. The policy plugin enforcing allowlists and budgets is the one from tenant isolation; onboarding only has to construct it with the tenant's scope.
The smoke run that gates activation
The smoke run is the step that turns configuration into evidence. It runs through the real runner, with admission bypassed by an explicit internal flag, and checks four things:
- One scripted turn completes. Send a fixed prompt that should call one harmless allowlisted tool and produce a final response. Assert the final event exists and that the tool's name appears in the event stream.
- Data lands under the right key. After the turn, list sessions for
(appName, smokeUserId)and assert exactly one exists; list the same user under the default app name and assert none do. - A disallowed tool is refused. A second prompt asks for a tool outside the allowlist; assert the policy plugin's refusal appears instead of a call.
- Usage is recorded. The quota ledger shows tokens charged to this tenant id, which is the same number billing will invoice.
Afterwards, delete the smoke session so it does not appear in the tenant's history or its usage report. Keep the smoke run cheap: one or two model calls on the tenant's configured model. If your model provider is billed per tenant project, this is also the first proof that the tenant's credentials work.
Worked example: one tenant, two retries
Walk one tenant through it. A contract for acme-eu on the Team plan in an EU region creates a record in REQUESTED. The worker reserves t_acme_eu and advances to RESERVED. It creates a key with alias tenant/acme-eu in the EU key ring, then crashes before writing KEYED.
On restart the resume loop reads RESERVED and runs the key step again. The alias lookup finds the key created before the crash and returns it, so there is still exactly one key. Configuration fails validation because the contract enabled a calendar tool that this region's tool registry does not offer; the record shows lastError=config: tool calendar_book not in registry. Support removes the tool from the overrides and reruns. Configuration passes, wiring registers the runner, and the smoke run catches one more problem: the tenant's instruction overlay told the agent to always ask a clarifying question, so the scripted turn never called its tool. The overlay is fixed, the smoke run passes, the state is VERIFIED, and an account manager activates it. Every one of those problems would otherwise have been found by the customer.
Failure modes
- Default application name. A factory falls back to a shared app name when the tenant's is missing. Make the builder throw on a blank or default name, and keep the smoke check that lists the default app.
- Duplicate keys or app names after retries. Steps that create instead of create-or-get. Give every created asset a deterministic name derived from the tenant id and look it up first.
- Lost updates between workers. Two resumes advance the same tenant. The versioned compare-and-set on the state turns this into a harmless reload.
- Runner registry drift. Runners live in memory per replica; a replica that started after onboarding has none. Build runners lazily from the directory on first use, and treat
Wiredas "config stored", not "object exists". - Activation without verification. An admin API sets
ACTIVEdirectly. Enforce the allowed transitions in the directory itself, not in callers. - Smoke data leaking into reports. Use a reserved smoke user id, delete the session, and exclude that user from usage exports.
Trade-offs
A state machine costs more code than a handler, and it is worth it once more than a few tenants a month arrive or more than two systems are involved. Fully automatic activation shortens time-to-value but removes the last human look at a contract's special terms; many teams auto-activate self-serve plans and require approval for enterprise ones. Building runners eagerly at onboarding makes the first request fast but costs memory for tenants that never show up; lazy building, with the per-tenant runner cache bounded as described in tenant scaling, is usually the better default. Finally, a per-tenant data key adds a KMS dependency to every write path, which is a latency and availability cost you accept in exchange for a clean, provable offboarding later.
What to do next
- Write down your tenant states and the single durable fact each one records.
- Make every provisioning step create-or-get with a name derived from the tenant id, and test it by running each step twice.
- Put an admission gate plugin on every tenant runner that refuses anything not
ACTIVE. - Create a per-tenant data key at onboarding, before any tenant data is written.
- Build the four-check smoke run and make it the only path to
VERIFIED. - Route abandoned tenants into the offboarding lifecycle rather than writing a second cleanup path.
- Validate configuration with the collect-all approach in configuration validation at startup, and wire plan limits through quota management.