Sooner or later a tenant has to move. A customer upgrades from the pooled tier to a dedicated store, a data-residency contract requires another region, or you replace the database behind the session service. The agent must keep answering, and afterwards every conversation, state value, file version and memory must be where the tenant expects it, with nothing left behind and nothing copied twice.

ADK Java has no migration feature, so you build one on its storage interfaces. This article starts with what those interfaces let you read and write, because several of them behave in ways that silently corrupt a naive copy. It then builds a migration in four parts: a router with a drain fence, a replay copier, a fingerprint verifier and a rollback plan. The API details were read from the google/adk-java source on its main branch; re-check them against the version you run.

What a tenant migration moves

The design assumes the layout from the tenant data separation article: each tenant has its own app name, such as support.acme, and a runner per app name. A migration then changes where a tenant's bytes live, not who the tenant is. The app name, user ids and session ids stay the same, so links, client-side session ids and audit records survive the move.

DataRead withWrite withTrap
Sessions and eventsgetSession per idcreateSession, appendEventlistSessions in the in-memory service returns copies with empty events and state
Session, user: and app: statemerged into getSession's statestateDelta on replayed eventscreateSession stores its state map as session-local
ArtifactslistArtifactKeys, listVersions, loadArtifactsaveArtifactthe target assigns version numbers
Memoryno export APIaddSessionToMemoryBaseMemoryService has only add and search
Your indexes, caches, logsyour own toolsyour own toolsADK knows nothing about them

What the interfaces let you copy

Four details from the source decide how the copy works.

Listing is not reading. listSessions(appName, userId) gives you ids. The in-memory implementation documents that it returns copies with empty events and state. Copy from that list and every target session is empty. Use the list only for ids, then call getSession(appName, userId, sessionId, Optional.empty()) for each.

State lives in three scopes. getSession returns session state merged with the user: and app: scopes. appendEvent routes each key of an event's stateDelta to the right scope; the base implementation also skips temp: keys. But createSession stores whatever map you pass as the new session's own state. Seed a target session with the merged map and you copy every app: and user: value into every session, where later updates will not reach it. Seed only unprefixed keys and let replayed events rebuild the shared scopes.

Reused ids may overwrite. InMemorySessionService's no-argument constructor uses the deprecated OVERWRITE behaviour: creating a session with an id that exists replaces it, discarding its events. REJECT is opt-in. Persistent implementations differ. The copier therefore deletes any target session explicitly before recreating it, so a re-run behaves the same everywhere.

Artifact versions are renumbered. saveArtifact returns a Single<Integer> with the version the target assigned. Copy versions in ascending order and record source-to-target pairs; anything that stored a version number, such as a tool's state value, must be checked against that map.

The architecture: phases and a fence

Tenant migration: route, fence, copy, verify, then fliprequeststenant from tokenMigrationRouterphase + in-flight countsource runnerpooled storetarget runnersilo storeSOURCETARGETdirty setsessions touchedCopiergetSession, replay eventsVerifierfingerprints both sidesversion mapartifact renumberingPREPAREBACKFILLDRAINDELTAVERIFYTARGETSOAKwrites are refused only during DRAIN, DELTA and VERIFY; the source stays read-only through SOAK
The router decides which runner serves a tenant and records dirty sessions. The copier and verifier work from the source to the target; the phases at the bottom are per tenant.

The migration runs per tenant through seven phases. PREPARE provisions the target store and a target runner wrapped in the same guards as the source. BACKFILL copies every session while the tenant stays live; the router records each session id it serves into a dirty set. DRAIN closes the fence: new invocations get a retry-after response, and the router waits for in-flight invocations to finish. DELTA re-copies only dirty sessions and then reconciles user: and app: state. VERIFY compares fingerprints for every session. TARGET flips the router, which opens the fence on the new store. SOAK keeps the source read-only for a fixed period as the rollback path, after which offboarding deletes it.

Why a fence rather than dual writes? An ADK invocation appends many events and state deltas, and tool calls can take seconds. Writing every event to two stores, in order, with failure handling on each, doubles the surface of the hottest path. A fence makes the tenant unavailable for a short, measured interval and keeps the serving code unchanged.

The router and the drain fence

The fence must not let an invocation slip in after DRAIN begins. The usual bug is to read the phase, then increment the in-flight counter: a request that reads SOURCE just before the flip increments after the drain loop saw zero. Increment first, then re-check:

public final class MigrationRouter {
  public enum Phase { SOURCE, DRAINING, TARGET }

  private final ConcurrentHashMap<String, Phase> phase = new ConcurrentHashMap<>();
  private final ConcurrentHashMap<String, AtomicInteger> inFlight = new ConcurrentHashMap<>();
  private final Set<String> dirty = ConcurrentHashMap.newKeySet();   // "tenant/user/session"
  private final TenantRunners source, target;    // per-tenant runner caches

  public Flowable<Event> run(String tenant, String userId, String sessionId, Content msg) {
    AtomicInteger n = inFlight.computeIfAbsent(tenant, k -> new AtomicInteger());
    n.incrementAndGet();                          // announce first ...
    Phase ph = phase.getOrDefault(tenant, Phase.SOURCE);
    if (ph == Phase.DRAINING) {                   // ... then look
      n.decrementAndGet();
      return Flowable.error(new RetryLater(tenant, Duration.ofSeconds(30)));
    }
    if (ph == Phase.SOURCE) dirty.add(tenant + "/" + userId + "/" + sessionId);
    Runner r = (ph == Phase.TARGET ? target : source).forTenant(tenant);
    return r.runAsync(userId, sessionId, msg, RunConfig.builder().build())
        .doFinally(n::decrementAndGet);
  }

  public void startDrain(String tenant) { phase.put(tenant, Phase.DRAINING); }
  public boolean drained(String tenant) {
    return inFlight.getOrDefault(tenant, new AtomicInteger()).get() == 0;
  }
  public void flip(String tenant) { phase.put(tenant, Phase.TARGET); }
}

Bound the drain: if in-flight count is not zero after your longest tool timeout plus a margin, abort and reopen on the source rather than wait. Clients must treat RetryLater, our own exception type, as retryable. If you run several instances, the phase and counters belong in a shared store, and the drain waits for every instance to acknowledge the new phase.

The copier

The copier works on one session at a time. It deletes any previous target copy, creates the session with only session-scoped state, replays complete events in order, then copies artifact versions in ascending order and records the version map.

Single<Long> copySession(String app, String user, String sid) {
  return src.getSession(app, user, sid, Optional.empty())
      .switchIfEmpty(Single.error(new IllegalStateException("vanished: " + sid)))
      .flatMap(s -> dst.deleteSession(app, user, sid).onErrorComplete()
          .andThen(dst.createSession(app, user, sessionScoped(s.state()), sid))
          .flatMap(t -> Flowable.fromIterable(s.events())
              .filter(e -> !e.partial().orElse(false))
              .concatMapSingle(e -> dst.appendEvent(t, e))
              .count()))
      .flatMap(events -> copyArtifacts(app, user, sid).map(v -> events + v));
}

static ConcurrentMap<String, Object> sessionScoped(Map<String, Object> merged) {
  ConcurrentMap<String, Object> out = new ConcurrentHashMap<>();
  merged.forEach((k, v) -> {
    if (!shared(k) && !k.startsWith(State.TEMP_PREFIX)) out.put(k, v);
  });
  return out;
}

static boolean shared(String k) {
  return k.startsWith(State.APP_PREFIX) || k.startsWith(State.USER_PREFIX);
}

Single<Long> copyArtifacts(String app, String user, String sid) {
  return srcArt.listArtifactKeys(app, user, sid)
      .flattenAsFlowable(ListArtifactsResponse::filenames)
      .filter(name -> !name.startsWith("user:") || copiedUserFiles.add(app + "/" + user + "/" + name))
      .concatMapSingle(name -> srcArt.listVersions(app, user, sid, name)
          .flattenAsFlowable(vs -> vs.stream().sorted().toList())
          .concatMapSingle(v -> srcArt.loadArtifact(app, user, sid, name, v).toSingle()
              .flatMap(part -> dstArt.saveArtifact(app, user, sid, name, part))
              .doOnSuccess(nv -> versionMap.put(app, user, sid, name, v, nv)))
          .count())
      .reduce(0L, Long::sum);
}

The filter on user: filenames copies user-scoped files once per user, whichever session lists them first; check how your artifact service lists them.

Replay is right for session-scoped keys and wrong for shared ones. If an old session sets user:tier to basic and a newer one sets it to pro, copying the newer session first leaves basic on the target, and every DELTA re-copy of an old session re-applies old values over new ones. Re-copying never converges. So after DELTA, while the fence is still closed, write the shared scopes once from the source, per user, in one migration-authored event:

Completable reconcileShared(String app, String user, String anySid) {
  return Single.zip(
          src.getSession(app, user, anySid, Optional.empty()).toSingle(),
          dst.getSession(app, user, anySid, Optional.empty()).toSingle(),
          (s, t) -> {
            Map<String, Object> delta = new HashMap<>();
            s.state().forEach((k, v) -> { if (shared(k)) delta.put(k, v); });
            t.state().keySet().stream()
                .filter(k -> shared(k) && !s.state().containsKey(k))
                .forEach(k -> delta.put(k, State.REMOVED));
            Event fix = Event.builder().id(Event.generateEventId())
                .invocationId("tenant-migration").author("migration")
                .actions(EventActions.builder().stateDelta(delta).build()).build();
            return dst.appendEvent(t, fix);
          })
      .flatMap(x -> x)
      .ignoreElement();
}

Both getSession calls return state with the app: and user: scopes merged in under their prefixes, which is what makes the diff possible. The fix event becomes part of one session's history on the target, so the verifier skips events authored by migration.

Memory has no export API. After the delta pass, call addSessionToMemory on the target memory service for each migrated session, or move the underlying index with its own tools, keeping the same embedding model; see the memory service article. Your own vector indexes, caches and logs need the same treatment.

Verification by fingerprint

Verification compares a fingerprint of each session on both sides. It must include everything a user could see and exclude what the target is allowed to change. Whether a persistent target preserves event ids and timestamps is implementation-specific, so leave them out unless you have read your service's code.

String fingerprint(Session s, List<String> artifactDigests) {
  StringBuilder b = new StringBuilder();
  for (Event e : s.events()) {
    if (e.partial().orElse(false) || "migration".equals(e.author())) continue;
    Map<String, Object> delta = e.actions() == null || e.actions().stateDelta() == null
        ? Map.of() : e.actions().stateDelta();
    b.append(e.author()).append('|')
     .append(JsonBaseModel.toJsonString(e.content().orElse(null))).append('|')
     .append(JsonBaseModel.toJsonString(new TreeMap<>(delta))).append('\n');
  }
  b.append(JsonBaseModel.toJsonString(new TreeMap<>(s.state())));
  artifactDigests.stream().sorted().forEach(b::append);   // name, version count, bytes hash
  return sha256Hex(b.toString());
}

Sorting map keys makes the digest independent of hash-map order; nested maps inside state need the same canonical treatment. A mismatch is a blocker: re-copy that session, re-run the reconciliation for its user, or abort the cutover. Also compare the session count and the per-user artifact counts, which catch sessions missing on one side entirely.

Worked example: sizing the fence

Plan a move for a tenant with 12,600 sessions across 1,840 users. A test copy of 500 sessions runs at 40 sessions per second against the target, so BACKFILL takes about 12,600 / 40 = 315 seconds of copy time; spread over an evening it stays below the store's write budget. Suppose 900 sessions are touched during BACKFILL. The fence then costs the drain, bounded by the 60-second tool timeout, plus 900 / 40 = 22.5 seconds of delta copy, plus verifying 900 sessions on both sides at a similar rate: roughly two minutes of retry-after responses in total. If that is too long, run a second BACKFILL pass over the dirty set before draining, which shrinks the final delta, or schedule the fence in the tenant's quietest hour.

Failure modes

  • Empty sessions on the target. The copier read from listSessions. Every count matches and every conversation is gone; fingerprints catch it.
  • Shared state frozen into sessions. The target was seeded with merged state, so app: and user: values exist per session and diverge from the real scopes.
  • Old shared values win. Replay order put an old user: assignment last. Only the reconciliation event fixes it; re-copying does not.
  • Silent overwrite on re-run. A second copy with the same id replaced a session that had already received new traffic after cutover. Never re-copy after TARGET.
  • Stale artifact versions. State or messages refer to version 7, but the target renumbered from 1. Consult the version map or keep version references out of state.
  • Fence leak. Phase read before the counter increment lets one invocation write to the source after the delta copy. Its events are lost at cutover.
  • Forgotten side channels. Memory, vector indexes and log routing still point at the pooled store, so search returns old data or nothing.

Trade-offs and rollback

ApproachGainsCosts
Fence plus delta copy (this article)Serving path unchanged; works across store enginesShort unavailability per tenant
Dual writes during migrationNo fenceTwo writes per event on the hot path; ordering and partial failure
Database-level replicationFastest for large tenantsSame engine on both sides; ADK-level checks still needed
Rename on move (new app name)Clean separation of old and newEvery stored reference to the app name must change

Rollback is simple only while the source is untouched: during SOAK, keep it read-only, and if the target misbehaves, flip the router back and replay the sessions created since cutover in the opposite direction with the same copier. The resume after crash article covers what an interrupted invocation looks like in the event log, which matters when you decide what a drained session contains, and tenant isolation covers the runtime boundary the target runner must keep.

What to do next

  1. Read your session and artifact services' list, create and save methods, and write down how each handles listing, duplicate ids and versions.
  2. Implement the router with increment-then-check and a bounded drain, and test it with a request that arrives between phase change and drain check.
  3. Build the copier and verifier, then migrate a synthetic tenant with app:, user: and session state, user-scoped files and multi-version artifacts. Require identical fingerprints.
  4. Time a 500-session copy against the real target and compute your fence length from the dirty-set size, as in the worked example.
  5. List every side channel keyed by tenant, including memory, vector indexes, caches and logs, and give each a copy and a verify step.
  6. Write the rollback runbook before the first production move, and rehearse it once.
Key takeaway: Treat a tenant migration as copy, fence, delta, reconcile, verify and flip. Read sessions with getSession, not listSessions; seed only session-scoped state and let replayed events rebuild user: and app: state; delete before re-creating; map renumbered artifact versions; rebuild memory explicitly. Prove the move with fingerprints, and keep the source read-only until you no longer need to roll back.