Most blue/green guides treat the release as a container image. You bring up a second copy, check it, and move traffic. For an agent that answers from a knowledge base or recalls long-term memories, that misses half the release. If the new version embeds queries with a different model, re-chunks documents, or changes what it stores in memory, then the vector index and memory store belong to that version too. Swap the containers and keep one shared index, and one colour searches vectors it cannot interpret. Usually it raises no error. It just retrieves the wrong passages.

This article extends blue/green to an ADK Java agent's data plane: a binding that ties code to its index, a guard against mismatched pairs, building green's index while blue serves, memories that survive rollback, and the cutover order. The container switch and session compatibility are covered in ADK Java blue-green deployment.

Why the container is the wrong unit

An embedding model maps text to a point in a vector space. Two different models produce two different spaces, and distances between a query from one and a document from the other mean nothing, even when the dimensions happen to match. So the query embedder and the index are a pair. They have to change together and be switched together.

Four parts of a typical retrieval agent belong to a release in this sense:

  • The knowledge index, built with an embedding model and a chunker. Change either and you need a new index, not an update to the old one.
  • The long-term memory store, if memories are retrieved by similarity. Memories written during the window have to be searchable by whichever colour serves next.
  • Identifiers in conversation history. Tool results recorded in session events carry document or chunk ids. If green re-chunks, a chunk id in a green-written history may not exist in blue's index.
  • Tool servers whose schemas the model sees. Here they matter only for the cutover order.

Parts that change only additively stay shared: new documents in an unchanged embedding space, the session store and the model quota. Colour what changes the meaning of stored data. Share what only adds to it.

The release binding

Blue/green over the agent's data plane, not only its containerRelease pointeractive = blueBlue: agent release 41model M1, query embedder A (768)Green: agent release 42model M1, query embedder B (1024)all usersverifier onlykb_v7vectors in space Amem_Amemories, space Akb_v8vectors in space Bmem_Bmemories, space BDocument change logone ordered stream of upserts and deletesindexer A (live)indexer B (backfill, then live)Memory write pathembeds into both spacesEach colour reads only the index and memory namespace built in its own embedding space.
Each colour pairs its query embedder with an index and memory namespace in the same space. One change log feeds both indexers, and the memory write path embeds into both spaces while the window is open.

Make the pairing explicit with one immutable record per colour, loaded at startup and never edited after the colour is built. The router or traffic split chooses a colour. Each process serves exactly one binding, so no request can mix blue's embedder with green's index.

/** Everything that must change together. One per colour, immutable once built. */
public record ReleaseBinding(
    String colour,           // "blue" or "green"
    String release,          // "41", "42"
    String modelVersion,     // a pinned model version, never a floating alias
    String instructionSha,   // hash of the instruction text in the config repo
    String embeddingModel,   // embeds queries for this colour; documents were embedded with it too
    int embeddingDim,
    String knowledgeIndex,   // physical index name, e.g. "kb_v8", never "kb_current"
    String memoryNamespace,  // e.g. "mem_B"
    URI toolServer) {}       // tool backend this colour was verified against

Name the physical index. An alias such as kb_current that a re-index job can swap moves the data without moving the code, which is exactly the mismatch this design prevents. The release pointer should be the only switch.

A startup guard against mismatched pairs

A mismatched pair passes every HTTP health check, so the agent must check the pairing itself before it reports ready. Store the embedding model name and dimension as metadata on the index when you create it. Most vector stores allow index-level labels; if yours does not, keep a small manifest table. Then compare them with the binding at startup. Below, Embedder, VectorIndex, IndexInfo and Hit are your own thin interfaces over whatever store you use. The ADK parts are FunctionTool and the @Schema annotation from com.google.adk.tools.Annotations.

public final class KnowledgeTools {
  private final ReleaseBinding b;
  private final Embedder embedder;
  private final VectorIndex index;

  public KnowledgeTools(ReleaseBinding b, Embedder embedder, VectorIndex index) {
    this.b = b; this.embedder = embedder; this.index = index;
  }

  /** Runs before readiness. A mismatch is a startup failure, never a warning. */
  public void verifyPairing() {
    IndexInfo info = index.describe(b.knowledgeIndex());
    if (!info.embeddingModel().equals(b.embeddingModel()) || info.dimension() != b.embeddingDim())
      throw new IllegalStateException(b.knowledgeIndex() + " holds " + info.embeddingModel()
          + "/" + info.dimension() + " vectors; release " + b.release() + " embeds with "
          + b.embeddingModel() + "/" + b.embeddingDim());
    if (embedder.embed("pairing probe").length != b.embeddingDim())
      throw new IllegalStateException("embedder dimension differs from binding");
  }

  public Map<String, Object> searchKnowledge(
      @Schema(name = "query", description = "What to look up in the product knowledge base")
      String query) {
    List<Hit> hits = index.search(b.knowledgeIndex(), embedder.embed(query), 8);
    return Map.of("results", hits.stream()
        .map(h -> Map.of("doc_id", h.docId(), "title", h.title(), "text", h.text()))
        .toList());
  }
}

// Startup: one colour per process.
ReleaseBinding b = BindingLoader.load(System.getenv("AGENT_COLOUR"));
KnowledgeTools kt = new KnowledgeTools(b, embedders.forModel(b.embeddingModel()), vectorIndex);
kt.verifyPairing();
LlmAgent agent = LlmAgent.builder()
    .name("support_agent")
    .model(b.modelVersion())
    .instruction(instructions.load(b.instructionSha()))
    .tools(FunctionTool.create(kt, "searchKnowledge"))
    .build();

Compare model names, not only dimensions. Two models with the same output size fail silently: every query returns confident, irrelevant passages, and the model answers from them.

Building the green index while blue serves

Green's index has to be complete and current before any user reaches it. Blue keeps serving the whole time. Drive both indexes from one ordered change log, such as an outbox table, a Pub/Sub topic with ordering keys or a Kafka topic, and run one indexer per colour with its own checkpoint. A fresh green indexer starts from position zero, which is the backfill, and then keeps following the log, which keeps it live. If your log does not retain history, snapshot the document store at log position P, bulk-index the snapshot, and start the green indexer at P. The same pattern, applied to search clusters, appears in designing search at scale.

/** One instance per colour; both read the same ordered log. Upserts replace all chunks of a doc. */
void runIndexer(ChangeLog log, String indexName, Embedder emb, Chunker chunker) {
  long pos = checkpoints.get(indexName);              // 0 for a fresh index = full backfill
  while (running) {
    List<Change> batch = log.read(pos, 500);
    for (Change c : batch) {
      if (c.deleted()) index.deleteDoc(indexName, c.docId());
      else index.upsertDoc(indexName, c.docId(), embedAll(emb, chunker.split(c.doc())));
    }
    if (!batch.isEmpty()) {
      pos = batch.get(batch.size() - 1).seq();
      checkpoints.put(indexName, pos);                // after the writes: replays are idempotent
    }
    lagGauge.set(indexName, log.head() - pos);
  }
}

Upserts replace every chunk of a document and the checkpoint is written after the index writes, so a crash replays work harmlessly rather than skipping it. Before green is eligible for traffic, gate on index facts as well as agent behaviour: document count equal to blue's within the deletes still in flight, indexer lag under a minute, and recall at 10 on a labelled query set at least equal to blue's. Then run the agent-level golden replays from the companion article against green's private address.

Memories written during the window

Long-term memories are the hard case because users create them during the window. A user tells green their preferred contact hours, green stores a memory, and next morning you roll back. If that memory exists only as a vector in space B, blue cannot find it and the user has to repeat themselves.

The fix is to separate the memory from its embedding. Store the memory text once, keyed by a memory id, in a store both colours read. Then, for the length of the window, enqueue an embedding job for every active namespace: mem_A and mem_B. Each colour searches its own namespace and resolves ids to the shared text. When the window closes, stop embedding into the losing space and drop it. If green loses, the same step removes mem_B. Dual embedding roughly doubles embedding spend on the memory write path for the length of the window. That is usually small next to the backfill.

Stable ids so history survives a switch

Session history outlives the switch. A tool result recorded on Monday may be read by the other colour on Tuesday. If searchKnowledge returns chunk ids and a follow-up tool such as openDocument(doc_id) fetches by id, then chunk ids from green's chunker will not resolve in blue's index. Return stable document ids, which come from the source system and are independent of chunking. Resolve documents against the document store, not the vector index. If you need finer references, derive them from the document plus a character offset, not from a chunk number that changes whenever the chunker does.

The cutover order

The cutover is a sequence. Each step can be reversed until the last one.

  1. Build. Create kb_v8 with its metadata, start indexer B and let it backfill. Start dual embedding of new memories into mem_B. Blue is unaffected.
  2. Deploy green dark. Deploy release 42 with its binding. Readiness depends on verifyPairing(). If green's tool server version differs, deploy it first and give it its own address in the binding.
  3. Verify. Check index gates (count, lag, recall), then agent golden replays, then continuation turns on copied real sessions.
  4. Flip the pointer. Users move to green. Indexer A and dual embedding keep running, so blue stays a current, complete rollback target.
  5. Hold the window. Keep it open for as long as you need to catch the regressions that matter, typically one to three business days for answer quality.
  6. Contract. Stop indexer A, stop embedding into mem_A, delete kb_v7 and scale blue to zero. Only after this step is rollback no longer instant.

The common mistake is to stop indexer A at step 4 to save money. Rollback still flips traffic in seconds, but blue then answers from a stale knowledge base.

Worked example: a new embedding model and chunker

Take a support agent over 1.2 million chunks that moves from embedding model A (768 dimensions) to model B (1,024 dimensions), with a new chunker that keeps tables whole. The figures below are illustrative. The arithmetic is what carries over to your system.

QuantityCalculationResult
Backfill tokens1.2M chunks x 350 tokens average420M tokens
Backfill embedding cost420M x a placeholder 0.02 to 0.15 USD per 1M8.40 to 63 USD
Backfill time1.2M chunks at 400 chunks/s (quota-limited)3,000 s, about 50 min
Raw vector storage, blue1.2M x 768 x 4 bytesabout 3.7 GB
Raw vector storage, green1.2M x 1,024 x 4 bytesabout 4.9 GB
Both during the windowbefore index overheadabout 8.6 GB

Verification used 600 labelled queries. Recall at 10 was 0.78 on kb_v7 and 0.84 on kb_v8, and 114 of 120 golden conversations passed on green against 112 on blue, so the pointer flipped on Tuesday morning. On Wednesday, support flagged answers about plan limits that quoted the wrong tier. The new chunker kept a pricing table whole, and the resulting chunk was too long to rank well for narrow questions. Rollback was one pointer change. Because indexer A had kept following the log, blue answered from documents updated on Tuesday. Because memories had been dual-embedded, the 3,400 memories users created on green were searchable on blue. Release 43 shipped a table-splitting rule, built kb_v9 from position zero, and went through the same sequence.

Failure modes

  • Silent space mismatch. A shared alias, or two models with equal dimensions, lets a colour search the wrong space with no error. Bind physical index names, and check the model name at startup.
  • Stale green. The backfill finished but the indexer then stalled. Gate traffic on the lag gauge, and alert on it for both colours during the window.
  • Stale blue. Indexer A was stopped at the flip, so rollback serves yesterday's documents. Stop it only at the contract step.
  • One-sided memories. Memories embedded only in the active space are lost on rollback. Store text once and embed it into every live namespace.
  • Unresolvable ids in history. Chunk ids from one chunker appear in sessions read by the other. Return source document ids.
  • Quota contention. The backfill shares the embedding quota with live query embedding and pushes up blue's latency. Throttle the backfill, or run it under a separate project or key.

Trade-offs

ApproachRollbackExtra costUse when
Colour the data plane (this page)seconds, data includedsecond index, dual embedding for the windowthe embedding model or chunker changes
Code-only blue/green, shared indexseconds, code onlynonethe embedding space does not change
Re-index in place, one colournone for the datanoneadditive document changes only
Canary with a per-variant indexgradualsecond index, longeryou need a statistical quality comparison

Data-plane blue/green costs storage and some embedding spend, and it buys an honest rollback. When the embedding space is unchanged, the extra index is pure overhead, so fall back to code-only blue/green. To compare answer quality on live traffic before committing, use the session-pinned canary approach, with each variant bound to its own index in the same way.

What to do next

  1. List every store your agent searches by similarity, and record which embedding model and chunker produced each one.
  2. Add the embedding model and dimension as index metadata, and add a startup pairing check that fails readiness on mismatch.
  3. Replace any floating index alias in agent configuration with a physical index name in a per-release binding.
  4. Drive indexing from one ordered change log with a per-index checkpoint and a lag gauge.
  5. Split long-term memory into shared text plus per-namespace embeddings, and dual-embed during the window.
  6. Write the six-step cutover into your runbook, with the contract step as a separate, approved change.
  7. Wire green verification into the deploy pipeline, and review long-term memory retrieval for the memory read path.
Key takeaway: When an agent retrieves by similarity, its query embedder and its index are one unit, and blue/green has to switch them together. Bind each colour to physical index and memory names, refuse to start on a mismatched pair, build green from the same change log blue uses, and dual-embed memories. Keep blue's data current until the window closes. Rollback is then one pointer change, and neither stale documents nor lost memories come with it.