Most blue/green guides treat the release as a container image. You bring up a second copy, check it, and move traffic. For an agent that answers from a knowledge base or recalls long-term memories, that misses half the release. If the new version embeds queries with a different model, re-chunks documents, or changes what it stores in memory, then the vector index and memory store belong to that version too. Swap the containers and keep one shared index, and one colour searches vectors it cannot interpret. Usually it raises no error. It just retrieves the wrong passages.
This article extends blue/green to an ADK Java agent's data plane: a binding that ties code to its index, a guard against mismatched pairs, building green's index while blue serves, memories that survive rollback, and the cutover order. The container switch and session compatibility are covered in ADK Java blue-green deployment.
Why the container is the wrong unit
An embedding model maps text to a point in a vector space. Two different models produce two different spaces, and distances between a query from one and a document from the other mean nothing, even when the dimensions happen to match. So the query embedder and the index are a pair. They have to change together and be switched together.
Four parts of a typical retrieval agent belong to a release in this sense:
- The knowledge index, built with an embedding model and a chunker. Change either and you need a new index, not an update to the old one.
- The long-term memory store, if memories are retrieved by similarity. Memories written during the window have to be searchable by whichever colour serves next.
- Identifiers in conversation history. Tool results recorded in session events carry document or chunk ids. If green re-chunks, a chunk id in a green-written history may not exist in blue's index.
- Tool servers whose schemas the model sees. Here they matter only for the cutover order.
Parts that change only additively stay shared: new documents in an unchanged embedding space, the session store and the model quota. Colour what changes the meaning of stored data. Share what only adds to it.
The release binding
Make the pairing explicit with one immutable record per colour, loaded at startup and never edited after the colour is built. The router or traffic split chooses a colour. Each process serves exactly one binding, so no request can mix blue's embedder with green's index.
/** Everything that must change together. One per colour, immutable once built. */
public record ReleaseBinding(
String colour, // "blue" or "green"
String release, // "41", "42"
String modelVersion, // a pinned model version, never a floating alias
String instructionSha, // hash of the instruction text in the config repo
String embeddingModel, // embeds queries for this colour; documents were embedded with it too
int embeddingDim,
String knowledgeIndex, // physical index name, e.g. "kb_v8", never "kb_current"
String memoryNamespace, // e.g. "mem_B"
URI toolServer) {} // tool backend this colour was verified againstName the physical index. An alias such as kb_current that a re-index job can swap moves the data without moving the code, which is exactly the mismatch this design prevents. The release pointer should be the only switch.
A startup guard against mismatched pairs
A mismatched pair passes every HTTP health check, so the agent must check the pairing itself before it reports ready. Store the embedding model name and dimension as metadata on the index when you create it. Most vector stores allow index-level labels; if yours does not, keep a small manifest table. Then compare them with the binding at startup. Below, Embedder, VectorIndex, IndexInfo and Hit are your own thin interfaces over whatever store you use. The ADK parts are FunctionTool and the @Schema annotation from com.google.adk.tools.Annotations.
public final class KnowledgeTools {
private final ReleaseBinding b;
private final Embedder embedder;
private final VectorIndex index;
public KnowledgeTools(ReleaseBinding b, Embedder embedder, VectorIndex index) {
this.b = b; this.embedder = embedder; this.index = index;
}
/** Runs before readiness. A mismatch is a startup failure, never a warning. */
public void verifyPairing() {
IndexInfo info = index.describe(b.knowledgeIndex());
if (!info.embeddingModel().equals(b.embeddingModel()) || info.dimension() != b.embeddingDim())
throw new IllegalStateException(b.knowledgeIndex() + " holds " + info.embeddingModel()
+ "/" + info.dimension() + " vectors; release " + b.release() + " embeds with "
+ b.embeddingModel() + "/" + b.embeddingDim());
if (embedder.embed("pairing probe").length != b.embeddingDim())
throw new IllegalStateException("embedder dimension differs from binding");
}
public Map<String, Object> searchKnowledge(
@Schema(name = "query", description = "What to look up in the product knowledge base")
String query) {
List<Hit> hits = index.search(b.knowledgeIndex(), embedder.embed(query), 8);
return Map.of("results", hits.stream()
.map(h -> Map.of("doc_id", h.docId(), "title", h.title(), "text", h.text()))
.toList());
}
}
// Startup: one colour per process.
ReleaseBinding b = BindingLoader.load(System.getenv("AGENT_COLOUR"));
KnowledgeTools kt = new KnowledgeTools(b, embedders.forModel(b.embeddingModel()), vectorIndex);
kt.verifyPairing();
LlmAgent agent = LlmAgent.builder()
.name("support_agent")
.model(b.modelVersion())
.instruction(instructions.load(b.instructionSha()))
.tools(FunctionTool.create(kt, "searchKnowledge"))
.build();Compare model names, not only dimensions. Two models with the same output size fail silently: every query returns confident, irrelevant passages, and the model answers from them.
Building the green index while blue serves
Green's index has to be complete and current before any user reaches it. Blue keeps serving the whole time. Drive both indexes from one ordered change log, such as an outbox table, a Pub/Sub topic with ordering keys or a Kafka topic, and run one indexer per colour with its own checkpoint. A fresh green indexer starts from position zero, which is the backfill, and then keeps following the log, which keeps it live. If your log does not retain history, snapshot the document store at log position P, bulk-index the snapshot, and start the green indexer at P. The same pattern, applied to search clusters, appears in designing search at scale.
/** One instance per colour; both read the same ordered log. Upserts replace all chunks of a doc. */
void runIndexer(ChangeLog log, String indexName, Embedder emb, Chunker chunker) {
long pos = checkpoints.get(indexName); // 0 for a fresh index = full backfill
while (running) {
List<Change> batch = log.read(pos, 500);
for (Change c : batch) {
if (c.deleted()) index.deleteDoc(indexName, c.docId());
else index.upsertDoc(indexName, c.docId(), embedAll(emb, chunker.split(c.doc())));
}
if (!batch.isEmpty()) {
pos = batch.get(batch.size() - 1).seq();
checkpoints.put(indexName, pos); // after the writes: replays are idempotent
}
lagGauge.set(indexName, log.head() - pos);
}
}Upserts replace every chunk of a document and the checkpoint is written after the index writes, so a crash replays work harmlessly rather than skipping it. Before green is eligible for traffic, gate on index facts as well as agent behaviour: document count equal to blue's within the deletes still in flight, indexer lag under a minute, and recall at 10 on a labelled query set at least equal to blue's. Then run the agent-level golden replays from the companion article against green's private address.
Memories written during the window
Long-term memories are the hard case because users create them during the window. A user tells green their preferred contact hours, green stores a memory, and next morning you roll back. If that memory exists only as a vector in space B, blue cannot find it and the user has to repeat themselves.
The fix is to separate the memory from its embedding. Store the memory text once, keyed by a memory id, in a store both colours read. Then, for the length of the window, enqueue an embedding job for every active namespace: mem_A and mem_B. Each colour searches its own namespace and resolves ids to the shared text. When the window closes, stop embedding into the losing space and drop it. If green loses, the same step removes mem_B. Dual embedding roughly doubles embedding spend on the memory write path for the length of the window. That is usually small next to the backfill.
Stable ids so history survives a switch
Session history outlives the switch. A tool result recorded on Monday may be read by the other colour on Tuesday. If searchKnowledge returns chunk ids and a follow-up tool such as openDocument(doc_id) fetches by id, then chunk ids from green's chunker will not resolve in blue's index. Return stable document ids, which come from the source system and are independent of chunking. Resolve documents against the document store, not the vector index. If you need finer references, derive them from the document plus a character offset, not from a chunk number that changes whenever the chunker does.
The cutover order
The cutover is a sequence. Each step can be reversed until the last one.
- Build. Create
kb_v8with its metadata, start indexer B and let it backfill. Start dual embedding of new memories intomem_B. Blue is unaffected. - Deploy green dark. Deploy release 42 with its binding. Readiness depends on
verifyPairing(). If green's tool server version differs, deploy it first and give it its own address in the binding. - Verify. Check index gates (count, lag, recall), then agent golden replays, then continuation turns on copied real sessions.
- Flip the pointer. Users move to green. Indexer A and dual embedding keep running, so blue stays a current, complete rollback target.
- Hold the window. Keep it open for as long as you need to catch the regressions that matter, typically one to three business days for answer quality.
- Contract. Stop indexer A, stop embedding into
mem_A, deletekb_v7and scale blue to zero. Only after this step is rollback no longer instant.
The common mistake is to stop indexer A at step 4 to save money. Rollback still flips traffic in seconds, but blue then answers from a stale knowledge base.
Worked example: a new embedding model and chunker
Take a support agent over 1.2 million chunks that moves from embedding model A (768 dimensions) to model B (1,024 dimensions), with a new chunker that keeps tables whole. The figures below are illustrative. The arithmetic is what carries over to your system.
| Quantity | Calculation | Result |
|---|---|---|
| Backfill tokens | 1.2M chunks x 350 tokens average | 420M tokens |
| Backfill embedding cost | 420M x a placeholder 0.02 to 0.15 USD per 1M | 8.40 to 63 USD |
| Backfill time | 1.2M chunks at 400 chunks/s (quota-limited) | 3,000 s, about 50 min |
| Raw vector storage, blue | 1.2M x 768 x 4 bytes | about 3.7 GB |
| Raw vector storage, green | 1.2M x 1,024 x 4 bytes | about 4.9 GB |
| Both during the window | before index overhead | about 8.6 GB |
Verification used 600 labelled queries. Recall at 10 was 0.78 on kb_v7 and 0.84 on kb_v8, and 114 of 120 golden conversations passed on green against 112 on blue, so the pointer flipped on Tuesday morning. On Wednesday, support flagged answers about plan limits that quoted the wrong tier. The new chunker kept a pricing table whole, and the resulting chunk was too long to rank well for narrow questions. Rollback was one pointer change. Because indexer A had kept following the log, blue answered from documents updated on Tuesday. Because memories had been dual-embedded, the 3,400 memories users created on green were searchable on blue. Release 43 shipped a table-splitting rule, built kb_v9 from position zero, and went through the same sequence.
Failure modes
- Silent space mismatch. A shared alias, or two models with equal dimensions, lets a colour search the wrong space with no error. Bind physical index names, and check the model name at startup.
- Stale green. The backfill finished but the indexer then stalled. Gate traffic on the lag gauge, and alert on it for both colours during the window.
- Stale blue. Indexer A was stopped at the flip, so rollback serves yesterday's documents. Stop it only at the contract step.
- One-sided memories. Memories embedded only in the active space are lost on rollback. Store text once and embed it into every live namespace.
- Unresolvable ids in history. Chunk ids from one chunker appear in sessions read by the other. Return source document ids.
- Quota contention. The backfill shares the embedding quota with live query embedding and pushes up blue's latency. Throttle the backfill, or run it under a separate project or key.
Trade-offs
| Approach | Rollback | Extra cost | Use when |
|---|---|---|---|
| Colour the data plane (this page) | seconds, data included | second index, dual embedding for the window | the embedding model or chunker changes |
| Code-only blue/green, shared index | seconds, code only | none | the embedding space does not change |
| Re-index in place, one colour | none for the data | none | additive document changes only |
| Canary with a per-variant index | gradual | second index, longer | you need a statistical quality comparison |
Data-plane blue/green costs storage and some embedding spend, and it buys an honest rollback. When the embedding space is unchanged, the extra index is pure overhead, so fall back to code-only blue/green. To compare answer quality on live traffic before committing, use the session-pinned canary approach, with each variant bound to its own index in the same way.
What to do next
- List every store your agent searches by similarity, and record which embedding model and chunker produced each one.
- Add the embedding model and dimension as index metadata, and add a startup pairing check that fails readiness on mismatch.
- Replace any floating index alias in agent configuration with a physical index name in a per-release binding.
- Drive indexing from one ordered change log with a per-index checkpoint and a lag gauge.
- Split long-term memory into shared text plus per-namespace embeddings, and dual-embed during the window.
- Write the six-step cutover into your runbook, with the contract step as a separate, approved change.
- Wire green verification into the deploy pipeline, and review long-term memory retrieval for the memory read path.