Most teams have three environments for an ADK Java agent and assume they mean something: dev for building, staging for proving, production for users. In practice the three drift apart quietly. Staging points at a different model version because someone was testing a preview. Production's MCP server was upgraded on its own schedule and now exposes two extra tools. Staging's payment tool talks to a sandbox that always succeeds, so the refusal path was never exercised. When the release is promoted, the evidence collected in staging describes an agent that does not exist in production.
This article is about the environments, not the pipeline. Building once, release manifests and canaries are covered in the CI/CD pipeline article, and ramps and rollback in the deploy pipeline article. Here you will define an environment contract that says what may differ, bind tools so lower environments cannot touch production systems, attach evidence to a configuration hash so it expires when staging changes, detect drift, and handle hotfixes without breaking any of it.
What may differ between environments
Start by listing every input that shapes what the agent does, then sort each one into three buckets. The sorting is the design; the code below only enforces it.
| Input | Bucket | Reason |
|---|---|---|
| Image digest, ADK version, instruction text, tool schemas | Must be identical | These are the agent; evidence about one does not transfer to another |
| Model id and generation settings (temperature, max tokens) | Must be identical | A different model is a different agent, even with the same prompt |
| Model endpoint region, project, quota | May differ | Same model, different capacity; affects latency and 429s, not decisions |
| Tool backend base URLs, MCP server addresses | Must differ | Lower environments must never reach production systems |
| Credentials and service accounts | Must differ | A staging pod holding a production key is a production pod |
| Session and artifact store | Must differ | Conversations and files never cross environments |
| Log sampling, trace rates, replica counts | May differ | Operational, not behavioural |
Two entries surprise people. The model id belongs with the release, not the environment, because a staging run on one model says little about another. And the MCP server version is behavioural even though it looks like infrastructure: ADK fetches the tool list from the server at runtime, so a server upgrade changes the tools the model can see without any change to your image.
Architecture
Every promotion asks three questions: is the release identical, do the overlays differ only where allowed, and is the evidence fresh and collected against the configuration staging has now. A no to any of them stops the promotion.
The environment contract as code
Make the contract a file in the repository and check it on every promotion. Resolve each environment's full configuration (release plus overlay, with secrets replaced by their reference names) into a flat map of dotted keys, then compare. Use the layered resolver from the configuration layering article so the map is exactly what the process will read.
# environment-contract.yaml
identical: [release.*, model.id, model.temperature, model.max_output_tokens, agent.instruction_sha256]
must_differ: [tools.*.base_url, mcp.*.url, credentials.*, sessions.store, artifacts.bucket]
may_differ: [model.region, model.project, telemetry.*, scaling.*]
# any key not matched above is treated as "identical" by defaultpublic record Violation(String key, String rule, Object lower, Object upper) {}
public final class EnvironmentContract {
private final List<String> identical, mustDiffer, mayDiffer; // glob patterns from the YAML
public List<Violation> check(Map<String, Object> lower, Map<String, Object> upper) {
var keys = new TreeSet<String>(); keys.addAll(lower.keySet()); keys.addAll(upper.keySet());
var out = new ArrayList<Violation>();
for (String k : keys) {
Object a = lower.get(k), b = upper.get(k);
String rule = matches(mustDiffer, k) ? "must_differ" : matches(mayDiffer, k) ? "may_differ" : "identical";
if (rule.equals("identical") && !Objects.equals(a, b)) out.add(new Violation(k, rule, a, b));
if (rule.equals("must_differ") && Objects.equals(a, b)) out.add(new Violation(k, rule, a, b));
}
return out;
}
}Defaulting unknown keys to identical is deliberate. A new setting that nobody classified fails the check the first time it differs, which forces a decision instead of letting drift in silently. Run the check in CI against the committed overlays and again at promotion time against what is actually deployed, because the two can disagree.
Binding tools per environment
The contract says tool backends must differ; the code has to make that impossible to get wrong. Build tools from an environment binding object, never from scattered environment variables, and refuse to start if a lower environment resolves to anything that looks like production.
public record ToolBindings(String env, URI ordersApi, URI paymentsApi, Set<String> prodHosts) {
public ToolBindings {
if (!env.equals("prod")) {
for (URI u : List.of(ordersApi, paymentsApi))
if (prodHosts.contains(u.getHost()))
throw new IllegalStateException(env + " is bound to production host " + u.getHost());
}
}
}
static List<BaseTool> tools(ToolBindings b, HttpClient http) {
var orders = new OrderTools(http, b.ordersApi());
var payments = new PaymentTools(http, b.paymentsApi());
return List.of(
FunctionTool.create(orders, "lookupOrder"),
FunctionTool.create(payments, "issueRefund", true)); // confirmation in every environment
}The host check is a tripwire, not the defence. The defence is that staging's service account, issued as described in the secrets management article, has no grant on production APIs and its egress policy only allows sandbox hosts, so a misconfiguration fails closed at the network. The tripwire turns that failure into a clear startup error instead of a stream of confusing 403s inside tool results.
Sandbox backends have their own trap: they are too kind. A payments sandbox that approves everything means staging never exercises the refusal branch of your instruction. Give staging fakes a fault profile (a fixed share of declines, timeouts and 429s) so the evidence covers the paths production will hit. Keep confirmation turned on in every environment; turning it off in staging to make tests easier means you are promoting an agent whose confirmation flow you never ran.
Evidence that expires
A promotion is a claim that evidence collected in one environment still describes the release that is about to run in the next. Record the evidence with the configuration hash it was collected against, and let it expire.
{
"release": "sha256:9f3c...e1",
"from": "staging", "to": "prod",
"evidence": [
{"kind": "replay_eval", "run": "eval-2026-10-05-17", "config_hash": "c41a...", "at": "2026-10-05T09:12Z",
"result": {"task_success": 0.912, "baseline": 0.905, "tool_error_rate": 0.011}},
{"kind": "fault_suite", "run": "faults-311", "config_hash": "c41a...", "at": "2026-10-05T10:40Z", "result": "pass"}
],
"approved_by": "oncall-support-agents", "decision": "promote"
}boolean admissible(Evidence e, String currentStagingHash, Instant now, Duration maxAge) {
return e.configHash().equals(currentStagingHash) // staging has not changed since
&& Duration.between(e.at(), now).compareTo(maxAge) <= 0; // and the evidence is recent
}The config hash is the hash of the resolved staging map, including the MCP tool list fetched from the running server. If anyone changes staging after the run, the hash moves and the evidence stops counting. The age limit catches the other drift: dependencies you do not hash, such as a provider changing a model's behaviour behind a stable id. A few days is a common choice; pick it from how often your upstreams change. Store records append-only, so the answer to "why is this running in production" is a lookup, not an investigation.
Sessions and state across environments
Sessions are never promoted. A conversation started in staging stays in staging's session store, and production starts empty for every new release. What does move between environments is the shape of state. If release N+1 reads a new key or reinterprets an old one, the change ships as code that handles both shapes, and staging must contain sessions written by release N to prove it. Keep a small corpus of recorded sessions from the previous release, scrubbed of personal data, and load it into staging before each evaluation.
Keep scopes in mind. ADK state keys prefixed app: are shared by every user of the app, user: keys follow a user across sessions, and temp: keys are stripped before events are persisted, so they never reach the session store. App-scoped values in production, such as feature flags held in state, will not match staging, so either classify them in the contract or move them into configuration where the contract can see them.
Models, quotas and what staging cannot tell you
Keeping the model id identical does not make staging a performance test. Staging usually runs on a smaller quota and possibly another region, so its latency and rate-limit numbers describe staging. Use staging for decision quality (did the agent pick the right tool, refuse the right request, produce the right arguments) and leave latency, cost per session and 429 rates to the production canary, where they are real. If the provider lets you pin a dated model version rather than a floating alias, pin it in the release; an alias that moves under you is unrecorded drift in exactly the input that matters most.
Detecting drift
Drift is any difference between what the promotion record says is running and what is running. Find it by asking the running process, not the repository. Expose an internal endpoint that returns the resolved configuration hash, the instruction hash and the tool names and schema hashes the agent currently sees, including MCP tools, and compare it nightly with the latest record for that environment.
| Drift | How it happens | Response |
|---|---|---|
| Hand-edited variable | Someone patched a value during an incident | Open a ticket; reconcile into the overlay or revert |
| MCP server upgraded | Server team released on their own schedule | Invalidate evidence; re-run staging; pin server versions per environment |
| Model alias moved | Provider repointed a floating name | Pin a dated version; treat as a new release |
| Contract key unclassified | New setting added in code | Classify it; CI fails until then |
Worked example: a refund instruction change
The numbers in this example are illustrative. A support agent's release r42 changes the refund instruction. On Monday it passes staging: task success 0.912 against a production baseline of 0.905 on 600 replayed conversations, and the fault suite is clean. The record stores staging config hash c41a.
On Wednesday the MCP team upgrades staging's ticketing server, which adds a merge_tickets tool. Staging's hash becomes 7be0. On Thursday the promotion job finds that the evidence was collected against c41a, not 7be0, and refuses. The point is not that merge_tickets is dangerous; it is that the model now sees a tool list nobody evaluated r42 against. The re-run shows the agent calling the new tool in 3 of 600 conversations where it should not. The instruction is fixed and r43 is evaluated.
The second gate then catches something else: production's overlay sets model.temperature to 0.2 while staging uses 0.7, left over from an experiment. That key is in the identical bucket, so the contract check fails before any traffic moves. Without it, r43 would have been promoted on evidence from a differently configured agent.
Hotfixes and break-glass
Hotfixes are where environment discipline dies, because the fastest fix is always an edit in production. Give them a real path instead. A hotfix is still a release and still passes through staging, with a reduced evidence set declared in advance: the fault suite and a smoke set, not the full replay. If you truly need break-glass access, make it time-boxed and logged, and have the drift scanner open a reconciliation ticket automatically. A production edit that is never folded back into the overlay will be silently reverted by the next promotion, which turns one incident into two.
Failure modes
| Failure | Symptom | Guard |
|---|---|---|
| Evidence from a different config | Production behaves unlike staging | Config hash on every evidence item |
| Staging bound to production | Test refunds hit real customers | No grants, egress allowlist, startup tripwire |
| Kind sandboxes | Refusal and retry paths untested | Fault profiles on staging fakes |
| Floating model alias | Quality shifts with no deploy | Pinned dated version in the release |
| Old session shapes | Errors only on resumed conversations | Previous-release session corpus in staging |
| Unreconciled hotfix | Fix vanishes on next promotion | Drift scanner and reconciliation ticket |
Trade-offs
A strict contract slows teams down at first: every new setting needs a classification and every staging change invalidates evidence, so evaluations re-run more often. That cost is the price of knowing what you promoted. Teams with very cheap evaluations sometimes skip the hash and simply re-run evidence on every promotion; that is fine as long as the re-run is mandatory. A fourth environment, such as a pre-production with production data access, can catch data-shape problems staging cannot, but it doubles what you must keep in parity. Add one only when a class of incident proves you need it.
What to do next
- List every input to your agent and sort it into identical, must differ and may differ; commit the result as the environment contract.
- Build a resolver that emits each environment's flat configuration map, with secret names instead of values, and run the contract check in CI.
- Construct tools from one bindings object, add the production-host tripwire, and confirm staging's credentials have no production grants.
- Give every staging fake a fault profile and keep confirmation enabled in all environments.
- Attach a config hash and timestamp to each evidence item and refuse promotion on a mismatch or expiry.
- Expose a running-config endpoint, including MCP tool hashes, and schedule a nightly drift comparison.
- Write down the hotfix path and its reduced evidence set before the next incident needs it.