Most teams have three environments for an ADK Java agent and assume they mean something: dev for building, staging for proving, production for users. In practice the three drift apart quietly. Staging points at a different model version because someone was testing a preview. Production's MCP server was upgraded on its own schedule and now exposes two extra tools. Staging's payment tool talks to a sandbox that always succeeds, so the refusal path was never exercised. When the release is promoted, the evidence collected in staging describes an agent that does not exist in production.

This article is about the environments, not the pipeline. Building once, release manifests and canaries are covered in the CI/CD pipeline article, and ramps and rollback in the deploy pipeline article. Here you will define an environment contract that says what may differ, bind tools so lower environments cannot touch production systems, attach evidence to a configuration hash so it expires when staging changes, detect drift, and handle hotfixes without breaking any of it.

What may differ between environments

Start by listing every input that shapes what the agent does, then sort each one into three buckets. The sorting is the design; the code below only enforces it.

InputBucketReason
Image digest, ADK version, instruction text, tool schemasMust be identicalThese are the agent; evidence about one does not transfer to another
Model id and generation settings (temperature, max tokens)Must be identicalA different model is a different agent, even with the same prompt
Model endpoint region, project, quotaMay differSame model, different capacity; affects latency and 429s, not decisions
Tool backend base URLs, MCP server addressesMust differLower environments must never reach production systems
Credentials and service accountsMust differA staging pod holding a production key is a production pod
Session and artifact storeMust differConversations and files never cross environments
Log sampling, trace rates, replica countsMay differOperational, not behavioural

Two entries surprise people. The model id belongs with the release, not the environment, because a staging run on one model says little about another. And the MCP server version is behavioural even though it looks like infrastructure: ADK fetches the tool list from the server at runtime, so a server upgrade changes the tools the model can see without any change to your image.

Architecture

One release, three overlays, one contract check per promotionRelease (immutable)digest, instruction hashmodel id, tool schemasdev overlayfakes, laptopstaging overlaysandbox toolsprod overlayreal toolsContract checkdiff only in allowed keysresolvedEvidenceeval run, config hash, timePromotion gatefresh? same hash? contract?Promotion recordappend-only ledgerpass/failapproveDrift scannernightly, per environmentinvalidatesThe release never changes between environments. Overlays may differ only where the contract allows, and evidence is tied to a config hash.
The release is fixed; overlays are checked against the contract and evidence is bound to the configuration it was collected on.

Every promotion asks three questions: is the release identical, do the overlays differ only where allowed, and is the evidence fresh and collected against the configuration staging has now. A no to any of them stops the promotion.

The environment contract as code

Make the contract a file in the repository and check it on every promotion. Resolve each environment's full configuration (release plus overlay, with secrets replaced by their reference names) into a flat map of dotted keys, then compare. Use the layered resolver from the configuration layering article so the map is exactly what the process will read.

# environment-contract.yaml
identical:   [release.*, model.id, model.temperature, model.max_output_tokens, agent.instruction_sha256]
must_differ: [tools.*.base_url, mcp.*.url, credentials.*, sessions.store, artifacts.bucket]
may_differ:  [model.region, model.project, telemetry.*, scaling.*]
# any key not matched above is treated as "identical" by default
public record Violation(String key, String rule, Object lower, Object upper) {}

public final class EnvironmentContract {
  private final List<String> identical, mustDiffer, mayDiffer;   // glob patterns from the YAML

  public List<Violation> check(Map<String, Object> lower, Map<String, Object> upper) {
    var keys = new TreeSet<String>(); keys.addAll(lower.keySet()); keys.addAll(upper.keySet());
    var out = new ArrayList<Violation>();
    for (String k : keys) {
      Object a = lower.get(k), b = upper.get(k);
      String rule = matches(mustDiffer, k) ? "must_differ" : matches(mayDiffer, k) ? "may_differ" : "identical";
      if (rule.equals("identical") && !Objects.equals(a, b)) out.add(new Violation(k, rule, a, b));
      if (rule.equals("must_differ") && Objects.equals(a, b)) out.add(new Violation(k, rule, a, b));
    }
    return out;
  }
}

Defaulting unknown keys to identical is deliberate. A new setting that nobody classified fails the check the first time it differs, which forces a decision instead of letting drift in silently. Run the check in CI against the committed overlays and again at promotion time against what is actually deployed, because the two can disagree.

Binding tools per environment

The contract says tool backends must differ; the code has to make that impossible to get wrong. Build tools from an environment binding object, never from scattered environment variables, and refuse to start if a lower environment resolves to anything that looks like production.

public record ToolBindings(String env, URI ordersApi, URI paymentsApi, Set<String> prodHosts) {
  public ToolBindings {
    if (!env.equals("prod")) {
      for (URI u : List.of(ordersApi, paymentsApi))
        if (prodHosts.contains(u.getHost()))
          throw new IllegalStateException(env + " is bound to production host " + u.getHost());
    }
  }
}

static List<BaseTool> tools(ToolBindings b, HttpClient http) {
  var orders = new OrderTools(http, b.ordersApi());
  var payments = new PaymentTools(http, b.paymentsApi());
  return List.of(
      FunctionTool.create(orders, "lookupOrder"),
      FunctionTool.create(payments, "issueRefund", true));   // confirmation in every environment
}

The host check is a tripwire, not the defence. The defence is that staging's service account, issued as described in the secrets management article, has no grant on production APIs and its egress policy only allows sandbox hosts, so a misconfiguration fails closed at the network. The tripwire turns that failure into a clear startup error instead of a stream of confusing 403s inside tool results.

Sandbox backends have their own trap: they are too kind. A payments sandbox that approves everything means staging never exercises the refusal branch of your instruction. Give staging fakes a fault profile (a fixed share of declines, timeouts and 429s) so the evidence covers the paths production will hit. Keep confirmation turned on in every environment; turning it off in staging to make tests easier means you are promoting an agent whose confirmation flow you never ran.

Evidence that expires

A promotion is a claim that evidence collected in one environment still describes the release that is about to run in the next. Record the evidence with the configuration hash it was collected against, and let it expire.

{
  "release": "sha256:9f3c...e1",
  "from": "staging", "to": "prod",
  "evidence": [
    {"kind": "replay_eval", "run": "eval-2026-10-05-17", "config_hash": "c41a...", "at": "2026-10-05T09:12Z",
     "result": {"task_success": 0.912, "baseline": 0.905, "tool_error_rate": 0.011}},
    {"kind": "fault_suite", "run": "faults-311", "config_hash": "c41a...", "at": "2026-10-05T10:40Z", "result": "pass"}
  ],
  "approved_by": "oncall-support-agents", "decision": "promote"
}
boolean admissible(Evidence e, String currentStagingHash, Instant now, Duration maxAge) {
  return e.configHash().equals(currentStagingHash)            // staging has not changed since
      && Duration.between(e.at(), now).compareTo(maxAge) <= 0;  // and the evidence is recent
}

The config hash is the hash of the resolved staging map, including the MCP tool list fetched from the running server. If anyone changes staging after the run, the hash moves and the evidence stops counting. The age limit catches the other drift: dependencies you do not hash, such as a provider changing a model's behaviour behind a stable id. A few days is a common choice; pick it from how often your upstreams change. Store records append-only, so the answer to "why is this running in production" is a lookup, not an investigation.

Sessions and state across environments

Sessions are never promoted. A conversation started in staging stays in staging's session store, and production starts empty for every new release. What does move between environments is the shape of state. If release N+1 reads a new key or reinterprets an old one, the change ships as code that handles both shapes, and staging must contain sessions written by release N to prove it. Keep a small corpus of recorded sessions from the previous release, scrubbed of personal data, and load it into staging before each evaluation.

Keep scopes in mind. ADK state keys prefixed app: are shared by every user of the app, user: keys follow a user across sessions, and temp: keys are stripped before events are persisted, so they never reach the session store. App-scoped values in production, such as feature flags held in state, will not match staging, so either classify them in the contract or move them into configuration where the contract can see them.

Models, quotas and what staging cannot tell you

Keeping the model id identical does not make staging a performance test. Staging usually runs on a smaller quota and possibly another region, so its latency and rate-limit numbers describe staging. Use staging for decision quality (did the agent pick the right tool, refuse the right request, produce the right arguments) and leave latency, cost per session and 429 rates to the production canary, where they are real. If the provider lets you pin a dated model version rather than a floating alias, pin it in the release; an alias that moves under you is unrecorded drift in exactly the input that matters most.

Detecting drift

Drift is any difference between what the promotion record says is running and what is running. Find it by asking the running process, not the repository. Expose an internal endpoint that returns the resolved configuration hash, the instruction hash and the tool names and schema hashes the agent currently sees, including MCP tools, and compare it nightly with the latest record for that environment.

DriftHow it happensResponse
Hand-edited variableSomeone patched a value during an incidentOpen a ticket; reconcile into the overlay or revert
MCP server upgradedServer team released on their own scheduleInvalidate evidence; re-run staging; pin server versions per environment
Model alias movedProvider repointed a floating namePin a dated version; treat as a new release
Contract key unclassifiedNew setting added in codeClassify it; CI fails until then

Worked example: a refund instruction change

The numbers in this example are illustrative. A support agent's release r42 changes the refund instruction. On Monday it passes staging: task success 0.912 against a production baseline of 0.905 on 600 replayed conversations, and the fault suite is clean. The record stores staging config hash c41a.

On Wednesday the MCP team upgrades staging's ticketing server, which adds a merge_tickets tool. Staging's hash becomes 7be0. On Thursday the promotion job finds that the evidence was collected against c41a, not 7be0, and refuses. The point is not that merge_tickets is dangerous; it is that the model now sees a tool list nobody evaluated r42 against. The re-run shows the agent calling the new tool in 3 of 600 conversations where it should not. The instruction is fixed and r43 is evaluated.

The second gate then catches something else: production's overlay sets model.temperature to 0.2 while staging uses 0.7, left over from an experiment. That key is in the identical bucket, so the contract check fails before any traffic moves. Without it, r43 would have been promoted on evidence from a differently configured agent.

Hotfixes and break-glass

Hotfixes are where environment discipline dies, because the fastest fix is always an edit in production. Give them a real path instead. A hotfix is still a release and still passes through staging, with a reduced evidence set declared in advance: the fault suite and a smoke set, not the full replay. If you truly need break-glass access, make it time-boxed and logged, and have the drift scanner open a reconciliation ticket automatically. A production edit that is never folded back into the overlay will be silently reverted by the next promotion, which turns one incident into two.

Failure modes

FailureSymptomGuard
Evidence from a different configProduction behaves unlike stagingConfig hash on every evidence item
Staging bound to productionTest refunds hit real customersNo grants, egress allowlist, startup tripwire
Kind sandboxesRefusal and retry paths untestedFault profiles on staging fakes
Floating model aliasQuality shifts with no deployPinned dated version in the release
Old session shapesErrors only on resumed conversationsPrevious-release session corpus in staging
Unreconciled hotfixFix vanishes on next promotionDrift scanner and reconciliation ticket

Trade-offs

A strict contract slows teams down at first: every new setting needs a classification and every staging change invalidates evidence, so evaluations re-run more often. That cost is the price of knowing what you promoted. Teams with very cheap evaluations sometimes skip the hash and simply re-run evidence on every promotion; that is fine as long as the re-run is mandatory. A fourth environment, such as a pre-production with production data access, can catch data-shape problems staging cannot, but it doubles what you must keep in parity. Add one only when a class of incident proves you need it.

What to do next

  1. List every input to your agent and sort it into identical, must differ and may differ; commit the result as the environment contract.
  2. Build a resolver that emits each environment's flat configuration map, with secret names instead of values, and run the contract check in CI.
  3. Construct tools from one bindings object, add the production-host tripwire, and confirm staging's credentials have no production grants.
  4. Give every staging fake a fault profile and keep confirmation enabled in all environments.
  5. Attach a config hash and timestamp to each evidence item and refuse promotion on a mismatch or expiry.
  6. Expose a running-config endpoint, including MCP tool hashes, and schedule a nightly drift comparison.
  7. Write down the hotfix path and its reduced evidence set before the next incident needs it.
Key takeaway: Promotion is only meaningful if the environments differ where you intend and nowhere else. Write a contract that classifies every input, build tools from bindings that cannot reach production from staging, tie each piece of evidence to the configuration hash it was collected against, and scan running processes for drift. Then a promotion record answers what is running and why, and staging evidence actually describes the agent your users will meet.