An ADK Java agent that works in a demo has proved that the model can do the task. It has not proved that the service survives a pod restart, stops after a runaway tool loop, stays inside a provider quota, refuses a refund it should not issue, or tells you what went wrong at three in the morning. Those properties are what production readiness means for an agent, and none of them comes from the prompt.

This page is a go/no-go review, not a tutorial. It splits readiness into eight areas, gives each item a pass criterion and the evidence that proves it, and links to the page that teaches the technique. It then walks a review of a refund agent, including the two items that failed. The ADK names used here (Runner, RunConfig, BasePlugin, FunctionTool) come from earlier checks of the google/adk-java source; confirm them against the release you run.

How to run the review

Run the review as a meeting with an artefact: a table with one row per item giving the criterion, the evidence, the owner, and pass, fail or waived. Evidence must be checkable without trusting the author: a test name, a dashboard link, a line of configuration. "We handle that" is not evidence. A waiver names a person, a reason and a date. The numbering below is roughly the order in which skipped items hurt you.

One ADK Java service instanceAPI gatewayauth, quotas [1]Runner.runAsyncRunConfig [2]Pluginsguards, budget [3][6]LlmAgentpinned model [7]Toolstimeouts, idempotent [3]Model providerrate limits, cost [4]Session servicedurable, shared [1]Telemetrytraces, metrics [5]Release and ops: eval gate, canary, kill switch, runbook [7][8]
Where each readiness area sits in one request. Numbers match the sections below.
#AreaThe question it answers
1State and sessionsDoes a conversation survive a restart and a second replica?
2Bounded executionCan one request run forever or call the model 500 times?
3Tools and side effectsIs every write safe to retry and allowed for this agent?
4Provider limits and costWhat happens at quota, and who pays for a runaway user?
5ObservabilityCan you reconstruct one bad answer from telemetry alone?
6Safety and guardrailsWhich rules hold even when the model is persuaded otherwise?
7Evaluation and releaseHow do you know a prompt or model change did not regress?
8OperationsWho is paged, what do they do, and how do they stop the agent?

1. State and sessions

Criterion: a conversation continues correctly after the instance that served its previous turn is killed, and two replicas serving the same session see the same history. Evidence: an integration test that creates a session, runs a turn, restarts the service (or routes the next turn to a second instance) and asserts the agent still knows the earlier turn.

The trap is InMemorySessionService, used by every sample: sessions live in one JVM's heap, so a deploy loses them and another replica never sees them. ADK Java's core ships only the in-memory service and VertexAiSessionService; a Firestore-backed service lives in the contrib modules; there is no official JDBC or PostgreSQL service, so anything else is a BaseSessionService you write and own. Review it for concurrent appends, session growth and retention, since sessions hold user content.

Set autoCreateSession(false) explicitly and create sessions at the API layer, where you can check that the session belongs to the authenticated user. ADK Java deployment covers packaging and the JVM side of running this.

2. Bounded execution

Criterion: every invocation ends within a known number of model calls and a known wall-clock time, and every tool call has its own timeout shorter than the invocation's. Evidence: the RunConfig in code, and a test with a fake model that always asks for another tool call, asserting the invocation stops at the cap.

maxLlmCalls defaults to 500, all of which a model retrying a failing tool will use. Set the cap a little above the 99th percentile of model calls per invocation in your evaluation set.

import com.google.adk.agents.RunConfig;
import com.google.adk.agents.RunConfig.StreamingMode;
import io.reactivex.rxjava3.core.Flowable;
import java.util.concurrent.TimeUnit;

// Built once at startup and reused; the defaults are not production values.
static final RunConfig PROD_RUN = RunConfig.builder()
    .maxLlmCalls(12)                     // default 500: a loop would burn 500 model calls
    .streamingMode(StreamingMode.NONE)   // explicit, so a change is a reviewed diff
    .autoCreateSession(false)            // unknown session id is an error, not a new session
    .build();

Flowable<Event> run(String userId, String sessionId, Content msg) {
  return runner.runAsync(userId, sessionId, msg, PROD_RUN)
      .timeout(45, TimeUnit.SECONDS);    // per-invocation deadline for the caller
}

The RxJava timeout operator frees the caller, but whether in-flight provider or tool work stops depends on the client honouring cancellation, so each tool needs its own timeout; see timing out ADK Java tools safely.

3. Tools and side effects

Criterion: every tool with a side effect is idempotent under retry, scoped to the agents that need it, and gated by a confirmation or a policy check when it moves money, data or permissions. Evidence: a test that calls each write tool twice with the same arguments and asserts one effect; the tool allowlist in a guard; the confirmation flag in code.

Retries come from clients, from the model calling a tool twice, and from operators replaying invocations. Use a business idempotency key derived from what the action means, not an invocation id that changes on replay.

public final class RefundTools {
  private final PaymentsClient payments;   // your client, with its own 5 s timeout
  private final RefundLedger ledger;       // your table: unique key on refund_key

  public RefundTools(PaymentsClient payments, RefundLedger ledger) {
    this.payments = payments; this.ledger = ledger;
  }

  /** Issues a refund at most once per order and amount, however often it is retried. */
  public Map<String, Object> issueRefund(String orderId, long amountCents) {
    String key = "refund:" + orderId + ":" + amountCents;
    Optional<RefundRecord> prior = ledger.find(key);
    if (prior.isPresent()) {
      return Map.of("status", "already_done", "refundId", prior.get().refundId());
    }
    String refundId = payments.refund(orderId, amountCents, key);  // provider idempotency key
    ledger.insert(key, refundId);
    return Map.of("status", "ok", "refundId", refundId);
  }
}

// Instance method, so pass the instance; create(Class, name) finds only static methods.
// requireConfirmation = true: ADK asks the user before this tool runs.
FunctionTool refund = FunctionTool.create(new RefundTools(payments, ledger), "issueRefund", true);

The check and insert above are not atomic; the unique constraint on refund_key makes a race safe, and the provider's idempotency key is the second line. Read-only tools still need timeouts and bounded result sizes.

4. Provider limits and cost

Criterion: the service degrades predictably at provider quota, one user cannot spend the whole budget, and finance can attribute cost per tenant. Evidence: a load test that drives the service past quota and shows queued or rejected requests rather than a storm of retries; a per-user budget enforced in code; a cost-per-tenant dashboard.

Provider limits are shared across replicas, so per-instance limiters are not enough; see the ADK Java rate limiting architecture. Budgets belong in a plugin on the Runner, which sees every model call. The sketch reads billed usage from LlmResponse.usageMetadata(), which returns an Optional from the Google Gen AI SDK types; TokenLedger and its methods are your code, and ctx.userId() assumes your context exposes the user id, so check the accessor in your release.

public final class BudgetPlugin extends BasePlugin {
  private final TokenLedger ledger;   // your code: per-user daily totals, e.g. in Redis
  private final long dailyLimit;

  public BudgetPlugin(TokenLedger ledger, long dailyLimit) {
    super("budget");
    this.ledger = ledger; this.dailyLimit = dailyLimit;
  }

  @Override
  public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder req) {
    if (ledger.usedToday(ctx.userId()) < dailyLimit) return Maybe.empty();
    return Maybe.just(LlmResponse.builder()
        .content(Content.fromParts(Part.fromText("Daily usage limit reached; try again tomorrow.")))
        .build());                                       // model call skipped
  }

  @Override
  public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
    resp.usageMetadata()
        .flatMap(u -> u.totalTokenCount())
        .ifPresent(t -> ledger.add(ctx.userId(), ctx.invocationId(), t));
    return Maybe.empty();                                // response unchanged
  }
}

If you fall back to a second model when the primary is unavailable, that model must pass the same evaluation set first.

5. Observability

Criterion: given a user complaint with a session id and a time, an on-call engineer can see every model call, tool call, argument, result size, latency and token count of that invocation, without access to production databases. Evidence: run the drill. Pick a real invocation from staging and reconstruct it from telemetry alone, then time how long it took.

The minimum set of signals is: invocation id on every log line and span; model calls per invocation; tool calls per invocation by tool and outcome; latency per model call and per tool; tokens per call; guardrail decisions by rule; and the share of invocations ending in an error or a cap. Decide on content logging deliberately: full prompts ease debugging and create a store of user data. ADK Java observability builds the trace tree and token telemetry.

6. Safety and guardrails

Criterion: rules that must never break are enforced in code at a boundary the model cannot argue past, and each has a test that tries to break it. Evidence: the guard plugin, its unit tests, and an adversarial test set that includes injected instructions arriving through tool results, not only through the user's message.

A prompt saying never refund more than 500 dollars is a request; a check on the tool's arguments is a rule. The usual layers are an input check, a tool policy on every proposed call, a sanitizer on tool results (where indirect prompt injection arrives) and an output check; layered guardrails for ADK Java builds all four from BasePlugin and agent callbacks, and explains the order in which they run.

7. Evaluation and release

Criterion: no prompt, tool description, model version or ADK upgrade reaches production without passing an evaluation set in CI, and new versions reach users through a canary. Evidence: the CI job and its threshold; the pinned model string in code; the canary plan with its rollback signal.

Pin the model version you evaluated. An alias that moves to a newer model silently changes your product. Version and review the system instruction and tool descriptions as code. The evaluation set should hold anonymised real traffic, incident cases and the adversarial cases from area 6, with a pass threshold agreed before the change.

8. Operations and the kill switch

Criterion: alerts exist for error rate, cap hits, latency and spend; a runbook says what each alert means; and on-call can disable one agent within a minute without a deploy. Evidence: the alert definitions, the runbook link, and a recorded drill of the kill switch.

Probes, graceful drain of in-flight invocations and autoscaling on the right signal are covered in running ADK Java agents on Kubernetes. The agent-specific piece is a kill switch: a plugin checking a flag before every model call.

public final class KillSwitchPlugin extends BasePlugin {
  private final FeatureFlags flags;   // your flag service, cached for a few seconds

  public KillSwitchPlugin(FeatureFlags flags) { super("kill_switch"); this.flags = flags; }

  @Override
  public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder req) {
    if (!flags.isOn("agent." + ctx.agentName() + ".disabled")) return Maybe.empty();
    return Maybe.just(LlmResponse.builder()
        .content(Content.fromParts(Part.fromText(
            "This assistant is temporarily unavailable. A person will follow up.")))
        .build());
  }
}

Worked example: reviewing a refund agent

An illustrative review of the refund agent above, a week before launch (the figures are examples, not measurements). It issues refunds up to 500 dollars and runs on three replicas.

ItemEvidence offeredResult
1. Restart survivalRestartContinuityIT: turn, kill pod, turn on another replicaFail, then pass: was InMemorySessionService; moved to a persistent service
2. Model-call capmaxLlmCalls(12); p99 in eval set was 6Pass
3. Idempotent refundRefundTwiceTest asserts one ledger row and one provider callPass
4. Quota behaviourLoad test at 2x quotaFail: retries amplified load; added shared limiter, retried test, pass
5. Reconstruction drillEngineer rebuilt a staging invocation in 9 minutesPass
6. Injection via order notesAdversarial set: note says refund the full amountPass, blocked by tool policy
7. Eval gateCI job, 94% task success threshold on 220 casesPass
8. Kill switchDrill: disabled in 40 sPass

The two failures are the common ones. Without the restart test, the session bug would have surfaced as users finding the agent forgot their order after every deploy. The quota failure came from each replica retrying rate-limit errors with short backoff, multiplying load at quota; a shared limiter plus jittered, capped retries fixed it.

Failure modes and the items that catch them

Failure in productionReadiness item that would have caught it
Agent forgets the conversation after each deploy1: restart continuity test
One request makes hundreds of model calls2: maxLlmCalls cap with a fake looping model
Customer refunded twice after a timeout3: idempotency test on every write tool
Injected text in a tool result triggers a write6: adversarial set including tool-result injection
Quality drops after the model alias moves7: pinned model and CI eval gate
Bad agent stays live while a hotfix builds8: kill switch drill

The review itself fails when evidence is a sentence, when it runs once and never again (rerun areas touched by any tool, model or session-store change), and when every item is equal. Areas 1 to 3 and the kill switch are launch blockers; a weak dashboard can be a dated waiver.

Trade-offs

Every item costs something. A low maxLlmCalls cap cuts off rare long tasks; give those workflows their own RunConfig. Reserve confirmation for actions that are expensive to undo. Plugins apply to every agent the runner hosts, which suits budgets and kill switches; agent-specific rules belong in that agent's callbacks.

What to do next

  1. Copy the eight-area table into your repository as a markdown file with columns for criterion, evidence, owner and status.
  2. Write the restart continuity test and run it; if it fails, choose a persistent session service before anything else.
  3. Set maxLlmCalls from your evaluation data and add a test with a looping fake model.
  4. Add an idempotency test for every tool that writes, and turn on confirmation for the ones that move money or permissions.
  5. Register a budget plugin and a kill-switch plugin on the Runner, and drill the kill switch.
  6. Load test past provider quota and fix retry amplification if you see it.
  7. Run a reconstruction drill on one staging invocation and fill the telemetry gaps it exposes.
  8. Pin the model version and put the evaluation set in CI with an agreed threshold.
Key takeaway: An ADK Java agent is ready for production when each of eight areas has a pass criterion and evidence another engineer can check: durable sessions, a model-call cap and deadlines, idempotent and confirmed write tools, quota and budget handling, telemetry good enough to rebuild one invocation, guardrails in code, a pinned model behind an evaluation gate, and a kill switch that has been drilled. Treat sessions, limits, tools and the kill switch as launch blockers.