Most production problems with ADK for Java agents are not wrong answers. They are a request that never finishes, an invocation that fails after hundreds of model calls, state that a tool set and the next turn cannot see, an exception that disappears without a useful log line, or a pod whose heap climbs until it is killed. Each leaves evidence in the same places: the OpenTelemetry spans ADK emits, the session's event log and the JVM itself. Debugging is quick when you know which of those to open first for a given symptom, and slow when you start from whatever dashboard happens to be open.

This article is a triage method for those non-semantic failures. It maps symptoms to evidence, adds a small diagnostics plugin so every failure carries an invocation id, shows how to read a session's event log, and walks five failure classes with their causes and fixes. When the agent runs fine but says something false, the method is different and is covered in debugging hallucinations. Containing an incident while you debug it is the job of the incident response playbook. API names below were checked against the google-adk 1.11.0 jar on Maven Central; confirm them against the release you run.

Start from the symptom

Let the reported symptom choose the first evidence source; opening the wrong one first is where most debugging time goes.

SymptomOpen firstUsual causes
Spinner never ends; client times outDiagnostics log: run.start without run.end; then a thread dumpBlocking call with no timeout in a tool or callback; a reactive chain that never completes
Fails after a long runcall_llm spans per invocation; LlmCallsLimitExceededExceptionAgents transferring back and forth; a tool that keeps returning a retryable error
A value set last turn is missingEvent log: the stateDelta of each eventtemp: key; state mutated outside the context; two invocations racing on one session
HTTP 500 with no useful logonRunErrorCallback and the RxJava error handlerError raised after the subscriber was disposed; an exception logged without ids
Latency and restarts creep up over hoursHeap histogram, JFR recording, thread dumpUnbounded caches of runners or sessions; an in-memory session service in production

The evidence chain

The evidence chain: one id threads through four sourcesSymptomticket, alert, SLORequest logrequestId to invocationIdTraceinvocation span treeEvent loglistEvents / getSessioninvoke_agentone per agent runcall_llmmodel latency, tokensexecute_tooltool latency, argsEventscalls, deltas, errorsJVM evidencethread dump, JFR, heap histogramDiagnostics pluginrun.start / run.end / errorswhen spans never endsame invocationIdUnfinished spans are never exported, so a hung invocation isvisible in logs and thread dumps before it is visible in traces
Request ids lead to invocation ids, which lead to the span tree and the session's events. When spans never end, logs and the JVM are the only evidence.

ADK wraps each run in a span named invocation, each agent run in invoke_agent, each model call in call_llm and each tool execution in execute_tool. Its spans carry attributes under the gcp.vertex.agent. prefix, including invocation_id and session_id, and the model spans can carry the serialized request and response. That last point matters twice: it is the richest evidence you have, and it puts prompts and user data into your tracing backend, so decide deliberately what is exported, as discussed in ADK Java observability.

The second source is the event log. Every non-partial event the runner produces is appended to the session: the user message, each model response, each function call and function response, and the state changes each one made. It is keyed by app name, user id and session id, and it outlives the trace retention window in most deployments. The Event contract explains each field. The third source is the JVM, which you need exactly when the first two are silent: a request that is stuck has an open span that will never be exported, and no final event.

A diagnostics plugin that makes failures findable

A plugin sees every run, model call and tool call, so it can guarantee one structured line with ids per failure. Every callback below exists on Plugin in 1.11.0. Each returns an empty Maybe or a completed Completable, so it observes without changing behaviour. Returning a non-empty response from an error callback would replace the error with that response, which is a decision, not a logging change.

public final class DiagnosticsPlugin extends BasePlugin {
  private static final Logger log = LoggerFactory.getLogger(DiagnosticsPlugin.class);
  private record Run(long startNanos, AtomicInteger llmCalls) {}
  private final ConcurrentMap<String, Run> runs = new ConcurrentHashMap<>();

  public DiagnosticsPlugin() { super("diagnostics"); }

  @Override
  public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
    runs.put(ctx.invocationId(), new Run(System.nanoTime(), new AtomicInteger()));
    log.info("run.start inv={} app={} user={} session={}",
        ctx.invocationId(), ctx.appName(), ctx.userId(), ctx.session().id());
    return Maybe.empty();
  }

  @Override
  public Maybe<LlmResponse> beforeModelCallback(CallbackContext cb, LlmRequest.Builder req) {
    Run r = runs.get(cb.invocationId());
    if (r != null && r.llmCalls().incrementAndGet() % 20 == 0) {
      log.warn("run.many_llm_calls inv={} agent={} calls={}",
          cb.invocationId(), cb.agentName(), r.llmCalls().get());
    }
    return Maybe.empty();
  }

  @Override
  public Maybe<LlmResponse> onModelErrorCallback(
      CallbackContext cb, LlmRequest.Builder req, Throwable err) {
    log.error("model.error inv={} agent={}", cb.invocationId(), cb.agentName(), err);
    return Maybe.empty();                       // empty: the error propagates unchanged
  }

  @Override
  public Maybe<Map<String, Object>> onToolErrorCallback(
      BaseTool tool, Map<String, Object> args, ToolContext tc, Throwable err) {
    log.error("tool.error inv={} tool={} call={}",
        tc.invocationId(), tool.name(), tc.functionCallId().orElse("?"), err);
    return Maybe.empty();
  }

  @Override
  public Completable onRunErrorCallback(InvocationContext ctx, Throwable err) {
    finish(ctx, "error");
    log.error("run.error inv={}", ctx.invocationId(), err);
    return Completable.complete();
  }

  @Override
  public Completable afterRunCallback(InvocationContext ctx) {
    finish(ctx, "ok");
    return Completable.complete();
  }

  private void finish(InvocationContext ctx, String outcome) {
    Run r = runs.remove(ctx.invocationId());        // remove on both paths, or the map leaks
    if (r != null) {
      log.info("run.end inv={} outcome={} ms={} llmCalls={}", ctx.invocationId(), outcome,
          (System.nanoTime() - r.startNanos()) / 1_000_000, r.llmCalls().get());
    }
  }
}

Two details are deliberate. The ids are passed explicitly on every line instead of through SLF4J's MDC, because MDC is thread-local and an RxJava chain hops threads; a value set at the HTTP edge is usually gone by the time a tool runs. And a run.start with no run.end is itself the signal for a hung invocation, which no trace backend will show you while the span is still open.

Reading a session&#x27;s event log

When a specific session misbehaved, dump its events before forming a theory. Run this from an admin endpoint, never on a request thread.

Session s = sessionService
    .getSession(appName, userId, sessionId, Optional.empty())
    .blockingGet();
if (s == null) throw new IllegalArgumentException("no such session");
for (Event e : s.events()) {
  System.out.printf("%d inv=%s author=%s calls=%s responses=%s delta=%s transfer=%s err=%s final=%b%n",
      e.timestamp(), e.invocationId(), e.author(),
      e.functionCalls().stream().map(f -> f.name().orElse("?")).toList(),
      e.functionResponses().stream().map(f -> f.name().orElse("?")).toList(),
      e.actions().stateDelta().keySet(),
      e.actions().transferToAgent().orElse("-"),
      e.errorCode().map(Object::toString).orElse("-"),
      e.finalResponse());
}

Read it as a sequence of invocations. Each user turn starts a new invocation id. Inside one, look for four shapes. A function call with no matching function response is a tool that never returned or whose result was lost. Alternating transfer values between two agents is a routing loop. A stateDelta that contains the key you expected tells you the write happened, so the problem is on the read side. An errorCode on a model event means the model call itself reported a problem, such as a safety block, without any exception being thrown. How state prefixes and context objects behave is covered in session context in depth.

Runaway invocations

Every invocation counts its model calls, and when the count exceeds RunConfig.maxLlmCalls() ADK throws LlmCallsLimitExceededException. The builder default in 1.11.0 is 500, which is a backstop, not a budget: a healthy conversational turn makes a handful of calls, so an invocation that reaches 500 has been looping for minutes and spending money the whole time. Set the limit from data, roughly three times the p99 calls per invocation you measure with the plugin above.

RunConfig runConfig = RunConfig.builder()
    .setMaxLlmCalls(30)          // measured p99 was 9 calls per invocation
    .build();
runner.runAsync(userId, sessionId, userMessage, runConfig);

The causes show up clearly in the event log. Two agents whose descriptions both claim the same task hand the conversation back and forth, so transferToAgent alternates. A tool returns an error string such as "try again later" and the model obliges, so the same function call repeats with identical arguments. A loop agent's exit condition depends on state that no sub-agent writes. Fix the cause, then keep the lower limit so the next loop costs thirty calls instead of five hundred.

Hung invocations

A hung invocation produces no error, so error-rate alerts stay green while users stare at a spinner. The diagnostics log shows run.start without run.end, and the event log's last event is usually a model function call with no response. Take two thread dumps a few seconds apart with jcmd <pid> Thread.print and look for threads in the same frame both times. RxJava's own pools are easy to spot by name, RxComputationThreadPool-N and RxCachedThreadScheduler-N.

The common finding is a tool making a blocking call with no read timeout to a slowed dependency; if the thread belongs to a bounded pool, other invocations queue behind it. The fix has three parts: a connect and read timeout on every client a tool uses, a reactive .timeout(...) on any Single or Maybe a callback returns, and returning a structured error result, for example {"status": "unavailable"}, so the model can tell the user. Also dispose the subscription when the client disconnects, or abandoned invocations keep running and billing.

Errors that vanish

RxJava has one place for errors it cannot deliver: an error raised after the subscriber was disposed, a second error on a finished stream, or a subscribe() call with no error consumer. Those go to the global handler in RxJavaPlugins, and by default end up on standard error, often without your log format and without ids. Install a handler at startup so they become first-class log lines and a metric.

RxJavaPlugins.setErrorHandler(e -> {
  Throwable cause = (e instanceof UndeliverableException) ? e.getCause() : e;
  undeliverableCounter.increment();
  log.error("rx.undeliverable", cause);
});

The second way errors vanish is that they are not exceptions at all. A model response can carry an error code and message that ADK copies onto the event, and the run then completes normally. Alert on events with errorCode set, not only on thrown exceptions.

State that does not stick

"The agent forgot" usually means the write never became durable. Four causes cover nearly every case. Keys with the temp: prefix are skipped when an event's state delta is applied to the session, by design. Partial events, the streaming chunks, are not appended at all. Code that mutates session.state() directly instead of writing through toolContext.state() or callbackContext.state() produces no state delta; an in-memory session service hides this, a database-backed one exposes it. Finally, two invocations on one session at once, from a double submit or a retrying client, start from the same snapshot, and the second write either wins silently or is rejected, depending on the session service.

The event dump distinguishes them: a key present in a delta was overwritten later or raced; a key absent everywhere was never durable. Serialize invocations per session at your API layer.

Slow death: memory and threads

Some failures take hours. Heap that only grows, a full GC every few minutes and restarts by the orchestrator point at something retained per tenant, per session or per invocation. jcmd <pid> GC.class_histogram taken twice an hour apart shows which classes grow; jcmd <pid> JFR.start duration=120s filename=agent.jfr records allocation, locks and threads at low overhead. The usual culprits are unbounded maps of runners or agent trees, the in-memory session service in a production profile, and per-invocation maps not cleaned on the error path. Bound every cache and give it a metric.

Worked example: the 03:10 latency cliff

A support agent's p95 latency jumps from 4 seconds to the 60-second client timeout at 03:10, while the error rate stays flat. Triage card first: spinner, so start with the diagnostics log, not the trace UI. About 4 percent of run.start lines since 03:10 have no run.end. Picking three of those invocation ids and dumping their sessions shows the same final event each time: a model function call to order_lookup with no function response.

Two thread dumps ten seconds apart show 64 threads parked in a socket read under the order lookup tool's HTTP client, identical in both. The client was built with a connect timeout but no read timeout, and the order service started returning headers and then stalling during a database failover. Because the tool blocked a thread from a bounded pool, healthy requests queued behind the stuck ones, which is why latency rose for everyone.

The fix shipped in that order: a 5-second read timeout, a structured unavailable result the instruction tells the model to explain, a .timeout on the tool's reactive wrapper as a second guard, and an alert on invocations open longer than 45 seconds. No error alert could have fired: nothing failed, things only waited.

Operational guidance

Make the evidence exist before the incident. Keep traces with tail sampling that retains every errored or slow invocation and a small fraction of the rest. Keep event logs at least as long as your support ticket SLA, since tickets arrive days after the session. Alert on four signals that error rates miss: invocations open past the client timeout, p99 model calls per invocation, tool error rate per tool, and the undeliverable-error counter. Redact serialized prompts from spans where policy requires it.

Trade-offs

Rich spans and verbose logs make debugging fast and cost storage, money and privacy review; sample and redact rather than turning them off. A low maxLlmCalls limit caps runaway cost but can cut off a legitimately long task, so give long-running agents their own RunConfig instead of raising the global one. Aggressive timeouts free threads but abandon slow work that would have succeeded, so pair them with idempotent tools. Serializing per session removes races but makes a slow turn block the next.

What to do next

  1. Log invocation id, trace id and your request id together at the HTTP edge.
  2. Add the diagnostics plugin and alert on run.start lines with no run.end past the client timeout.
  3. Build an admin endpoint that dumps a session's events in the format above.
  4. Measure p99 model calls per invocation and set setMaxLlmCalls to about three times that.
  5. Audit every tool's HTTP and database clients for connect and read timeouts.
  6. Install an RxJavaPlugins error handler with a counter.
  7. Grep for direct session.state() mutation and temp: keys that are expected to persist.
  8. Serialize invocations per session id at the API layer.
Key takeaway: Production ADK Java failures that are not wrong answers fall into a few classes: runaway, hung, silent, forgetful and slowly leaking. Pick the first evidence source from the symptom, carry the invocation id everywhere, read the session's event log before theorising, and use thread dumps when spans are silent. Then turn each finding into a limit, a timeout or an alert.