Logs are how you answer 'what happened in this conversation' after the fact. Metrics show that tool errors rose, and traces show where time went. Logs hold the decisions, outcomes and error details for each invocation, in a form you can grep or query. An agent makes this harder than a normal service. One user turn fans out into several model calls and tool calls across RxJava threads, and the obvious thing to log, the prompt, is the thing you must not log.

This page starts with what ADK Java 1.11.0 logs by itself, read from its bytecode rather than assumed. It then covers the built-in LoggingPlugin, a JSON log line design, a Logback setup that also catches the parts of ADK that do not use SLF4J, a structured logging plugin, and the volume and sampling decisions that keep the pipeline affordable. For explicit-context logging helpers, see observability as a first-class concern. For records that must survive disputes, see audit logging in ADK Java. An audit trail is a separate stream from operational logs.

What ADK Java logs by itself

In google-adk 1.11.0, 57 core classes log through the SLF4J API. ADK does not choose a backend for you: its POM lists slf4j-simple only at test scope, so without a binding on your classpath SLF4J falls back to a no-op logger and you see nothing. The core flow class, BaseLlmFlow, logs almost entirely at DEBUG. Those messages are control-flow notes such as 'Ending flow execution because max steps reached.' and 'Pausing flow execution on a pending long-running call.', and also event: {} functionCalls: {}, which prints whole events, content included. It logs at ERROR for 'LLM calls limit exceeded.' and for an unknown transfer target.

That gives you two rules at once. Leave com.google.adk at INFO or WARN in production, because DEBUG writes user messages and model replies into your logs. And alert on the ERROR lines, because they mean an invocation ended abnormally.

The exception is the analytics package. The seven classes under com.google.adk.plugins.agentanalytics, the BigQuery Agent Analytics plugin, use java.util.logging. Without a bridge, their warnings, including 'BigQuery event queue is full, dropping event.', go to the JDK console handler. They skip your JSON pipeline, your log levels and your alerts.

LoggingPlugin: a development tool, not a log schema

com.google.adk.plugins.LoggingPlugin implements every plugin callback and logs each one at INFO, through SLF4J, as [{}] {} with the plugin name (logging_plugin by default) and a multi-line block. The blocks are labelled USER MESSAGE RECEIVED, INVOCATION STARTING, AGENT STARTING, LLM REQUEST, LLM RESPONSE, TOOL STARTING, TOOL COMPLETED, TOOL ERROR, EVENT YIELDED, AGENT COMPLETED and INVOCATION COMPLETED. Each block holds indented lines such as invocation id, agent name, model, available tools, system instruction, token usage, function call id, arguments and results. Content is truncated to 200 characters and arguments to 300.

Runner runner = new InMemoryRunner(rootAgent, "support_app", List.of(new LoggingPlugin()));

That is very useful on a laptop: you watch the agent think in order. In production it works against you. Multi-line messages break line-oriented shippers unless you configure multiline joining. The fields are text, not keys, so you cannot filter on tool without regular expressions. And 200 characters of user content and 300 of tool arguments is still personal data. Use it in development and in short debugging sessions on one instance. For production, write a plugin that emits structured fields and leaves content out.

Design the log line first

Decide the line before writing code. One event per line, a fixed set of event names, stable key names, and no free text from users or models by default:

FieldExampleWhy
eventtool_endEnumerated names: invocation_start/end, llm_end, tool_end, tool_error, run_error
invocation_ide-7c1...Joins every line from one turn
session_ids-41a...Joins turns of one conversation
useru_3fa91c0bPseudonymized; never the raw account id or email
agentrefund_agentWhich agent in a multi-agent tree
tool, call_idget_order, adk-...Pairs start and end of one call
duration_ms412Latency without a trace lookup
outcomeok / error / blockedWhat dashboards group by
error_classTimeoutExceptionClass name, not the message, which can contain data
tokens_in, tokens_out1840, 212Cost questions from logs alone
trace_id, span_idadded by appenderJump from a log line to the trace

Logback, JSON and trace ids

One JSON log pipeline for an ADK Java serviceADK coreSLF4J, mostly DEBUGagentanalyticsjava.util.loggingStructuredLogPluginSLF4J INFO, kv fieldsYour toolsSLF4J + MDCjul-to-slf4jbridge handlerOpenTelemetryAppenderadds trace_id, span_idAsyncAppenderneverBlock, queueJSON linesLogstashEncoderLog storequery by invocation_idThe MDC wrapper must sit outside the async appender: it reads the span on the calling thread.
SLF4J and bridged JUL output pass through the trace-id wrapper on the calling thread, then an async appender writes JSON lines.

Use Logback as the backend, logstash-logback-encoder for JSON, jul-to-slf4j for the analytics package, and the OpenTelemetry opentelemetry-logback-mdc-1.0 library to stamp trace ids. Pin versions from Maven Central for your build. The configuration that matters:

<configuration>
  <contextListener class="ch.qos.logback.classic.jul.LevelChangePropagator">
    <resetJUL>true</resetJUL>
  </contextListener>

  <appender name="JSON" class="ch.qos.logback.core.ConsoleAppender">
    <encoder class="net.logstash.logback.encoder.LogstashEncoder"/>
  </appender>
  <appender name="ASYNC" class="ch.qos.logback.classic.AsyncAppender">
    <queueSize>8192</queueSize>
    <neverBlock>true</neverBlock>
    <appender-ref ref="JSON"/>
  </appender>
  <!-- Outermost: reads the current span on the logging thread, before the async hop. -->
  <appender name="OTEL" class="io.opentelemetry.instrumentation.logback.mdc.v1_0.OpenTelemetryAppender">
    <appender-ref ref="ASYNC"/>
  </appender>

  <logger name="com.google.adk" level="INFO"/>
  <root level="INFO"><appender-ref ref="OTEL"/></root>
</configuration>
// Once, at startup, before building the runner:
SLF4JBridgeHandler.removeHandlersForRootLogger();
SLF4JBridgeHandler.install();

The order of the appenders matters. The MDC wrapper adds trace_id, span_id and trace_flags from the span that is current when the log call happens. Put it inside the async appender and it runs on Logback's worker thread, where no span is current. LevelChangePropagator keeps JUL level checks cheap once the bridge is installed. If you run the OpenTelemetry Java agent, it can inject the same MDC keys itself, so check its Logback support for your version instead of adding the library. See OpenTelemetry integration in ADK Java for setting up the tracer.

A structured logging plugin

The plugin below writes one structured line per lifecycle step. It uses StructuredArguments.kv so each value becomes a JSON key, not text in the message. Every callback returns Maybe.empty(), so the plugin observes and never changes the run, and every body is guarded so a logging bug cannot fail a turn.

public final class StructuredLogPlugin extends BasePlugin {
  private static final Logger LOG = LoggerFactory.getLogger("agent.events");
  private final Map<String, Long> toolStart = new ConcurrentHashMap<>();
  private final Map<String, Long> runStart = new ConcurrentHashMap<>();

  public StructuredLogPlugin() { super("structured_log"); }

  @Override public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
    runStart.put(ctx.invocationId(), System.nanoTime());
    safe(() -> LOG.info("invocation_start", kv("event", "invocation_start"),
        kv("invocation_id", ctx.invocationId()), kv("session_id", ctx.session().id()),
        kv("user", Pseudo.user(ctx.userId())), kv("agent", ctx.agent().name())));
    return Maybe.empty();
  }

  @Override public Maybe<LlmResponse> afterModelCallback(CallbackContext cc, LlmResponse r) {
    safe(() -> {
      var u = r.usageMetadata();
      LOG.info("llm_end", kv("event", "llm_end"), kv("invocation_id", cc.invocationId()),
          kv("agent", cc.agentName()),
          kv("finish", r.finishReason().map(Object::toString).orElse(null)),
          kv("error_code", r.errorCode().map(Object::toString).orElse(null)),
          kv("tokens_in", u.flatMap(m -> m.promptTokenCount()).orElse(null)),
          kv("tokens_out", u.flatMap(m -> m.candidatesTokenCount()).orElse(null)));
    });
    return Maybe.empty();
  }

  @Override public Maybe<Map<String, Object>> beforeToolCallback(BaseTool t, Map<String, Object> a, ToolContext tc) {
    tc.functionCallId().ifPresent(id -> toolStart.put(id, System.nanoTime()));
    return Maybe.empty();
  }

  @Override public Maybe<Map<String, Object>> afterToolCallback(BaseTool t, Map<String, Object> a, ToolContext tc, Map<String, Object> res) {
    safe(() -> LOG.info("tool_end", toolFields(t, tc, "ok", null)));
    return Maybe.empty();
  }

  @Override public Maybe<Map<String, Object>> onToolErrorCallback(BaseTool t, Map<String, Object> a, ToolContext tc, Throwable e) {
    safe(() -> LOG.warn("tool_error", toolFields(t, tc, "error", e.getClass().getSimpleName())));
    return Maybe.empty();   // empty = let ADK's normal error handling proceed
  }

  @Override public Completable afterRunCallback(InvocationContext ctx) {
    Long t0 = runStart.remove(ctx.invocationId());
    safe(() -> LOG.info("invocation_end", kv("event", "invocation_end"),
        kv("invocation_id", ctx.invocationId()),
        kv("duration_ms", t0 == null ? null : (System.nanoTime() - t0) / 1_000_000)));
    return Completable.complete();
  }

  private Object[] toolFields(BaseTool t, ToolContext tc, String outcome, String errorClass) {
    String id = tc.functionCallId().orElse(null);
    Long t0 = id == null ? null : toolStart.remove(id);
    return new Object[] {kv("event", outcome.equals("ok") ? "tool_end" : "tool_error"),
        kv("invocation_id", tc.invocationId()), kv("agent", tc.agentName()), kv("tool", t.name()),
        kv("call_id", id), kv("outcome", outcome), kv("error_class", errorClass),
        kv("duration_ms", t0 == null ? null : (System.nanoTime() - t0) / 1_000_000)};
  }

  private static void safe(Runnable r) {
    try { r.run(); } catch (RuntimeException e) { /* never fail a turn because of logging */ }
  }
}

Two details are deliberate. Arguments and results are not logged. If a field from them is needed, such as an order status, add it by name in the tool, after deciding it is safe. And the timing maps are keyed by function call id and invocation id, then removed on completion, so they cannot grow without bound. Also clear runStart in onRunErrorCallback, which the sketch omits. These are hand-written lines, so they carry no span unless one is current. The appender adds trace ids when the callback runs inside ADK's spans.

Context across RxJava threads

The plugin passes ids explicitly, so it does not depend on the MDC. Your own tool code often does depend on it, through log lines that expect invocation_id from the MDC. MDC values are thread locals and RxJava moves work between threads. One global fix covers work scheduled through RxJava schedulers: copy the MDC when a task is scheduled and restore it when the task runs.

RxJavaPlugins.setScheduleHandler(task -> {
  Map<String, String> captured = MDC.getCopyOfContextMap();
  return () -> {
    Map<String, String> previous = MDC.getCopyOfContextMap();
    if (captured == null) MDC.clear(); else MDC.setContextMap(captured);
    try { task.run(); }
    finally { if (previous == null) MDC.clear(); else MDC.setContextMap(previous); }
  };
});

This does not cover executors outside RxJava, such as an HTTP client's callback pool, and a stale MDC value is worse than none. Prefer explicit fields for anything you will query on, and treat the MDC as a convenience.

Volume, levels and sampling

Estimate volume before shipping. Take 50 turns per second, with each turn writing an invocation start and end, two model calls and three tool calls: about 7 lines. At roughly 500 bytes per JSON line, that is 175 KB/s, about 15 GB a day before compression. This is manageable, but it grows linearly with tool calls, and an agent stuck in a loop multiplies it.

  • Keep every WARN and ERROR. They are rare and they are the point.
  • Sample INFO by invocation, not by line. Hash the invocation id and keep, for example, 20% of invocations in full. A sampled invocation that is missing its tool lines is useless for debugging.
  • Know the async queue's behaviour. By default Logback's AsyncAppender drops TRACE, DEBUG and INFO events once less than 20% of the queue is free, and with neverBlock it drops instead of stalling the caller. That trades log completeness for request latency, which is the right choice for operational logs and the wrong one for audit records.
  • Set retention by purpose. Debugging needs days to weeks. Anything kept longer belongs in the analytics or audit stream.

Worked example: a timeout storm

Suppose support reports 'the agent says it cannot find orders' at 14:05. A query for event=tool_error AND tool=get_order over the last hour, grouped by error_class, returns 312 TimeoutException lines, all after 13:58. Choose one invocation_id and pull every line for it, ordered by time:

{"event":"invocation_start","invocation_id":"e-9f2","agent":"support_root","user":"u_3fa91c0b"}
{"event":"llm_end","invocation_id":"e-9f2","finish":"STOP","tokens_in":1712,"tokens_out":38}
{"event":"tool_error","invocation_id":"e-9f2","tool":"get_order","call_id":"adk-71c","error_class":"TimeoutException","duration_ms":10003,"trace_id":"4bf9..."}
{"event":"llm_end","invocation_id":"e-9f2","finish":"STOP","tokens_in":1790,"tokens_out":61}
{"event":"invocation_end","invocation_id":"e-9f2","duration_ms":11420}

The tool took 10,003 ms, so it hit a 10-second client timeout. The model then made a second call and answered politely without the data. The trace_id opens the trace, which shows the order service's database span waiting for a connection. Logs found the failing dependency in two queries, and no customer text was needed. For the tool metrics that should have alerted first, see tool observability and metrics.

Failure modes

  • No SLF4J binding. ADK logs nowhere, and nobody notices until an incident.
  • DEBUG on com.google.adk in production. Whole events, with user text, go into the log store.
  • Unbridged JUL. Analytics drop warnings go to stderr and never reach an alert.
  • LoggingPlugin left on. Multi-line blocks are split by the shipper into fragments that cannot be queried.
  • Logging exception messages. e.getMessage() from a tool often includes the request, so log the class and a sanitized code instead.
  • Trace wrapper inside the async appender. Every line has an empty trace_id.

What to do next

  1. Add Logback, the JSON encoder, jul-to-slf4j and the OpenTelemetry MDC library, and copy the configuration above.
  2. Install SLF4JBridgeHandler at startup, and confirm that an analytics plugin warning appears as JSON.
  3. Set com.google.adk to INFO and add a test that fails if it is configured at DEBUG in the production profile.
  4. Write down your event names and field table, then register the structured plugin on the runner.
  5. Remove LoggingPlugin from production configuration and keep it behind a dev profile.
  6. Add invocation-level sampling for INFO and alert on ERROR lines from BaseLlmFlow.
  7. Practise the worked example: from one tool_error line, reach the trace in two steps. Then read the ADK Java observability architecture to fit logs alongside metrics and traces.
Key takeaway: ADK Java logs through SLF4J but brings no backend. Its analytics package uses java.util.logging, and at DEBUG its flow logs whole events. Give it a JSON backend, bridge JUL, keep com.google.adk at INFO, stamp trace ids on the calling thread, and replace the human-readable LoggingPlugin with a plugin that writes one keyed line per lifecycle step and no content. Sample by invocation, keep every error, and you can go from one bad line to its trace in two queries.