An agent built with the Agent Development Kit for Java is a loop: the runner sends the conversation to a model, the model either answers or asks for tools, the tools run, and their results go back to the model until it produces a final response. Observing that loop is different from observing a normal service. The same request can take one model call or seven, cost a fraction of a cent or several dollars, and fail because of a prompt, a tool, stale session state or the model itself.
This article describes an observability architecture for ADK Java built from what the library actually provides. Class names, span names, attribute keys and the environment variable below were checked against the google/adk-java sources; generic OpenTelemetry GenAI conventions documented for other ADK languages are not assumed to apply to Java unless the Java source shows them. The architecture has four parts: the spans ADK emits by itself, a plugin that turns callbacks into metrics, correlation ids that join traces to session events and logs, and a collector that enforces privacy and sampling.
What ADK Java emits by itself
ADK Java instruments its own loop with OpenTelemetry. The tracing code lives in com.google.adk.telemetry.Tracing and uses a tracer named gcp.vertex.agent. It records four operations, set as the gen_ai.operation.name attribute: invoke_agent for an agent run, call_llm for a model request, execute_tool for a tool call, and send_data for data sent on a live connection. Spans carry standard GenAI attributes such as gen_ai.agent.name, gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons and gen_ai.tool.name, plus ADK-specific keys including gcp.vertex.agent.invocation_id, gcp.vertex.agent.session_id, gcp.vertex.agent.llm_request, gcp.vertex.agent.llm_response, gcp.vertex.agent.tool_call_args and gcp.vertex.agent.tool_response.
Those last four matter for privacy. The source reads an environment variable, ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS, which defaults to true. Left alone, full model requests and responses, tool arguments and tool results are written into span attributes and exported to whatever trace backend you use. That is invaluable in development and a data-protection problem in production, where prompts contain user messages and tool results contain records from your systems. Set it to false in production so those attributes are written as empty JSON, and decide deliberately where full content may live.
What ADK Java emits by itself
ADK Java instruments its own loop with OpenTelemetry. The tracing code lives in com.google.adk.telemetry.Tracing and uses a tracer named gcp.vertex.agent. It records four operations, set as the gen_ai.operation.name attribute: invoke_agent for an agent run, call_llm for a model request, execute_tool for a tool call, and send_data for data sent on a live connection. Spans carry standard GenAI attributes such as gen_ai.agent.name, gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons and gen_ai.tool.name, plus ADK-specific keys including gcp.vertex.agent.invocation_id, gcp.vertex.agent.session_id, gcp.vertex.agent.llm_request, gcp.vertex.agent.llm_response, gcp.vertex.agent.tool_call_args and gcp.vertex.agent.tool_response.
Those last four matter for privacy. The source reads an environment variable, ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS, which defaults to true. Left alone, full model requests and responses, tool arguments and tool results are written into span attributes and exported to whatever trace backend you use. That is invaluable in development and a data-protection problem in production, where prompts contain user messages and tool results contain records from your systems. Set it to false in production so those attributes are written as empty JSON, and decide deliberately where full content may live.
Wiring OpenTelemetry
ADK creates spans, but something must export them. The usual pattern is to build the OpenTelemetry SDK at startup and register it as the global instance before you build the Runner, or to attach the OpenTelemetry Java agent, which registers a global instance for you. Confirm in your ADK version how Tracing.getTracer() obtains its tracer; if spans never appear, an unregistered SDK is the first suspect.
Resource resource = Resource.getDefault().merge(Resource.create(
Attributes.of(AttributeKey.stringKey("service.name"), "support-agent",
AttributeKey.stringKey("service.version"), BuildInfo.version())));
SdkTracerProvider tracerProvider = SdkTracerProvider.builder()
.setResource(resource)
.setSampler(Sampler.parentBased(Sampler.alwaysOn())) // sample in the collector instead
.addSpanProcessor(BatchSpanProcessor.builder(
OtlpGrpcSpanExporter.builder().setEndpoint("http://otel-collector:4317").build())
.build())
.build();
OpenTelemetrySdk.builder().setTracerProvider(tracerProvider).buildAndRegisterGlobal();
Runner runner = Runner.builder()
.agent(rootAgent)
.appName("support")
.sessionService(sessionService)
.plugins(new TelemetryPlugin(meterRegistry))
.build();Export to a local OpenTelemetry Collector rather than directly to a vendor. The collector is where you drop or hash sensitive attributes, apply tail sampling (keep every trace with an error or high cost, a fraction of the rest), derive request and latency metrics from spans with the span-metrics connector, and switch backends without redeploying agents. Add the shutdown hook that flushes the batch processor, or the last traces before a pod stops are lost.
A plugin as the metrics tap
Spans answer "what happened in this conversation". Dashboards need aggregates by agent, model and version, and business numbers that ADK cannot know, such as cost per tenant. ADK Java's Plugin interface (in com.google.adk.plugins, with BasePlugin as a named base class) is the right tap point because the runner calls it for every agent, model call and tool, so instrumentation lives in one place. Its callbacks include beforeRunCallback, afterRunCallback, onRunErrorCallback, beforeModelCallback, afterModelCallback, onModelErrorCallback, beforeToolCallback, afterToolCallback, onToolErrorCallback and onEventCallback.
The rule that matters most: the model and tool callbacks return a Maybe, and a non-empty value replaces the real work. A beforeModelCallback that returns an LlmResponse skips the model call entirely. A telemetry plugin must therefore always return Maybe.empty() and must never throw, or a metrics bug becomes an outage.
public final class TelemetryPlugin extends BasePlugin {
private final MeterRegistry meters;
public TelemetryPlugin(MeterRegistry meters) {
super("telemetry");
this.meters = meters;
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
try {
String agent = ctx.agentName();
resp.usageMetadata().ifPresent(u -> {
u.promptTokenCount().ifPresent(n ->
meters.counter("agent.tokens", "agent", agent, "dir", "input").increment(n));
u.candidatesTokenCount().ifPresent(n ->
meters.counter("agent.tokens", "agent", agent, "dir", "output").increment(n));
});
resp.finishReason().ifPresent(r ->
meters.counter("agent.finish", "agent", agent, "reason", r.toString()).increment());
} catch (RuntimeException e) {
// telemetry must never break a turn
}
return Maybe.empty();
}
@Override
public Maybe<Map<String, Object>> onToolErrorCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx, Throwable error) {
try {
meters.counter("agent.tool.errors", "tool", tool.name(),
"type", error.getClass().getSimpleName()).increment();
} catch (RuntimeException ignored) { }
return Maybe.empty(); // let ADK's normal error handling proceed
}
@Override
public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
try {
meters.counter("agent.run.errors", "type", error.getClass().getSimpleName()).increment();
} catch (RuntimeException ignored) { }
return Completable.complete();
}
}Keep tag values bounded: agent name, model, tool name, error class and release version are fine; user id, session id or prompt text as tags will explode your metrics backend. Per-user figures belong in traces or in the session store, not in metric labels.
Correlation: traces, session events and logs
Three records describe one conversation: the trace, the session events ADK appends for every user message, model response, tool call and state change, and your application logs. They are only useful together if they share ids. ADK already puts the invocation id and session id on its spans. Put the trace id and invocation id on every log line (through the logging MDC), and store the trace id in your own records where you persist results, so a support engineer can go from a user complaint to the session, then to the exact trace.
Treat the session store as the durable record and traces as the fast, sampled index. Exporters drop data under backpressure and samplers discard most traces by design, but session events are written as part of the agent's work. When content capture is off in spans, the session store is also where the full conversation lives, under the access controls and retention rules you already apply to user data.
ADK's runner is built on RxJava, so work moves between threads. MDC values and OpenTelemetry context set on the request thread do not follow automatically into tool code on another scheduler. Set log context inside the plugin callbacks, which run with the invocation's context objects, rather than relying on a servlet filter, and test that a tool's log lines carry the right ids.
The signals worth keeping
| Signal | Source | Why it matters |
|---|---|---|
| Tokens by agent, model and direction | afterModelCallback | cost and context growth |
| Model calls per invocation | span counts or a plugin counter | detects runaway loops |
| Invocation and model-call latency | span metrics from the collector | user-facing speed |
| Finish reasons | afterModelCallback | truncation and safety stops |
| Tool error rate by tool and class | onToolErrorCallback | broken integrations |
| Model errors by class | onModelErrorCallback | quota, timeouts, provider incidents |
| Run errors | onRunErrorCallback | failed conversations |
| Quality score per release | eval pipeline over sampled traces | silent regressions |
Tag everything with the release version, including prompt and model version if they ship separately. Agent behaviour changes with prompt edits that never touch Java code, and a dashboard split by version is how "the agent feels worse" becomes a measurable difference. For dashboard layouts and cost attribution see metrics dashboards and Gemini cost tracking.
Sampling and evaluation
Agent traces are large. One invocation can hold a dozen spans, and with content capture on each model span carries the whole conversation so far, so trace volume grows roughly with the square of conversation length. Keeping every trace is rarely affordable, but deciding at the start of a request, with a head sampler, throws away failures at the same rate as successes. Sample on outcome instead. Let the SDK keep everything, and configure the collector's tail sampling to keep every trace that contains an error, every trace whose token total or model-call count crosses a threshold, every trace from a new release for its first hours, and a small random share of the rest as a baseline.
That retained set has a second job: evaluation. Agent failures are often not errors at all but plausible, wrong answers. Feed a steady sample of production conversations, taken from the session store with the trace id attached, into the same scorers you use before release, and publish the score per release next to latency and cost. A drop in quality with flat error rates is exactly the regression that operational metrics cannot see, and the trace id lets a reviewer open the conversation behind any low score.
Worked example: a doubled token bill
On a Tuesday, the daily token bill for a support agent doubles. The token counter split by agent shows the increase is all input tokens on the billing_specialist sub-agent, starting with release 4.12. Model calls per invocation have risen from a median of 2 to 5. The collector kept every trace over a cost threshold, so there are thousands to choose from; one shows call_llm spans alternating with execute_tool spans for lookup_invoice, each returning a tool error that the model then retries with slightly different arguments.
The tool error counter confirms it: lookup_invoice errors by class IllegalArgumentException began with the same release, which changed the tool's date parameter format without updating its description. Each retry re-sends the whole growing conversation, so input tokens climb with every loop. The fix is to restore the description and add a loop limit. The lesson for the architecture is that three signals working together, a cost counter by agent, a loop-depth metric and tail-sampled traces, turned a vague bill increase into a root cause in minutes.
Failure modes
- Content leakage:
ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANSleft at its default ships user messages and tool results to a trace vendor. See PII redaction for scrubbing patterns. - Accidental short circuit: a telemetry callback that returns a value instead of
Maybe.empty()replaces the model or tool result. - Telemetry exceptions: an uncaught exception in a plugin fails the turn.
- No spans at all: the SDK was never registered globally, or was registered after the runner started.
- Lost spans at shutdown: no flush of the batch processor.
- Cardinality explosion: session or user ids used as metric tags.
- Broken correlation: log lines from tools on other threads have no trace id.
- Head sampling hides failures: a 1% head sampler keeps 1% of the errors; sample in the collector on outcome instead.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Content in spans | full replay of any trace | sensitive data in the trace backend |
| Content only in session store | one governed copy | two lookups to debug |
| Tail sampling in collector | keeps every error and expensive trace | collector memory and complexity |
| Plugin metrics | one place, every agent | must be defensive and bounded |
| Span-derived metrics | no code changes | only what spans already record |
For the per-call view of prompts and completions see prompt and completion tracing, and for where callbacks fit in the agent lifecycle see ADK Java callbacks.
What to do next
- Register the OpenTelemetry SDK globally before building the Runner and confirm ADK spans reach a backend.
- Set ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false in production and decide where full content may live.
- Run a collector with attribute redaction, tail sampling on errors and cost, and span metrics.
- Add a telemetry plugin that records tokens, finish reasons and errors, returns Maybe.empty() and never throws.
- Put trace id and invocation id on every log line, including tool code on other threads.
- Tag metrics with release, prompt and model versions; keep user and session ids out of tags.
- Alert on model calls per invocation and tokens per agent, not only on latency.
- Feed a sample of production traces into your eval pipeline and chart quality per release.