An agent built with the Agent Development Kit for Java is a loop: the runner sends the conversation to a model, the model either answers or asks for tools, the tools run, and their results go back to the model until it produces a final response. Observing that loop is different from observing a normal service. The same request can take one model call or seven, cost a fraction of a cent or several dollars, and fail because of a prompt, a tool, stale session state or the model itself.

This article describes an observability architecture for ADK Java built from what the library actually provides. Class names, span names, attribute keys and the environment variable below were checked against the google/adk-java sources; generic OpenTelemetry GenAI conventions documented for other ADK languages are not assumed to apply to Java unless the Java source shows them. The architecture has four parts: the spans ADK emits by itself, a plugin that turns callbacks into metrics, correlation ids that join traces to session events and logs, and a collector that enforces privacy and sampling.

ADK Java telemetry: built-in spans, a plugin for metrics, a collector for policyRunner.runAsyncone invocationinvoke_agentspan per agent runcall_llmmodel, tokens, finishexecute_tooltool name, args, resultTelemetry pluginbefore/after callbackstapsMicrometer / OTelcounters, histogramsSession eventsdurable recordidsOTel SDKglobal tracer, batch exporttracer gcp.vertex.agentCollectorredact, tail-sample, span metricsOTLPTrace backendDashboards and alertsEval pipelinesampled traces scored offlineADK creates the spans; your plugin adds business metrics; the collector enforces privacy and sampling.
ADK Java observability: the library emits invoke_agent, call_llm and execute_tool spans; a plugin adds metrics; the collector applies policy.

What ADK Java emits by itself

ADK Java instruments its own loop with OpenTelemetry. The tracing code lives in com.google.adk.telemetry.Tracing and uses a tracer named gcp.vertex.agent. It records four operations, set as the gen_ai.operation.name attribute: invoke_agent for an agent run, call_llm for a model request, execute_tool for a tool call, and send_data for data sent on a live connection. Spans carry standard GenAI attributes such as gen_ai.agent.name, gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons and gen_ai.tool.name, plus ADK-specific keys including gcp.vertex.agent.invocation_id, gcp.vertex.agent.session_id, gcp.vertex.agent.llm_request, gcp.vertex.agent.llm_response, gcp.vertex.agent.tool_call_args and gcp.vertex.agent.tool_response.

Those last four matter for privacy. The source reads an environment variable, ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS, which defaults to true. Left alone, full model requests and responses, tool arguments and tool results are written into span attributes and exported to whatever trace backend you use. That is invaluable in development and a data-protection problem in production, where prompts contain user messages and tool results contain records from your systems. Set it to false in production so those attributes are written as empty JSON, and decide deliberately where full content may live.

What ADK Java emits by itself

ADK Java instruments its own loop with OpenTelemetry. The tracing code lives in com.google.adk.telemetry.Tracing and uses a tracer named gcp.vertex.agent. It records four operations, set as the gen_ai.operation.name attribute: invoke_agent for an agent run, call_llm for a model request, execute_tool for a tool call, and send_data for data sent on a live connection. Spans carry standard GenAI attributes such as gen_ai.agent.name, gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons and gen_ai.tool.name, plus ADK-specific keys including gcp.vertex.agent.invocation_id, gcp.vertex.agent.session_id, gcp.vertex.agent.llm_request, gcp.vertex.agent.llm_response, gcp.vertex.agent.tool_call_args and gcp.vertex.agent.tool_response.

Those last four matter for privacy. The source reads an environment variable, ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS, which defaults to true. Left alone, full model requests and responses, tool arguments and tool results are written into span attributes and exported to whatever trace backend you use. That is invaluable in development and a data-protection problem in production, where prompts contain user messages and tool results contain records from your systems. Set it to false in production so those attributes are written as empty JSON, and decide deliberately where full content may live.

Wiring OpenTelemetry

ADK creates spans, but something must export them. The usual pattern is to build the OpenTelemetry SDK at startup and register it as the global instance before you build the Runner, or to attach the OpenTelemetry Java agent, which registers a global instance for you. Confirm in your ADK version how Tracing.getTracer() obtains its tracer; if spans never appear, an unregistered SDK is the first suspect.

Resource resource = Resource.getDefault().merge(Resource.create(
        Attributes.of(AttributeKey.stringKey("service.name"), "support-agent",
                      AttributeKey.stringKey("service.version"), BuildInfo.version())));

SdkTracerProvider tracerProvider = SdkTracerProvider.builder()
        .setResource(resource)
        .setSampler(Sampler.parentBased(Sampler.alwaysOn()))   // sample in the collector instead
        .addSpanProcessor(BatchSpanProcessor.builder(
                OtlpGrpcSpanExporter.builder().setEndpoint("http://otel-collector:4317").build())
                .build())
        .build();

OpenTelemetrySdk.builder().setTracerProvider(tracerProvider).buildAndRegisterGlobal();

Runner runner = Runner.builder()
        .agent(rootAgent)
        .appName("support")
        .sessionService(sessionService)
        .plugins(new TelemetryPlugin(meterRegistry))
        .build();

Export to a local OpenTelemetry Collector rather than directly to a vendor. The collector is where you drop or hash sensitive attributes, apply tail sampling (keep every trace with an error or high cost, a fraction of the rest), derive request and latency metrics from spans with the span-metrics connector, and switch backends without redeploying agents. Add the shutdown hook that flushes the batch processor, or the last traces before a pod stops are lost.

A plugin as the metrics tap

Spans answer "what happened in this conversation". Dashboards need aggregates by agent, model and version, and business numbers that ADK cannot know, such as cost per tenant. ADK Java's Plugin interface (in com.google.adk.plugins, with BasePlugin as a named base class) is the right tap point because the runner calls it for every agent, model call and tool, so instrumentation lives in one place. Its callbacks include beforeRunCallback, afterRunCallback, onRunErrorCallback, beforeModelCallback, afterModelCallback, onModelErrorCallback, beforeToolCallback, afterToolCallback, onToolErrorCallback and onEventCallback.

The rule that matters most: the model and tool callbacks return a Maybe, and a non-empty value replaces the real work. A beforeModelCallback that returns an LlmResponse skips the model call entirely. A telemetry plugin must therefore always return Maybe.empty() and must never throw, or a metrics bug becomes an outage.

public final class TelemetryPlugin extends BasePlugin {
    private final MeterRegistry meters;

    public TelemetryPlugin(MeterRegistry meters) {
        super("telemetry");
        this.meters = meters;
    }

    @Override
    public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
        try {
            String agent = ctx.agentName();
            resp.usageMetadata().ifPresent(u -> {
                u.promptTokenCount().ifPresent(n ->
                        meters.counter("agent.tokens", "agent", agent, "dir", "input").increment(n));
                u.candidatesTokenCount().ifPresent(n ->
                        meters.counter("agent.tokens", "agent", agent, "dir", "output").increment(n));
            });
            resp.finishReason().ifPresent(r ->
                    meters.counter("agent.finish", "agent", agent, "reason", r.toString()).increment());
        } catch (RuntimeException e) {
            // telemetry must never break a turn
        }
        return Maybe.empty();
    }

    @Override
    public Maybe<Map<String, Object>> onToolErrorCallback(
            BaseTool tool, Map<String, Object> args, ToolContext ctx, Throwable error) {
        try {
            meters.counter("agent.tool.errors", "tool", tool.name(),
                    "type", error.getClass().getSimpleName()).increment();
        } catch (RuntimeException ignored) { }
        return Maybe.empty();      // let ADK's normal error handling proceed
    }

    @Override
    public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
        try {
            meters.counter("agent.run.errors", "type", error.getClass().getSimpleName()).increment();
        } catch (RuntimeException ignored) { }
        return Completable.complete();
    }
}

Keep tag values bounded: agent name, model, tool name, error class and release version are fine; user id, session id or prompt text as tags will explode your metrics backend. Per-user figures belong in traces or in the session store, not in metric labels.

Correlation: traces, session events and logs

Three records describe one conversation: the trace, the session events ADK appends for every user message, model response, tool call and state change, and your application logs. They are only useful together if they share ids. ADK already puts the invocation id and session id on its spans. Put the trace id and invocation id on every log line (through the logging MDC), and store the trace id in your own records where you persist results, so a support engineer can go from a user complaint to the session, then to the exact trace.

Treat the session store as the durable record and traces as the fast, sampled index. Exporters drop data under backpressure and samplers discard most traces by design, but session events are written as part of the agent's work. When content capture is off in spans, the session store is also where the full conversation lives, under the access controls and retention rules you already apply to user data.

ADK's runner is built on RxJava, so work moves between threads. MDC values and OpenTelemetry context set on the request thread do not follow automatically into tool code on another scheduler. Set log context inside the plugin callbacks, which run with the invocation's context objects, rather than relying on a servlet filter, and test that a tool's log lines carry the right ids.

The signals worth keeping

SignalSourceWhy it matters
Tokens by agent, model and directionafterModelCallbackcost and context growth
Model calls per invocationspan counts or a plugin counterdetects runaway loops
Invocation and model-call latencyspan metrics from the collectoruser-facing speed
Finish reasonsafterModelCallbacktruncation and safety stops
Tool error rate by tool and classonToolErrorCallbackbroken integrations
Model errors by classonModelErrorCallbackquota, timeouts, provider incidents
Run errorsonRunErrorCallbackfailed conversations
Quality score per releaseeval pipeline over sampled tracessilent regressions

Tag everything with the release version, including prompt and model version if they ship separately. Agent behaviour changes with prompt edits that never touch Java code, and a dashboard split by version is how "the agent feels worse" becomes a measurable difference. For dashboard layouts and cost attribution see metrics dashboards and Gemini cost tracking.

Sampling and evaluation

Agent traces are large. One invocation can hold a dozen spans, and with content capture on each model span carries the whole conversation so far, so trace volume grows roughly with the square of conversation length. Keeping every trace is rarely affordable, but deciding at the start of a request, with a head sampler, throws away failures at the same rate as successes. Sample on outcome instead. Let the SDK keep everything, and configure the collector's tail sampling to keep every trace that contains an error, every trace whose token total or model-call count crosses a threshold, every trace from a new release for its first hours, and a small random share of the rest as a baseline.

That retained set has a second job: evaluation. Agent failures are often not errors at all but plausible, wrong answers. Feed a steady sample of production conversations, taken from the session store with the trace id attached, into the same scorers you use before release, and publish the score per release next to latency and cost. A drop in quality with flat error rates is exactly the regression that operational metrics cannot see, and the trace id lets a reviewer open the conversation behind any low score.

Worked example: a doubled token bill

On a Tuesday, the daily token bill for a support agent doubles. The token counter split by agent shows the increase is all input tokens on the billing_specialist sub-agent, starting with release 4.12. Model calls per invocation have risen from a median of 2 to 5. The collector kept every trace over a cost threshold, so there are thousands to choose from; one shows call_llm spans alternating with execute_tool spans for lookup_invoice, each returning a tool error that the model then retries with slightly different arguments.

The tool error counter confirms it: lookup_invoice errors by class IllegalArgumentException began with the same release, which changed the tool's date parameter format without updating its description. Each retry re-sends the whole growing conversation, so input tokens climb with every loop. The fix is to restore the description and add a loop limit. The lesson for the architecture is that three signals working together, a cost counter by agent, a loop-depth metric and tail-sampled traces, turned a vague bill increase into a root cause in minutes.

Failure modes

  • Content leakage: ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS left at its default ships user messages and tool results to a trace vendor. See PII redaction for scrubbing patterns.
  • Accidental short circuit: a telemetry callback that returns a value instead of Maybe.empty() replaces the model or tool result.
  • Telemetry exceptions: an uncaught exception in a plugin fails the turn.
  • No spans at all: the SDK was never registered globally, or was registered after the runner started.
  • Lost spans at shutdown: no flush of the batch processor.
  • Cardinality explosion: session or user ids used as metric tags.
  • Broken correlation: log lines from tools on other threads have no trace id.
  • Head sampling hides failures: a 1% head sampler keeps 1% of the errors; sample in the collector on outcome instead.

Trade-offs

ChoiceGainCost
Content in spansfull replay of any tracesensitive data in the trace backend
Content only in session storeone governed copytwo lookups to debug
Tail sampling in collectorkeeps every error and expensive tracecollector memory and complexity
Plugin metricsone place, every agentmust be defensive and bounded
Span-derived metricsno code changesonly what spans already record

For the per-call view of prompts and completions see prompt and completion tracing, and for where callbacks fit in the agent lifecycle see ADK Java callbacks.

What to do next

  1. Register the OpenTelemetry SDK globally before building the Runner and confirm ADK spans reach a backend.
  2. Set ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false in production and decide where full content may live.
  3. Run a collector with attribute redaction, tail sampling on errors and cost, and span metrics.
  4. Add a telemetry plugin that records tokens, finish reasons and errors, returns Maybe.empty() and never throws.
  5. Put trace id and invocation id on every log line, including tool code on other threads.
  6. Tag metrics with release, prompt and model versions; keep user and session ids out of tags.
  7. Alert on model calls per invocation and tokens per agent, not only on latency.
  8. Feed a sample of production traces into your eval pipeline and chart quality per release.
Key takeaway: ADK Java already traces its loop under the gcp.vertex.agent tracer, so the architecture work is around it: export through a collector, switch off content capture in production, add a defensive plugin for tokens and errors, and share ids across traces, session events and logs so any bad answer can be traced to its cause.