An agent's tools are where it touches the world: databases, payment APIs, search, ticketing. When an agent is slow, wrong or expensive, the cause is usually a tool, and yet tool behaviour is the part teams instrument last. They trace the model call, see a turn take twelve seconds, and cannot say whether the time went to the model, to a slow downstream, or to a tool that failed three times and was retried by the model.

Recent ADK Java releases do more of this for you than most teams realise. The source of release v1.10.1 (published 18 September 2026) and of the main branch opens a span for every tool execution and records three histograms per call. This article shows exactly what is recorded, how to make sure it actually reaches your backend, what is missing, and how to add the outcome metrics and SLOs that turn raw telemetry into something you can page on.

Advertisement

What ADK records for every tool call

When the model returns a function call, the flow resolves the tool and wraps the whole execution, including the before and after callbacks, in a span named execute_tool <tool name>, for example execute_tool get_order. The tracer and the meter are both obtained from the global OpenTelemetry instance under the instrumentation name gcp.vertex.agent. The span carries GenAI semantic-convention attributes plus ADK-specific ones:

AttributeValue
gen_ai.operation.nameexecute_tool
gen_ai.tool.name, gen_ai.tool.descriptionFrom the tool declaration
gen_ai.tool.typeThe tool's Java class simple name, for example FunctionTool
gen_ai.tool_call.idThe id that ties the call to the model's function call
gcp.vertex.agent.tool_call_argsSerialised arguments, if content capture is on
gcp.vertex.agent.tool_responseSerialised response, if content capture is on
gcp.vertex.agent.event_idThe function-response event id in the session

When the model asks for several tools in one response, ADK subscribes to all of them eagerly (whether they actually overlap depends on whether each tool's work is asynchronous or on its own scheduler), each gets its own span, and ADK adds a span named execute_tool (merged) around the merged response event. Note that the README in the telemetry package still describes spans named tool_call [name] and tool_response [name]; the code is authoritative, so build dashboards on what your exporter actually receives. The ADK tracing walkthrough shows how these spans nest under the agent and model-call spans; it is written for the Python ADK, so its span names differ slightly.

When the span closes, ADK records three histograms, all with the attributes gen_ai.tool.name and gen_ai.agent.name:

InstrumentUnitWhat it measuresExtra attribute
gen_ai.tool.execution.durationmsWall time of the execution including callbackserror.type = exception class simple name when the call threw
gen_ai.tool.request.sizeBySize of the argumentsnone
gen_ai.tool.response.sizeBySize of the function responsenone

The same class records gen_ai.agent.invocation.duration and agent request and response sizes, so a single meter gives you agent-level and tool-level latency on the same axes. All of this was read from the release source; if you run an older version, check whether com.google.adk.telemetry.Metrics exists in your jar before relying on it.

The data flow, end to end

LLM responsefunction call partsexecute_tool get_orderspan opened by ADKOTel SDK (global)tracer + meter gcp.vertex.agentbeforeToolCallbackpolicy, cache, counttool bodyyour Java codeafterToolCallbackoutcome class, redactonToolErrorCallbackexceptionsthrowsgen_ai.tool.execution.durationhistogram, msgen_ai.tool.request.sizeand response.size, Bycustom: tool outcome countertool, agent, outcomeon closeexportADK records span attributes and three histograms per tool call; everything about business outcome is yours to add.
One tool call. The span wraps the before callback, the tool body and the after callback; histograms are recorded when it closes; custom outcome metrics come from your callbacks.

Two consequences follow from where the span sits. First, callback latency is charged to the tool: a slow policy check in a before callback shows up as a slow tool. Second, a before callback that returns a value short-circuits the tool, as the dispatch mechanics article explains, so a cache hit still produces a span and a duration sample, just a fast one. If you cache in a before callback, your p50 drops and your downstream's real latency becomes invisible unless you label hits separately.

Advertisement

Wiring OpenTelemetry so it is not a no-op

ADK depends only on the OpenTelemetry API. Without an SDK registered as the global instance, every span and histogram goes to a no-op implementation and nothing fails. That is the most common reason teams believe ADK emits no metrics. Order matters too. ADK obtains its meter and tracer from the global instance in static initialisers, and in OpenTelemetry Java the first read of an unset global installs a no-op global (unless SDK autoconfiguration is on the classpath and enabled). A later buildAndRegisterGlobal() then throws IllegalStateException saying set has already been called. Register the SDK before any ADK class loads, at the top of main.

import io.opentelemetry.exporter.otlp.metrics.OtlpGrpcMetricExporter;
import io.opentelemetry.exporter.otlp.trace.OtlpGrpcSpanExporter;
import io.opentelemetry.sdk.OpenTelemetrySdk;
import io.opentelemetry.sdk.metrics.Aggregation;
import io.opentelemetry.sdk.metrics.InstrumentSelector;
import io.opentelemetry.sdk.metrics.SdkMeterProvider;
import io.opentelemetry.sdk.metrics.View;
import io.opentelemetry.sdk.metrics.export.PeriodicMetricReader;
import io.opentelemetry.sdk.trace.SdkTracerProvider;
import io.opentelemetry.sdk.trace.export.BatchSpanProcessor;
import java.time.Duration;
import java.util.List;

public final class Telemetry {
  public static void init() {
    // Default SDK buckets stop at 10 s; tools that call other models or slow APIs need a longer tail.
    View toolLatency = View.builder()
        .setAggregation(Aggregation.explicitBucketHistogram(List.of(
            5.0, 10.0, 25.0, 50.0, 100.0, 250.0, 500.0, 1000.0,
            2500.0, 5000.0, 10000.0, 20000.0, 30000.0, 60000.0)))
        .build();

    SdkMeterProvider meters = SdkMeterProvider.builder()
        .registerView(InstrumentSelector.builder()
            .setName("gen_ai.tool.execution.duration").build(), toolLatency)
        .registerMetricReader(PeriodicMetricReader.builder(
            OtlpGrpcMetricExporter.getDefault()).setInterval(Duration.ofSeconds(30)).build())
        .build();

    SdkTracerProvider tracers = SdkTracerProvider.builder()
        .addSpanProcessor(BatchSpanProcessor.builder(OtlpGrpcSpanExporter.getDefault()).build())
        .build();

    OpenTelemetrySdk.builder()
        .setMeterProvider(meters)
        .setTracerProvider(tracers)
        .buildAndRegisterGlobal();          // must run before any ADK class touches telemetry
  }
}

The OpenTelemetry Java agent or SDK autoconfiguration achieves the same thing through environment variables; either is fine as long as the global instance exists first. Verify it the boring way: start the service, call one tool, and look for gen_ai.tool.execution.duration in your backend before writing any dashboard.

Content capture is on by default

ADK reads the environment variable ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS, and its default is true. With the default, tool arguments and tool responses are written into span attributes. For a get_order tool that means customer names, addresses and order contents end up in your trace store, which usually has wider access and longer retention than the database they came from. It also means large responses inflate span size, and many backends truncate or drop oversized attributes.

Set ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false in every environment that handles real user data, and turn it on deliberately in development. If you need arguments for debugging in production, log a redacted, allow-listed subset from a callback instead, so the decision about what leaves the process is explicit and reviewable.

What the built-in metrics miss

The built-in error.type is set only when the tool throws. Most well-behaved tools do not throw on business failures; they return a structured result such as {"status": "error", "retryable": false} so the model can explain it. Those calls record as successes. A tool that fails validation on 40 percent of calls can look perfectly healthy on the duration histogram.

The other gaps are outcome and cause: whether the call was a cache hit, a policy denial, a not-found, a downstream timeout, or a success. None of these are knowable to the framework. They are your contract, and the right place to record them is the tool callbacks, because they see every call to every tool without touching tool bodies.

Adding outcome metrics with callbacks

import com.google.adk.agents.LlmAgent;
import io.opentelemetry.api.GlobalOpenTelemetry;
import io.opentelemetry.api.common.AttributeKey;
import io.opentelemetry.api.common.Attributes;
import io.opentelemetry.api.metrics.LongCounter;
import io.reactivex.rxjava3.core.Maybe;
import java.util.Map;
import java.util.Optional;

final class ToolOutcomes {
  private static final AttributeKey<String> TOOL = AttributeKey.stringKey("gen_ai.tool.name");
  private static final AttributeKey<String> AGENT = AttributeKey.stringKey("gen_ai.agent.name");
  private static final AttributeKey<String> OUTCOME = AttributeKey.stringKey("tool.outcome");

  // Obtain after Telemetry.init() has registered the SDK.
  private static final LongCounter CALLS = GlobalOpenTelemetry.getMeter("com.example.agents")
      .counterBuilder("tool.calls").setDescription("Tool calls by outcome").build();

  static String classify(Object response) {
    if (response instanceof Map<?, ?> m) {
      Object status = m.get("status");
      if ("denied".equals(status)) return "denied";
      if ("not_found".equals(status)) return "not_found";
      if ("error".equals(status)) {
        return Boolean.TRUE.equals(m.get("retryable")) ? "error_retryable" : "error_terminal";
      }
    }
    return "ok";                               // a small, closed set of values
  }

  static LlmAgent withOutcomes(LlmAgent.Builder builder) {
    return builder
        .afterToolCallbackSync((ctx, tool, args, toolCtx, response) -> {
          CALLS.add(1, Attributes.of(TOOL, tool.name(),
              AGENT, ctx.agent().name(), OUTCOME, classify(response)));
          return Optional.empty();             // empty = keep the tool's own response
        })
        .onToolErrorCallback((ctx, tool, args, toolCtx, error) -> {
          CALLS.add(1, Attributes.of(TOOL, tool.name(),
              AGENT, ctx.agent().name(), OUTCOME, "exception"));
          return Maybe.empty();                // empty = let the error propagate as before
        })
        .build();
  }
}

The callback interfaces used here come from the release source: the synchronous after callback receives the invocation context, the tool, the arguments, the tool context and the response, and returns an Optional map where empty means no override; the error callback receives the exception and returns a Maybe. If you short-circuit calls in a before callback, such as a cache or a policy gate, count those there with outcome cache_hit or denied, because the after callback sees the substituted value and cannot tell it apart from a real result unless you mark it. If your version lets an error callback turn an exception into a result, check whether the built-in duration still carries error.type for those calls before relying on it.

Worked example: one slow, flaky tool

The numbers here are illustrative, not measured. Suppose a support agent has three tools: get_order, search_kb and issue_refund. Users report that refunds sometimes take a minute. The traces show turns with three execute_tool issue_refund spans in a row. The duration histogram for issue_refund has a p50 of 400 ms and a p99 of 9.8 seconds, and the default buckets stop at 10 seconds, so the tail is flattened into the overflow bucket. After adding the view above, the p99 turns out to be 28 seconds.

The outcome counter shows 22 percent of issue_refund calls ending in error_retryable, with error.type empty because the tool never throws. The model is doing exactly what the tool told it: the error was retryable, so it retried, three times, inside one turn. The fix has two halves. On the downstream side, a per-tool timeout below the payment provider's own timeout turns a 28-second hang into a fast failure. On the contract side, the tool returns a terminal error after the first timeout for the same idempotency key, and the retry happens in the tool with backoff, where it is visible in metrics, instead of in the model, where it costs a full model round trip each time.

Cardinality, sampling and cost

  • Keep metric attributes closed. Tool name, agent name, outcome class and error type are bounded. User ids, session ids, call ids, argument values and model-generated strings are not; each new value creates a new time series. Put those on spans, never on metrics.
  • Watch for dynamic tool names. Toolsets that generate tool names from remote schemas, such as MCP servers, can add names at runtime. Alert on the count of distinct gen_ai.tool.name values.
  • Sample traces, not metrics. Metrics are aggregated in process and are cheap at any call volume. Traces are not; use head sampling for normal traffic and keep all traces with errors if your pipeline supports tail sampling.
  • Use the size histograms. A tool response is replayed on every later model call in the turn, so gen_ai.tool.response.size is a leading indicator of token cost. A p95 over a few tens of kilobytes deserves a summarise-and-offload design.

SLOs and alerts

Pick one latency SLO and one success SLO per tool that has a user-visible effect. Exported to Prometheus, dots in names and attribute keys become underscores and the unit is appended to the name, so confirm the exact series names in your backend before copying these queries.

# p95 tool latency per tool over 5 minutes
histogram_quantile(0.95,
  sum by (le, gen_ai_tool_name) (
    rate(gen_ai_tool_execution_duration_milliseconds_bucket[5m])))

# share of calls that did not succeed, per tool (custom counter)
sum by (gen_ai_tool_name) (rate(tool_calls_total{tool_outcome!="ok"}[5m]))
  / sum by (gen_ai_tool_name) (rate(tool_calls_total[5m]))

Alert on burn rate against the SLO, not on single spikes. Add one agent-level signal that tools alone cannot show: tool calls per turn. A sudden rise usually means the model is looping on a failing tool, and it shows up there before it shows up in cost. The ADK Java observability article covers the model-call side of the same dashboard.

Failure modes

  • No SDK registered. Everything is a silent no-op. Test for the presence of the tool histogram in CI with an in-memory exporter.
  • SDK registered too late. If ADK touched the global first, a no-op is already installed and registration fails at startup with IllegalStateException. Initialise first thing in main.
  • PII in traces. The content-capture default writes arguments and responses into spans. Turn it off for real data.
  • Healthy-looking failures. Error maps are counted as successes by the built-in metric. Add the outcome counter.
  • Flattened tails. Default buckets hide anything above 10 seconds. Register a view.
  • Orphaned spans from your own async code. ADK propagates context across its RxJava operators; if your tool hops to its own executor, wrap tasks with Context.current().wrap(...) or your child spans lose their parent.

What to do next

  1. Register the OpenTelemetry SDK globally at the top of main and confirm gen_ai.tool.execution.duration reaches your backend after one tool call.
  2. Set ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false in every environment with real user data.
  3. Add a view with buckets up to at least 60 seconds for tool latency.
  4. Add an outcome counter from after and error callbacks, with a closed set of outcome values, and count before-callback short-circuits separately.
  5. Define a latency and a success SLO for every tool with a user-visible side effect, and alert on burn rate.
  6. Chart tool calls per turn and response size per tool, and review the top three tools by each monthly.
Key takeaway: ADK Java opens an execute_tool span for every tool call and records duration and request and response size histograms tagged with tool and agent name, but only if an OpenTelemetry SDK is registered globally before ADK runs, and with argument and response content captured into spans by default. The framework cannot see business outcomes: tools that return error maps look like successes. Add a closed-set outcome counter from the tool callbacks, widen the latency buckets, keep high-cardinality data on spans, and put SLOs on the tools that change the world.