Most teams add observability to an agent the way they add it to a script: a log line here, a timer there, added after the first incident. That works for one agent. It fails at five, because each agent records different fields under different names, and the question that matters in an incident, which invocation burned the tokens, called the failing tool and gave the user a wrong answer, cannot be answered across them. Treating observability as a first-class concern means the opposite: you decide what every agent run must emit, you emit it from the runtime rather than from agent code, and you test it like any other contract.
ADK Java makes this practical. Its Runner already opens a span tree for every invocation, and its plugin interface lets one class observe every agent, model call and tool call the runner executes. This article shows what ADK emits on its own (read from the current source on the main branch), how to define a telemetry contract, how to enforce it with a single plugin, how to test it in CI, and how to turn it into service level objectives you can alert on. Class and method names below were checked against the source; if you run an older release, confirm they exist in your jar before copying the code.
What ADK already emits
Start from what you get for free, so you do not rebuild it. ADK obtains an OpenTelemetry tracer named gcp.vertex.agent and opens four kinds of span. The Runner opens invocation as the root for each call to runAsync. Each agent run opens invoke_agent <agent name>. Each model request opens call_llm, and each tool execution opens execute_tool <tool name>. The spans carry OpenTelemetry GenAI semantic-convention attributes such as gen_ai.agent.name, gen_ai.conversation.id, gen_ai.request.model, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, plus ADK attributes such as gcp.vertex.agent.invocation_id and gcp.vertex.agent.session_id.
Two defaults decide whether any of that is useful. First, ADK depends only on the OpenTelemetry API. If no SDK is registered as the global instance before ADK classes load, every span goes to a no-op implementation and nothing fails, so teams conclude ADK emits nothing. Second, the environment variable ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS defaults to true, which writes prompts, model responses, tool arguments and tool results into span attributes. In production that copies user data into your trace store. Set it to false wherever real users are served.
The tool-level spans and the histograms recorded around them are covered in detail in tool observability and metrics in ADK Java. This article is about the layer above: making the whole run observable by design.
What ADK already emits
Start from what you get for free, so you do not rebuild it. ADK obtains an OpenTelemetry tracer named gcp.vertex.agent and opens four kinds of span. The Runner opens invocation as the root for each call to runAsync. Each agent run opens invoke_agent <agent name>. Each model request opens call_llm, and each tool execution opens execute_tool <tool name>. The spans carry OpenTelemetry GenAI semantic-convention attributes such as gen_ai.agent.name, gen_ai.conversation.id, gen_ai.request.model, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, plus ADK attributes such as gcp.vertex.agent.invocation_id and gcp.vertex.agent.session_id.
Two defaults decide whether any of that is useful. First, ADK depends only on the OpenTelemetry API. If no SDK is registered as the global instance before ADK classes load, every span goes to a no-op implementation and nothing fails, so teams conclude ADK emits nothing. Second, the environment variable ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS defaults to true, which writes prompts, model responses, tool arguments and tool results into span attributes. In production that copies user data into your trace store. Set it to false wherever real users are served.
The tool-level spans and the histograms recorded around them are covered in detail in tool observability and metrics in ADK Java. This article is about the layer above: making the whole run observable by design.
Write a telemetry contract
A telemetry contract is a short document, kept in the repository, that lists what every agent invocation must emit, under which names, with which attribute values allowed. It is the agent equivalent of an API schema. Without it, dashboards break whenever someone renames a field, and cardinality grows until the metrics bill becomes the incident. A workable first version has three parts.
| Signal | Name | Attributes (bounded) | Answers |
|---|---|---|---|
| Invocation count | agent.invocations | agent, outcome | Is the agent succeeding? |
| Invocation duration | agent.invocation.duration | agent, outcome | How long do users wait? |
| Tokens | agent.tokens | agent, model, direction | What does a run cost? |
| Tool outcomes | agent.tool.outcomes | agent, tool, outcome | Which dependency is failing? |
| Span tree | ADK built-in spans | ids on attributes, not metrics | What happened in this run? |
| Log lines | JSON, one event per line | invocation_id, trace_id | Why did it happen? |
The rule that keeps this affordable: identifiers such as invocation id, session id and user id go on spans and log lines, never on metric attributes. A metric attribute with a million distinct values creates a million time series. Outcome values are a closed list, for example ok, model_error, tool_error, policy_denied and timeout, and anything else maps to other. The contract also names who owns it, usually the platform team that owns the runner, so agent authors do not invent their own fields.
One plugin enforces it for every agent
The contract is enforced in one place: a plugin. ADK Java defines a Plugin interface whose callbacks fire at every stage of a run, and a BasePlugin class whose constructor takes a name and whose callbacks default to doing nothing. The callbacks this plugin uses are beforeRunCallback, afterRunCallback, onRunErrorCallback, afterModelCallback, afterToolCallback and onToolErrorCallback. Returning an empty Maybe (or a completed Completable) means the plugin observes without changing the run.
public final class ObservabilityPlugin extends BasePlugin {
private static final AttributeKey<String> AGENT = AttributeKey.stringKey("agent");
private static final AttributeKey<String> OUTCOME = AttributeKey.stringKey("outcome");
private static final AttributeKey<String> TOOL = AttributeKey.stringKey("tool");
private static final AttributeKey<String> DIRECTION = AttributeKey.stringKey("direction");
private final LongCounter invocations;
private final DoubleHistogram duration;
private final LongCounter tokens;
private final LongCounter toolOutcomes;
private final Map<String, Long> startNanos = new ConcurrentHashMap<>();
public ObservabilityPlugin(Meter meter) {
super("observability");
invocations = meter.counterBuilder("agent.invocations").build();
duration = meter.histogramBuilder("agent.invocation.duration").setUnit("s").build();
tokens = meter.counterBuilder("agent.tokens").build();
toolOutcomes = meter.counterBuilder("agent.tool.outcomes").build();
}
@Override
public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
startNanos.put(ctx.invocationId(), System.nanoTime());
return Maybe.empty();
}
@Override
public Completable afterRunCallback(InvocationContext ctx) {
finish(ctx, "ok");
return Completable.complete();
}
@Override
public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
finish(ctx, Outcomes.classify(error));
return Completable.complete();
}
private void finish(InvocationContext ctx, String outcome) {
Long start = startNanos.remove(ctx.invocationId());
if (start == null) return; // already recorded: never count a run twice
Attributes a = Attributes.of(AGENT, ctx.agent().name(), OUTCOME, outcome);
invocations.add(1, a);
duration.record((System.nanoTime() - start) / 1e9, a);
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext cb, LlmResponse response) {
response.usageMetadata().ifPresent(u -> {
u.promptTokenCount().ifPresent(n ->
tokens.add(n, Attributes.of(AGENT, cb.agentName(), DIRECTION, "input")));
u.candidatesTokenCount().ifPresent(n ->
tokens.add(n, Attributes.of(AGENT, cb.agentName(), DIRECTION, "output")));
});
return Maybe.empty();
}
@Override
public Maybe<Map<String, Object>> afterToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext tc, Map<String, Object> result) {
toolOutcomes.add(1, Attributes.of(AGENT, tc.agentName(), TOOL, tool.name(),
OUTCOME, Outcomes.fromResult(result)));
return Maybe.empty();
}
@Override
public Maybe<Map<String, Object>> onToolErrorCallback(
BaseTool tool, Map<String, Object> args, ToolContext tc, Throwable error) {
toolOutcomes.add(1, Attributes.of(AGENT, tc.agentName(), TOOL, tool.name(),
OUTCOME, Outcomes.classify(error)));
return Maybe.empty();
}
}Outcomes is your own small class: classify maps exception types to the closed outcome list, and fromResult reads a status field from a structured tool result, because well-behaved tools report business failures as results rather than exceptions. The model name for the token counter is left out here for brevity; read it from your agent configuration rather than from the response so it stays a bounded value. Register the plugin once, on the runner:
public static void main(String[] args) {
Telemetry.init(); // registers the OpenTelemetry SDK globally, before any ADK class loads
Meter meter = GlobalOpenTelemetry.getMeter("acme.agents");
Runner runner = Runner.builder()
.agent(rootAgent)
.appName("support-desk")
.sessionService(sessionService)
.plugins(new ObservabilityPlugin(meter))
.build();
// every agent in rootAgent's tree, including sub-agents, now emits the contract
}Check one behaviour in the release you run: whether afterRunCallback also fires after onRunErrorCallback. The remove guard in finish makes the plugin correct either way, which is cheaper than depending on the answer.
Correlating logs across RxJava threads
Metrics tell you that something is wrong; logs and traces tell you which run. The join key is the invocation id, which ADK already writes on its spans, plus the trace id. Every log line an agent, tool or plugin emits should carry both.
The trap is threads. ADK Java runs on RxJava, and work hops between schedulers. A logging MDC is a thread local, so a value put there in one callback is often missing, or worse, belongs to another invocation, by the time a tool logs on a different thread. ADK's Tracing.withContext transformer carries the OpenTelemetry context across those hops for its own spans, but it does not carry your MDC. Two patterns work. Pass the context explicitly: write a tiny logging helper that takes the ToolContext or CallbackContext and adds invocationId() and sessionId() to each event. Or read the trace id from Span.current().getSpanContext() at the moment of logging, which follows the OpenTelemetry context rather than the thread. Both are better than an MDC that is right on the happy path and wrong under load.
static void logEvent(ToolContext tc, String event, Map<String, Object> fields) {
SpanContext span = Span.current().getSpanContext();
Map<String, Object> line = new LinkedHashMap<>(fields);
line.put("event", event);
line.put("invocation_id", tc.invocationId());
line.put("session_id", tc.sessionId());
line.put("agent", tc.agentName());
line.put("trace_id", span.isValid() ? span.getTraceId() : null);
LOG.info(JSON.writeValueAsString(line)); // one JSON object per line
}Keep log field names in the same contract as metric names, and decide redaction there too: the helper is the one place where a tool argument can leave the process, so it is the one place to review. If an audit trail is also required, it is a separate stream with its own retention; see audit logging in ADK Java.
Testing the contract
A contract nobody tests decays. Two tests catch almost every regression, and both run without a backend.
The first asserts the span tree. ADK exposes Tracing.setTracerForTesting(Tracer), and both the agent and tool spans and the model-call span are created from that tracer. Point it at an SDK tracer backed by an in-memory exporter, run the agent against a fake model, and assert the shape.
@Test
void everyInvocationProducesTheContractedSpanTree() {
InMemorySpanExporter spans = InMemorySpanExporter.create();
SdkTracerProvider provider = SdkTracerProvider.builder()
.addSpanProcessor(SimpleSpanProcessor.create(spans)).build();
Tracing.setTracerForTesting(provider.get("test"));
runner.runAsync("u1", sessionId, userMessage("refund order 42")).blockingSubscribe();
List<String> names = spans.getFinishedSpanItems().stream().map(SpanData::getName).toList();
assertThat(names).contains("invocation", "invoke_agent support_agent", "call_llm",
"execute_tool get_order");
assertThat(spans.getFinishedSpanItems()).allSatisfy(s ->
assertThat(s.getAttributes().get(AttributeKey.stringKey("gcp.vertex.agent.tool_call_args"))).isNull());
}The last assertion runs with ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false in the test environment and proves no content leaks into spans. The second test does the same for metrics with the SDK's InMemoryMetricReader: construct the plugin with a meter from a test SdkMeterProvider, run one successful and one failing invocation, and assert that agent.invocations has exactly one ok and one error point and that no attribute outside the contract appears. That second check is what stops someone from adding user_id to a counter next quarter.
From signals to objectives
With the contract in place, objectives are queries rather than projects. Four cover most agents:
- Success rate: share of
agent.invocationswith outcomeok, per agent, for example 99 percent over 28 days. - Latency: 95th percentile of
agent.invocation.durationunder a target the product sets, for example 8 seconds for a chat turn. - Cost: tokens per invocation, the ratio of the
agent.tokensrate to the invocation rate, with an alert when it rises by more than a set fraction week over week. - Dependency health: per-tool error share from
agent.tool.outcomes, which points at the failing system before users describe it.
Alert on burn rate, not on single bad minutes: page when the error budget is being spent fast enough to run out within hours, and open a ticket when it would run out within days. Agents are noisy because models are; burn-rate alerting absorbs that noise. Dashboards for these views are covered in metrics dashboards for ADK Java.
Worked example: a cost regression with no errors
A support agent ships a prompt change on Monday. No errors appear and the success rate holds at 99.2 percent. On Wednesday the cost alert fires: tokens per invocation rose from about 6,000 to about 14,000. Because every signal shares the contract, the investigation takes minutes.
- The tokens counter, split by agent, shows the rise is entirely in
support_agent; sub-agents are flat. - The tool outcome counter for that agent shows
search_kbcalls per invocation tripled, all with outcomeok. - Filtering traces by agent and duration, one sample shows three
execute_tool search_kbspans and fourcall_llmspans under oneinvocation: the new prompt told the model to "verify with the knowledge base", and it now searches, answers, then searches again to verify. - The log lines for that invocation id show the three queries were near duplicates.
The fix is a prompt edit plus a tool description that says results are authoritative. Nothing about this was an error, which is the point: without cost and tool-outcome signals in the contract, the regression would have surfaced on the invoice.
Failure modes
- No SDK, silent no-op. The SDK is registered after an ADK class touched the global instance, so ADK holds a no-op tracer. Register in the first line of
mainand check for spans in a smoke test. - Content in traces. The capture variable was left at its default. Set it in deployment configuration, not in code, and assert it in the contract test.
- Cardinality blow-up. Someone tags a metric with session or user id. The metric attribute test catches it before production does.
- Double counting. Error and completion callbacks both record. Key the start time by invocation id and remove it on first use.
- Plugin latency. A plugin doing network I/O inside a callback slows every run; keep callbacks in-memory and let the SDK export in the background.
Trade-offs
Runtime-level instrumentation trades flexibility for consistency: agent authors cannot add a metric without going through the contract. Turning content capture off removes the fastest debugging aid; the compensating control is redacted logging through the helper above. Head sampling of traces keeps cost down but can drop the failing run you need; tail sampling in a collector keeps error traces at the price of running that collector. ADK's built-in names follow the still-evolving GenAI semantic conventions, so recheck dashboards built on them when you upgrade ADK. For the wider picture, see ADK Java observability architecture and the ADK Java callback architecture.
What to do next
- Write the telemetry contract: signal names, allowed attribute values and the owner, in the repository.
- Register the OpenTelemetry SDK on the first line of
mainand setADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=falsein production. - Add one
ObservabilityPluginto theRunnerand delete per-agent timers and counters. - Give every log line the invocation id, session id and trace id through an explicit helper, not an MDC.
- Add the span-tree test and the metric-attribute test to CI.
- Define success, latency, cost and tool-health objectives and alert on burn rate.
- After each ADK upgrade, rerun the contract tests and check built-in span and attribute names.