ADK for Java already records some metrics. On recent releases (it is present in v1.11.0), the class com.google.adk.telemetry.Metrics records seven OpenTelemetry histograms: agent invocation duration, tool execution duration, agent request and response size, workflow steps, and tool request and response size. Sooner or later you need numbers it does not have, such as model-call latency per agent, how often a cache short-circuits the model, runs by outcome, or a business event like an escalation to a human.
This page is about building that extension correctly. It covers what you can and cannot change in the built-ins, the exact dispatch rules of the plugin system, a complete metrics plugin, the correlation and streaming traps that make naive versions wrong, histogram Views, and a test that pins the behaviour. Which metrics belong on a dashboard is covered in Metrics + Dashboards, and tool outcome metrics in Tool Observability + Metrics.
What the built-ins give you, and what you can change
What you cannot change. Metrics is a final class with a private constructor and static methods. It holds its instruments in a static holder that it creates with GlobalOpenTelemetry.getMeter("gcp.vertex.agent"). You cannot subclass it or add instruments to it. The only public hook, setMeterForTesting(Meter), is documented as being for tests. The attribute keys it uses are gen_ai.agent.name, gen_ai.tool.name and, on failure, error.type, set to the exception's simple class name.
What you can change. Everything else lives in the OpenTelemetry SDK that you configure. Views can change the built-in histograms' buckets, rename them or drop attributes, because they select by meter name and instrument name. New measurements come from your own Meter, fed by a plugin that sees runtime events. Extending runtime metrics in ADK Java therefore means two things: configuring the SDK, and writing a plugin. There is no metrics registry inside ADK to inject or extend. Treat any example showing one as invented.
One wiring detail matters. The static holder calls GlobalOpenTelemetry when the class initializes. If no global SDK was ever registered, the built-ins get a no-op meter and record nothing, while your plugin, using a meter you passed in directly, works, and the dashboard is half populated. If you register too late, GlobalOpenTelemetry.get() has already installed a no-op global and the later set throws IllegalStateException at startup. Register the global SDK in the first lines of startup.
The plugin contract a metrics extension must respect
A plugin extends com.google.adk.plugins.BasePlugin and overrides default methods of the Plugin interface. Those cover a run (onUserMessageCallback, beforeRunCallback, onEventCallback, afterRunCallback, onRunErrorCallback), an agent (beforeAgentCallback, afterAgentCallback), a model call (beforeModelCallback, afterModelCallback, onModelErrorCallback) and a tool call (beforeToolCallback, afterToolCallback, onToolErrorCallback). The behaviour of PluginManager and BaseLlmFlow gives four rules a metrics plugin must respect.
- Order is the list order. Plugins run in the order passed to
Runner.builder().plugins(...). - First non-empty result wins. Most callbacks return a
Maybe. The manager stops at the first plugin that emits a value, and the plugins after it are not called. A cache plugin that answers inbeforeModelCallbackhides that callback from every later plugin. Plugin results are also checked before the agent's own callbacks, which run only when every plugin returned empty. A metrics plugin must therefore always returnMaybe.empty()and should be first in the list. - Errors propagate. The manager logs a plugin's error and passes it on. It does not swallow it. An exception thrown in your counter code fails the user's request. Wrap every callback body.
- A short-circuit skips the after-callback. If any before-model callback returns a response, the model is not called and
afterModelCallbackdoes not fire for that call. Any state you stored in the before-callback stays behind.
A complete metrics plugin
The plugin below records model-call latency and outcomes, counts model calls that start, and counts tool calls and runs by outcome. Token counters are left out because the dashboards page already shows them. They would go in afterModelCallback the same way.
public final class RuntimeMetricsPlugin extends BasePlugin {
private static final Logger log = LoggerFactory.getLogger(RuntimeMetricsPlugin.class);
private static final AttributeKey<String> AGENT = AttributeKey.stringKey("gen_ai.agent.name");
private static final AttributeKey<String> TOOL = AttributeKey.stringKey("gen_ai.tool.name");
private static final AttributeKey<String> ERROR = AttributeKey.stringKey("error.type");
private static final AttributeKey<String> OUTCOME = AttributeKey.stringKey("outcome");
private final DoubleHistogram modelDuration;
private final LongCounter modelStarts, modelCalls, toolCalls, runs;
private final Map<String, Long> starts = new ConcurrentHashMap<>();
public RuntimeMetricsPlugin(Meter meter) {
super("runtime_metrics");
modelDuration = meter.histogramBuilder("app.agent.model_call.duration").setUnit("ms")
.setExplicitBucketBoundariesAdvice(List.of(
100.0, 250.0, 500.0, 1000.0, 2500.0, 5000.0, 10000.0, 30000.0, 60000.0, 120000.0))
.build();
modelStarts = meter.counterBuilder("app.agent.model_call.starts").build();
modelCalls = meter.counterBuilder("app.agent.model_calls").build();
toolCalls = meter.counterBuilder("app.agent.tool_calls").build();
runs = meter.counterBuilder("app.agent.runs").build();
}
private static String key(CallbackContext ctx) {
return ctx.invocationId() + "/" + ctx.eventId();
}
@Override
public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder req) {
safely(() -> {
starts.put(key(ctx), System.nanoTime());
modelStarts.add(1, Attributes.of(AGENT, ctx.agentName()));
});
return Maybe.empty();
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
safely(() -> {
if (resp.partial().orElse(false)) {
return; // streaming chunk: wait for the final response
}
finishModelCall(ctx, Attributes.of(AGENT, ctx.agentName(), OUTCOME, "ok"));
});
return Maybe.empty();
}
@Override
public Maybe<LlmResponse> onModelErrorCallback(
CallbackContext ctx, LlmRequest.Builder req, Throwable error) {
safely(() -> finishModelCall(ctx, Attributes.of(
AGENT, ctx.agentName(), OUTCOME, "error", ERROR, error.getClass().getSimpleName())));
return Maybe.empty();
}
private void finishModelCall(CallbackContext ctx, Attributes attrs) {
modelCalls.add(1, attrs);
Long t0 = starts.remove(key(ctx));
if (t0 != null) {
modelDuration.record((System.nanoTime() - t0) / 1e6, attrs);
}
}
@Override
public Maybe<Map<String, Object>> afterToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx, Map<String, Object> result) {
safely(() -> toolCalls.add(1, Attributes.of(
TOOL, tool.name(), AGENT, ctx.agentName(), OUTCOME, "ok")));
return Maybe.empty();
}
@Override
public Maybe<Map<String, Object>> onToolErrorCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx, Throwable error) {
safely(() -> toolCalls.add(1, Attributes.of(
TOOL, tool.name(), AGENT, ctx.agentName(), OUTCOME, "exception")));
return Maybe.empty();
}
@Override
public Completable afterRunCallback(InvocationContext inv) {
safely(() -> {
runs.add(1, Attributes.of(AGENT, inv.agent().name(), OUTCOME, "ok"));
forget(inv);
});
return Completable.complete();
}
@Override
public Completable onRunErrorCallback(InvocationContext inv, Throwable error) {
safely(() -> {
runs.add(1, Attributes.of(AGENT, inv.agent().name(), OUTCOME, "error",
ERROR, error.getClass().getSimpleName()));
forget(inv);
});
return Completable.complete();
}
private void forget(InvocationContext inv) { // short-circuited calls never reach an after-callback
String prefix = inv.invocationId() + "/";
starts.keySet().removeIf(k -> k.startsWith(prefix));
}
private static void safely(Runnable body) {
try {
body.run();
} catch (RuntimeException e) {
log.warn("runtime metrics callback failed; request continues", e);
}
}
}Register it first, before any plugin that can answer on its own:
Runner runner = Runner.builder()
.agent(rootAgent)
.appName("support")
.sessionService(new InMemorySessionService())
.artifactService(new InMemoryArtifactService())
.plugins(new RuntimeMetricsPlugin(openTelemetry.getMeter("com.example.adk.metrics")),
new ResponseCachePlugin()) // your own plugin, listed after metrics
.build();
Correlation, short-circuits and streaming
Correlating before and after. Callbacks for one model call can run on different threads, as explained in the runtime thread model, so a thread-local timer is wrong. A key made of the agent name alone is also wrong, because parallel branches and several model calls per turn would overwrite each other. In the current flow code, the before, after and error callbacks for one model call are all built from the same event, so invocationId() + "/" + eventId() identifies the call. This is an implementation detail, not a documented promise, so the test in the next section checks it on every upgrade.
Short-circuits leak state. When the cache plugin answers, the before-callback has stored a start time that no after-callback will remove. Without forget, the map grows for as long as the process lives, and every cache hit is a small memory leak. The cleanup in afterRunCallback and onRunErrorCallback handles it. If you are not sure both fire on every path in your version, also cap the map: clear it and count an overflow metric when it passes a few thousand entries. The gap between app.agent.model_call.starts and app.agent.model_calls is then a useful number in its own right, because it measures short-circuited calls.
Streaming. The flow calls afterModelCallback once per response the model yields. In SSE streaming mode that means once per chunk, so counting every call inflates the counts. The plugin skips responses marked partial() and counts the final one. Confirm with a test against your model integration that a non-partial response really arrives at the end of a stream. If you also want time to first token, record it on the first partial response for a key, and keep the start time until the final response arrives.
Cost on the hot path. These callbacks run inline, so keep them to counter increments and a map operation. Never call a network exporter from a callback. The SDK's periodic reader batches and exports on its own thread.
Histogram buckets and Views
OpenTelemetry's default explicit histogram buckets end at 10,000. With millisecond units, that means 10 seconds. A multi-step agent run often takes longer, so on default buckets every slow run falls into the overflow bucket, and p95 or p99 can only be read as 'more than 10 seconds'. For your own instruments, the builder advice in the plugin sets the boundaries. For ADK's built-in histograms, which you cannot edit, register a View that selects them by meter name:
SdkMeterProvider meterProvider = SdkMeterProvider.builder()
.registerMetricReader(PeriodicMetricReader.builder(OtlpGrpcMetricExporter.getDefault()).build())
.registerView(
InstrumentSelector.builder()
.setMeterName("gcp.vertex.agent")
.setName("gen_ai.agent.invocation.duration")
.build(),
View.builder()
.setAggregation(Aggregation.explicitBucketHistogram(List.of(
250.0, 500.0, 1000.0, 2500.0, 5000.0, 10000.0, 20000.0, 40000.0, 80000.0, 160000.0)))
.build())
.build();
OpenTelemetry openTelemetry = OpenTelemetrySdk.builder()
.setMeterProvider(meterProvider)
.buildAndRegisterGlobal(); // before any agent runs; see the wiring note aboveViews are also where you control cardinality without touching code. A View can keep only an allowed set of attribute keys on an instrument, so a careless user_id attribute added later is dropped rather than multiplying your series. Do not put session IDs, user IDs or raw tool arguments on any metric. They belong on spans, where per-request detail is cheap.
Testing the extension
Metrics code is rarely tested, which is why it is often wrong. Use the SDK's in-memory reader and drive the plugin through a real Runner with a stub model:
@Test
void countsOneModelCallPerStreamAndForgetsShortCircuits() {
InMemoryMetricReader reader = InMemoryMetricReader.create();
SdkMeterProvider provider = SdkMeterProvider.builder().registerMetricReader(reader).build();
RuntimeMetricsPlugin plugin = new RuntimeMetricsPlugin(provider.get("test"));
// Stub BaseLlm: yields two partial chunks, then one final response.
// Run one turn through Runner with plugins(plugin), then a second turn
// where a cache plugin listed after it answers in beforeModelCallback.
Map<String, Long> totals = reader.collectAllMetrics().stream()
.filter(m -> m.getName().startsWith("app.agent."))
.filter(m -> !m.getLongSumData().getPoints().isEmpty())
.collect(toMap(MetricData::getName,
m -> m.getLongSumData().getPoints().stream().mapToLong(LongPointData::getValue).sum()));
assertEquals(2L, totals.get("app.agent.model_call.starts"));
assertEquals(1L, totals.get("app.agent.model_calls")); // chunks not counted, cache hit not completed
assertEquals(0, plugin.pendingStartsForTesting()); // forget() ran at the end of each run
}The last assertion needs a small package-private accessor returning starts.size(). This one test pins the three behaviours the plugin depends on: shared event IDs, final-response streaming and cleanup at the end of a run. When an ADK upgrade changes any of them, the test fails in CI rather than the dashboard drifting quietly. To test code that reads ADK's built-ins, Metrics.setMeterForTesting points them at a test meter.
Worked example: a cached support agent
Take a support agent with one tool, lookup_order, behind a response cache. A user asks about an order. The first model call takes 1.2 seconds and requests the tool, the tool returns in 300 ms, and a second model call writes the answer in 1.6 seconds. A second user asks the same question, and the cache answers before the model is called. After both runs, a correct plugin reports the following:
| Instrument | Value | Why |
|---|---|---|
app.agent.model_call.starts | 3 | two calls in run 1, one before the cache hit in run 2 |
app.agent.model_calls outcome=ok | 2 | the cache hit never completes a model call |
app.agent.model_call.duration | 2 samples: about 1,200 and 1,600 ms | only completed calls are timed |
app.agent.tool_calls lookup_order ok | 1 | the tool ran once |
app.agent.runs outcome=ok | 2 | both runs finished |
| pending start entries | 0 | run 2's entry removed by forget |
Now look at the common naive versions. A plugin listed after the cache shows 2 starts and a cache that seems to do nothing. A plugin without the partial() check, under SSE, reports many model calls per answer. A plugin without forget reports the same numbers but keeps one stale entry per cache hit, forever. Each of these produces a believable dashboard, which is what makes it dangerous.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Built-in histograms empty, plugin metrics fine | No global SDK registered; plugin given a Meter directly | Register the global SDK at startup |
Startup fails: GlobalOpenTelemetry.set has already been called | Global read before it was set | Register before any ADK class runs |
| User requests fail with a metrics stack trace | Exception escaped a callback | Wrap every body; return empty |
| Cache hit rate looks like zero | Metrics plugin listed after the cache plugin | Put metrics first |
| Heap grows slowly under cache load | Start times stored for short-circuited calls | Clean up per invocation; cap the map |
| Model calls about ten times answers | Counting partial streaming chunks | Skip partial() responses |
| p99 stuck at a round number | Default buckets end at 10,000 ms | View or bucket advice |
| Series count explodes | IDs used as attributes | Attribute-filter View; IDs on spans |
For how these metrics connect to the trace tree and the evaluation side of observability, see ADK Java observability architecture.
What to do next
- Check that
com.google.adk.telemetry.Metricsexists in your ADK version, and list which of its seven histograms reach your backend. - Move SDK registration to the first lines of startup and confirm the built-ins are no longer empty.
- Write the plugin with
safelywrappers andMaybe.empty()returns, and list it first inplugins(...). - Add the in-memory reader test covering streaming, a short-circuit and cleanup, and run it on every ADK upgrade.
- Register a View with longer buckets for
gen_ai.agent.invocation.duration, and an attribute filter for every custom instrument. - Graph starts minus completed model calls, and check that it matches your cache's own hit counter.