An agent in production fails in ways a normal service does not. It can stay up and fast while quietly calling the model eight times instead of three, doubling cost. It can keep a perfect error rate while a tool returns nonsense. A dashboard for an ADK Java agent therefore has to show more than requests and latency: it needs steps per invocation, tokens, model behaviour and cost, laid out so that an on-call engineer can go from "something is wrong" to "this agent, this model, since this deploy" in a minute.

This article lists exactly what ADK Java already measures, shows how to make sure those measurements are exported rather than dropped, fills the gaps for model latency, tokens and cost, and then builds the dashboard row by row with queries and alerts. Framework details were checked against the google/adk-java main branch on 2026-10-04. Tool-level metrics get a summary here; their full treatment, including tool SLOs, is in Tool Observability + Metrics in ADK Java.

Five questions, five rows

Design the dashboard from the questions it must answer, not from the metrics that happen to exist. Five questions cover almost every incident:

  1. Is it serving? Invocations per second and the share that fail, per agent.
  2. Is it fast enough? Invocation latency percentiles, and how many steps each invocation takes, since step count drives latency.
  3. Is the model behaving? Model call rate, model latency, tokens in and out, and how responses finish (normal stop, length limit, safety block).
  4. Are the tools healthy? The slowest and most failing tools, as a summary with a link to a tool dashboard.
  5. What does it cost? Spend per hour and per invocation, by agent and model.

Each question becomes one row. Rows are ordered so that the first one you look at tells you whether to keep looking.

What ADK Java records out of the box

ADK Java has a com.google.adk.telemetry.Metrics class that obtains a meter named gcp.vertex.agent from the global OpenTelemetry instance and records seven histograms:

InstrumentUnitAttributes
gen_ai.agent.invocation.durationmsgen_ai.agent.name, error.type when it failed
gen_ai.agent.workflow.steps1gen_ai.agent.name
gen_ai.agent.request.sizeBygen_ai.agent.name
gen_ai.agent.response.sizeBygen_ai.agent.name
gen_ai.tool.execution.durationmsagent name, gen_ai.tool.name, error.type on failure
gen_ai.tool.request.sizeByagent name, tool name
gen_ai.tool.response.sizeByagent name, tool name

The workflow steps histogram is derived from the events an invocation produces, which makes it the best built-in signal for loops and over-planning; read the method in your ADK version before treating one step as exactly one model call. Every histogram also gives you a count series, so the invocation duration histogram doubles as the request counter and, split by error.type, as the error counter.

What is not there matters as much. There is no metric for model call latency, tokens or cost. Tokens do appear as span attributes: the call_llm span carries gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.usage.cache_read.input_tokens and gen_ai.usage.reasoning.output_tokens. Spans are sampled and expensive to aggregate, so rows 3 and 5 need their own metrics. If you run an older release, check that the Metrics class exists in your jar before building on it.

Wiring the SDK so metrics are not a no-op

Because ADK reads the meter from GlobalOpenTelemetry, nothing is exported unless an SDK is registered globally. Without one, every histogram records into a no-op and the dashboard is empty, with no error anywhere. Register the SDK at startup, before you construct agents or the runner. The OpenTelemetry Java agent can do this for you instead; either way, verify by looking for the metric in your backend, not by reading config.

// Agent invocations routinely exceed 10 s, the top of the SDK's default
// millisecond buckets, so give them explicit boundaries.
View invocationBuckets = View.builder()
    .setAggregation(Aggregation.explicitBucketHistogram(List.of(
        250.0, 500.0, 1000.0, 2000.0, 4000.0, 8000.0, 15000.0,
        30000.0, 60000.0, 120000.0)))
    .build();

SdkMeterProvider meterProvider = SdkMeterProvider.builder()
    .setResource(Resource.getDefault().merge(Resource.create(
        Attributes.of(AttributeKey.stringKey("service.name"), "support-agent"))))
    .registerView(InstrumentSelector.builder()
        .setName("gen_ai.agent.invocation.duration").build(), invocationBuckets)
    .registerReader(PeriodicMetricReader.builder(
        OtlpGrpcMetricExporter.builder().build())
        .setInterval(Duration.ofSeconds(30)).build())
    .build();

OpenTelemetrySdk.builder()
    .setMeterProvider(meterProvider)
    .setTracerProvider(tracerProvider)      // spans for call_llm, built elsewhere
    .buildAndRegisterGlobal();

The bucket view is the change teams most often miss. The SDK's default boundaries stop at 10,000 ms, so every slow invocation lands in the overflow bucket and your p95 and p99 flatten at a meaningless ceiling exactly when things go wrong. Choose boundaries that bracket your latency objective, and add a second view for the steps histogram if your agents can take more than a handful of steps.

Filling the gaps: model latency, tokens and cost

Model latency is easiest to derive from spans in the OpenTelemetry Collector. The spanmetrics connector turns every span into a call counter and a duration histogram keyed by span name and chosen attributes, so filtering on call_llm gives model latency per model without touching code. It must see spans before any tail sampling, or the counts shrink with the sample rate. Output metric names differ between Collector versions, so look them up in your backend.

Tokens should be counters you record yourself, because you want exact sums over all traffic. An after-model callback sees every LlmResponse and its usageMetadata():

public final class ModelMetrics implements Callbacks.AfterModelCallbackSync {
  private static final AttributeKey<String> AGENT = AttributeKey.stringKey("gen_ai.agent.name");
  private static final AttributeKey<String> MODEL = AttributeKey.stringKey("model");
  private static final AttributeKey<String> TYPE = AttributeKey.stringKey("type");
  private static final AttributeKey<String> FINISH = AttributeKey.stringKey("finish_reason");
  private final LongCounter tokens;
  private final LongCounter responses;

  public ModelMetrics(Meter meter) {
    tokens = meter.counterBuilder("agent.model.tokens").setUnit("{token}").build();
    responses = meter.counterBuilder("agent.model.responses").build();
  }

  @Override
  public Optional<LlmResponse> call(CallbackContext ctx, LlmResponse resp) {
    if (resp.partial().orElse(false)) {
      return Optional.empty();              // count only complete responses
    }
    String agent = ctx.agentName();
    String model = resp.modelVersion().orElse("unknown");
    resp.usageMetadata().ifPresent(u -> {
      add(agent, model, "input", u.promptTokenCount());
      add(agent, model, "output", u.candidatesTokenCount());
      add(agent, model, "cached", u.cachedContentTokenCount());
      add(agent, model, "thoughts", u.thoughtsTokenCount());
    });
    String finish = resp.finishReason().map(Object::toString).orElse("none");
    responses.add(1, Attributes.of(AGENT, agent, MODEL, model, FINISH, finish));
    return Optional.empty();                // never replace the response
  }

  private void add(String agent, String model, String type, Optional<Integer> n) {
    n.ifPresent(v -> tokens.add(v, Attributes.of(AGENT, agent, MODEL, model, TYPE, type)));
  }
}
// LlmAgent.builder()...afterModelCallbackSync(new ModelMetrics(meter)).build();

Two cautions. When streaming, a call can produce several responses; skipping partial ones avoids double counting, but confirm on your model and ADK version that usage arrives on the final response. And the callback must return an empty Optional, because a non-empty value replaces the model's response.

Cost is tokens times price. Keep prices out of the agent: publish the token counters and multiply by a small price table in a recording rule or the dashboard, so a price change is a config edit rather than a deploy, and history can be recomputed. Label the result an estimate; the invoice is the source of truth.

Metric names and queries

Before writing a query, open your backend's metric browser and copy the real names. With an OTLP-to-Prometheus path, dots usually become underscores, the unit is often appended as a suffix (milliseconds for ms), counters gain _total and histograms appear as _bucket, _sum and _count. The exact result depends on the exporter version and its translation settings, so the queries below use the most common form and may need renaming.

# Row 1: invocations/s and error ratio per agent
sum by (gen_ai_agent_name) (rate(gen_ai_agent_invocation_duration_milliseconds_count[5m]))

sum by (gen_ai_agent_name) (rate(gen_ai_agent_invocation_duration_milliseconds_count{error_type!=""}[5m]))
  / sum by (gen_ai_agent_name) (rate(gen_ai_agent_invocation_duration_milliseconds_count[5m]))

# Row 2: p95 latency and p95 steps
histogram_quantile(0.95, sum by (le, gen_ai_agent_name)
  (rate(gen_ai_agent_invocation_duration_milliseconds_bucket[5m])))
histogram_quantile(0.95, sum by (le, gen_ai_agent_name)
  (rate(gen_ai_agent_workflow_steps_bucket[5m])))

# Row 3: tokens per invocation, by type
sum by (type) (rate(agent_model_tokens_total[5m]))
  / scalar(sum(rate(gen_ai_agent_invocation_duration_milliseconds_count[5m])))

The dashboard, row by row

Where each dashboard row gets its dataADK Metrics class7 built-in histogramscall_llm spanstoken usage attributesYour callbackstokens, finish reasonOTel Collectorspanmetrics connectorOTLP metricsperiodic readertraces1 Traffic and errors2 Invocation latency, steps3 Model calls and tokens4 Tools (summary)5 Costmodel latencyBuilt-in histograms cover rows 1, 2 and 4. Rows 3 and 5 need span-derived metrics or callbacks.
Three sources feed five rows. The built-in histograms arrive over OTLP; model latency is derived from call_llm spans in the Collector; token and finish-reason counters come from an after-model callback.
RowPanelsRead it as
1 Traffic and errorsInvocations/s, error ratio, errors by error.typeIs anything on fire, and which agent
2 Latency and stepsp50/p95/p99 invocation latency, p95 steps, heatmapDid invocations get slower because each step slowed or because there are more steps
3 ModelCalls/s and p95 latency per model, tokens per invocation by type, finish reasonsIs the model slower, chattier or being cut off
4 ToolsTop five tools by p95 and by error ratioWhich dependency to open next
5 CostEstimated spend per hour and per invocation, by agent and modelIs this incident also a budget incident

Put a deploy annotation on every panel; most agent regressions start with a prompt, model or tool change. Keep per-session detail out of the dashboard and link from each panel to traces filtered by agent and time, where the prompt and completion spans show what the model actually saw.

Worked example: a regression only the dashboard catches

Suppose a support agent averages 3 steps and 6,000 input tokens per invocation, with p95 latency of 7 s. A prompt change ships that tells the agent to "double-check order details". Error ratio stays at 0.4 percent and the service looks healthy by request metrics alone.

Row 2 shows the change. p95 steps rise from 4 to 9 and p95 latency from 7 s to 16 s. The question the row is designed to answer, slower steps or more steps, has a clear answer: the model latency panel in row 3 is flat, so each step is as fast as before and there are simply more of them. Row 3 also shows input tokens per invocation rising from 6,000 to about 15,000, because every extra step resends the growing conversation. At a hypothetical price of $1 per million input tokens and 200,000 invocations a day, input spend goes from about $1,200 to about $3,000 a day, which row 5 shows within minutes rather than at month end.

Row 4 completes the picture: the order lookup tool's call rate has tripled while its latency and errors are unchanged. The fix is in the prompt, not the tool. Without the steps histogram and token counters, this incident would have shown up first in the cloud bill.

Alerts worth having

  • Error ratio burn rate. Alert when the invocation error ratio consumes the error budget fast (for example 14 times the budget rate over 1 hour and 5 minutes), per agent.
  • Latency objective. Alert when the share of invocations above your latency threshold exceeds the objective for 15 minutes. Use bucket counts rather than a quantile, which averages badly across instances.
  • Runaway steps. Alert when p99 steps approaches your maxLlmCalls limit, a sign that invocations are about to be cut off. ADK Java Error Recovery covers what happens when they are.
  • Tokens per invocation. Alert on a sustained jump against the same hour last week. It catches prompt regressions and retrieval bloat that no error metric sees.
  • Telemetry absent. Alert when the invocation count series disappears while the service is up. That is the no-op SDK failure, and without this rule it looks like zero traffic.

Cardinality

Every unique label combination is a series, and every histogram multiplies that by its bucket count plus two. Agent name, model, tool name, token type and finish reason are bounded and safe. Session ids, user ids, invocation ids, prompts or raw error messages must never become labels; one per-user label can turn a few hundred series into millions. The built-in error.type uses the exception's class name, which is bounded as long as you do not generate exception classes dynamically. For per-user cost, aggregate in your billing pipeline from traces or logs, not from metric labels.

Failure modes

  • Empty dashboard, healthy service. No SDK registered globally, or registered after startup code already ran. Check for the metric in the backend after a test invocation.
  • p99 stuck at 10 s. Default buckets. Add the view.
  • Rates jump by a factor of the replica count or reset oddly. Delta and cumulative temporality mismatched with what the backend expects. Prometheus-style backends need cumulative.
  • Model latency undercounts. Spanmetrics placed after a sampler.
  • Tokens doubled. Partial streaming responses counted.
  • Zero errors while users complain. error.type is set only when an invocation or tool throws. A tool that returns an error map counts as success; use the outcome counters from the tool observability article.
  • Two agents, one name. Metrics are keyed by agent name, so reusing a name across sub-agents merges their series. Give every agent a unique name.

What to do next

  1. Register the OpenTelemetry SDK globally at startup, run one invocation and confirm all seven built-in histograms reach your backend.
  2. Add explicit bucket views for invocation duration and workflow steps that bracket your objectives.
  3. Add the after-model callback for tokens and finish reasons, and verify no double counting with streaming on.
  4. Enable the spanmetrics connector ahead of any sampling and filter it to call_llm for model latency.
  5. Copy the real metric names from your backend and build the five rows, with deploy annotations.
  6. Create the five alerts, including the telemetry-absent rule.
  7. Review label sets for anything unbounded before the first production deploy.
  8. Continue with ADK Java observability architecture to connect the dashboard with traces and evaluations.
Key takeaway: ADK Java already records invocation, step, size and tool histograms under the gcp.vertex.agent meter, but only if an OpenTelemetry SDK is registered globally, and its default buckets hide slow invocations. Add model latency from call_llm spans, token and finish-reason counters from an after-model callback, and cost as a recording rule. Lay the dashboard out as five rows: traffic and errors, latency and steps, model, tools, cost. Alert on burn rate, latency, runaway steps, tokens per invocation and missing telemetry, and keep every label bounded.