A metrics dashboard tells you that an agent got slower or more expensive. It rarely tells you why, because the why lives inside individual conversations: the sub-agent that keeps handing work back, the tool the model calls three times with the same arguments, the answer that never arrives because the run ended on a tool error. An agent dashboard worth having connects the two levels. Fleet panels show rates and distributions by agent version; one click opens the timeline of a single invocation that contributed to the bad number.
This article builds that for ADK Java. A small plugin turns the event stream of every invocation into one summary row and a handful of step rows. Five behaviour panels are plain SQL over those tables, and a timeline view renders one invocation from its step rows. The operational metrics (latency histograms, tokens, cost, alerts) are covered in the metrics and dashboards article, and product analytics through the BigQuery analytics plugin in agent telemetry events for analytics. This page fills the gap between them: behaviour you can only see by looking at the sequence of events.
Questions before panels
Start from the questions a dashboard must answer when someone says the agent got worse.
| Question | Signal | Where it comes from |
|---|---|---|
| Do runs end with an answer? | Share of invocations with a final response and no error | summary row |
| Is the agent looping? | Most repeats of an identical tool call within one run | summary row |
| Is it taking more steps? | Distribution of events and tool calls per run, by version | summary row |
| Who hands work to whom? | Counts of agent-to-agent transfers | step rows |
| Where does the time go? | Wall time split into model wait and tool wait | step rows |
None of these need conversation text, which matters for privacy and for cost: counts and identifiers are small, cheap to keep for months and safe to show to a wide audience. The text stays in the session store, behind its own access rules and retention, and the dashboard links to it by session id.
Architecture
ADK Java plugins are registered once on the runner and see every invocation, which is exactly the scope a dashboard needs; per-agent callbacks would have to be attached to each agent and would miss runner-level events. The plugin uses four hooks from the Plugin interface in v1.11.0: beforeRunCallback to open an accumulator, onEventCallback to fold in each event, and afterRunCallback and onRunErrorCallback to close it. See the plugin architecture article for how the plugin manager orders and runs these hooks.
The recorder plugin
public final class InvocationRecorder extends BasePlugin {
private final ConcurrentHashMap<String, Acc> live = new ConcurrentHashMap<>();
private final SummarySink sink; // bounded queue + background writer
private final String agentVersion;
public InvocationRecorder(SummarySink sink, String agentVersion) {
super("invocation_recorder");
this.sink = sink;
this.agentVersion = agentVersion;
}
@Override
public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
live.put(ctx.invocationId(), new Acc(ctx, agentVersion, System.currentTimeMillis()));
return Maybe.empty(); // never short-circuit the run
}
@Override
public Maybe<Event> onEventCallback(InvocationContext ctx, Event e) {
Acc acc = live.get(ctx.invocationId());
if (acc != null && !e.partial().orElse(false)) {
acc.add(e); // streaming chunks are skipped
}
return Maybe.empty(); // never replace the event
}
@Override
public Completable afterRunCallback(InvocationContext ctx) {
return Completable.fromAction(() -> emit(ctx.invocationId(), null));
}
@Override
public Completable onRunErrorCallback(InvocationContext ctx, Throwable err) {
return Completable.fromAction(() -> emit(ctx.invocationId(), err.getClass().getSimpleName()));
}
private void emit(String invocationId, String failure) {
Acc acc = live.remove(invocationId); // remove() makes a double close harmless
if (acc != null) sink.offer(acc.summary(failure), acc.steps());
}
}The accumulator does the counting. It records one step row per event and keeps running totals for the summary.
synchronized void add(Event e) {
long ts = e.timestamp(); // epoch milliseconds
String kind = !e.functionCalls().isEmpty() ? "tool_call"
: !e.functionResponses().isEmpty() ? "tool_result"
: e.actions().transferToAgent().isPresent() ? "transfer"
: "message";
steps.add(new Step(steps.size(), ts, ts - lastTs, e.author(), kind, toolName(e)));
lastTs = ts;
if (!e.author().equals("user") && !e.author().equals(lastAuthor)) {
agentPath.add(e.author());
lastAuthor = e.author();
}
for (FunctionCall fc : e.functionCalls()) {
toolCalls++;
String key = fc.name().orElse("?") + new TreeMap<>(fc.args().orElse(Map.of()));
maxRepeat = Math.max(maxRepeat, repeats.merge(key, 1, Integer::sum));
}
e.usageMetadata().ifPresent(u -> {
modelTurns++;
inputTokens += u.promptTokenCount().orElse(0);
outputTokens += u.candidatesTokenCount().orElse(0);
});
if (e.finalResponse()) finalResponse = true;
e.errorCode().ifPresent(code -> errorCode = code.toString());
}The repeat key sorts arguments with a TreeMap so the same call with keys in a different order still counts as a repeat. The summary row carries invocation id, session id, app, a hashed user id, agent version, start and end time, event count, model turns, tool calls, maximum repeat, the agent path, whether a final response arrived, error code, failure class and token totals. Register the plugin with Runner.builder().plugins(recorder) and confirm rows appear before building any panel on them.
Five behaviour panels in SQL
Every panel is a query over two tables. The SQL below is BigQuery dialect; adjust the quantile and window syntax for Postgres or ClickHouse.
-- Panel 1: outcomes and loops by version, last 24 hours
SELECT agent_version,
COUNT(*) AS runs,
AVG(CASE WHEN final_response AND failure IS NULL THEN 1.0 ELSE 0 END) AS answered,
AVG(CASE WHEN max_repeat >= 3 THEN 1.0 ELSE 0 END) AS looping,
APPROX_QUANTILES(tool_calls, 100)[OFFSET(95)] AS p95_tool_calls
FROM invocation_summary
WHERE start_ms >= @since
GROUP BY agent_version;
-- Panel 4: transfer graph edges (feed a Sankey or a table)
SELECT prev_author AS source, author AS target, COUNT(*) AS n
FROM (SELECT author,
LAG(author) OVER (PARTITION BY invocation_id ORDER BY seq) AS prev_author
FROM invocation_step WHERE ts_ms >= @since)
WHERE prev_author IS NOT NULL AND prev_author <> author
GROUP BY source, target ORDER BY n DESC;
-- Panel 5: where the time goes
SELECT CASE WHEN prev_kind = 'tool_call' THEN 'tool wait' ELSE 'model wait' END AS bucket,
SUM(gap_ms) / 1000.0 AS seconds
FROM (SELECT gap_ms, LAG(kind) OVER (PARTITION BY invocation_id ORDER BY seq) AS prev_kind
FROM invocation_step WHERE ts_ms >= @since)
WHERE prev_kind IS NOT NULL
GROUP BY bucket;Panel 2 is a heatmap of tool calls per run against time, split by version; a new version whose mass shifts upward is doing more work per question even if latency alarms have not fired yet. Panel 3 lists the top tool bigrams (pairs of consecutive tool names), which reveals habits such as a search always followed by the same lookup. Every panel row should link to a filtered list of invocation ids, and every id to its timeline.
The time split is an approximation worth stating on the panel itself. The gap before a tool result is attributed to the tool, and every other gap to the model, which also absorbs framework overhead and callback time. For exact model latency use the spans described in the metrics article.
The session timeline
The timeline view is the payoff. Load the step rows for one invocation, draw one lane per author and two for waiting, and place each step at its timestamp. Under the chart, list the steps with tool names and the gap before each. Add two links: one to the session in your session store for the conversation text (for people allowed to read it) and one to the trace for span-level detail. If you export spans, record the trace id in the summary row from the current span context so the link is exact; see prompt and completion tracing for what those spans contain.
Worked example: a loop the metrics could not explain
An illustrative incident shows the flow. After a prompt change ships as billing-v14, the metrics dashboard shows p95 latency up from 6 to 9 seconds with no errors. Panel 1 tells a sharper story: looping runs went from 0.4% to 7% for v14 only, while the answered rate barely moved. Clicking the v14 looping row lists invocations whose maximum repeat is 3 or more; nearly all repeat the same tool, lookup_invoice.
The timeline of one such invocation looks like the figure above. The router transfers to billing_agent, which calls lookup_invoice, waits for the model, then calls it again with identical arguments, twice. Reading the session text shows why: the tool returns a page of results with a has_more flag, and the new prompt tells the model to make sure it has all invoices, so it calls again, without a page token the tool never asked for. The fix is in the tool contract, a page_token argument and a description of it, and the panel confirms the looping rate falls back after the next release. The metrics dashboard said something was slower; the behaviour panels said what, and the timeline said why.
Operating it
- Record every run, sample nothing. Summary rows are a few hundred bytes, so keep all of them; sampling would make rare loops and rare transfer paths vanish from exactly the panels meant to find them. Step rows are larger; keep them for 30 to 90 days and summaries for a year or more.
- Test the plugin like code. Run an agent on InMemoryRunner with a scripted model that emits a known sequence (two identical tool calls, one transfer, a final answer) and assert the summary: tool calls 2, maximum repeat 2, the expected agent path, final response true. Add a case where the run throws, to prove the error hook emits a row.
- Refresh and alert. A five-minute refresh is enough for behaviour panels. Alert on the looping rate and the answered rate per version against the previous version's baseline, not against a fixed threshold, because agents differ.
- Access. Because the tables hold no text, product, support and engineering can share one dashboard; only the session-store link needs a permission check.
Failure modes
- Leaked accumulators. If a run is cancelled and neither close hook fires, its entry stays in the map forever. Evict entries older than your longest plausible run on a timer and count evictions as a metric.
- Blocking the event path. onEventCallback runs inline with the agent. Never write to a database there; offer to a bounded queue, drop on overflow and count drops, so a slow warehouse can never slow the agent.
- Double counting. Streaming emits partial events; skipping them is essential, or token and step counts inflate in streaming mode only.
- Missing nested runs. Agents wrapped as tools may run in their own invocation. Check in your release whether their events reach the outer plugin, and record the parent invocation id if they do not, or the transfer graph has holes.
- Version blind spots. A summary without agent_version cannot split panels by release, which is the split you need most. Inject it at startup from the build.
- Text creeping into summaries. Someone will add the final answer to the summary row for convenience. Do not: the table then needs the session store's access controls and retention, and the dashboard can no longer be widely shared.
Trade-offs
Building your own summary tables duplicates some of what the BigQuery analytics plugin and span-derived metrics already give you. The return is control: a schema designed for the panels you need, cheap queries over one row per run, and behaviour signals such as repeat counts that neither source computes. The cost is one more pipeline to own and test. If you already run the analytics plugin, consider computing the same summary as a scheduled query over its event table instead of adding a plugin.
What to do next
- Write down the five questions your team asks during agent incidents and map each to a summary or step field.
- Add the InvocationRecorder plugin with a bounded sink, an eviction timer and counters for drops and evictions; deploy it to one environment.
- Verify the rows against a few hand-inspected sessions: event counts, tool calls, repeats and the agent path should match what you see.
- Build panels 1, 4 and 5 first, split by agent_version, with links from rows to invocation lists and from ids to timelines.
- Add the trace id and a session-store link to the timeline, then run a release through it and check that a deliberately injected loop shows up within one refresh.