Operational dashboards tell you whether the agent is up and how fast it answers. They do not tell you whether people come back, whether they finish what they came to do, or where they give up. That is user analytics: product questions asked per end user, over days and weeks, rather than per request over minutes.
This article builds user analytics for an ADK Java agent on the BigQueryAgentAnalyticsPlugin that ships in google-adk 1.11.0, configured so the analytics table holds behaviour rather than conversation text. You will see what the plugin records and what its defaults do, how to pseudonymise user IDs before they reach the runner, how to add derived friction and outcome signals, the SQL for retention and task success, a worked example, and the ways this goes wrong. Per-tenant health scoring is a different problem and is covered in agent tenant analytics.
The questions user analytics answers
Start from the questions, because they decide what you record. A useful set for most agents:
- Activation. Of the users who started a first session, how many completed at least one task?
- Retention. Of a week's new users, what share return in week 1, week 2 and week 4?
- Task success. Per intent, what share of sessions end with an explicit or inferred success?
- Friction. How often do users rephrase, abandon mid-task, hit a tool error or get a refusal, and does any of that predict churn?
- Depth. Turns per session and sessions per active user, read together with success, since long sessions can mean engagement or struggle.
The unit hierarchy in ADK maps cleanly onto these questions. A user owns sessions; a session holds invocations, one per user message; an invocation produces events for agent starts, model calls, tool calls and state changes. Product metrics roll up from invocation to session to user to cohort. Keep that ladder in mind when writing SQL: most mistakes come from counting events where you meant invocations, or sessions where you meant users.
What the shipped analytics plugin records
ADK Java 1.11.0 includes com.google.adk.plugins.agentanalytics.BigQueryAgentAnalyticsPlugin, a plugin that implements the run, agent, model and tool callbacks and streams one row per event into BigQuery through the Storage Write API. Reading the 1.11.0 bytecode gives the event types it can emit: USER_MESSAGE_RECEIVED, INVOCATION_STARTING, INVOCATION_COMPLETED, AGENT_STARTING, AGENT_COMPLETED, LLM_REQUEST, LLM_RESPONSE, LLM_ERROR, TOOL_STARTING, TOOL_COMPLETED, TOOL_ERROR, TOOL_PAUSED, STATE_DELTA, AGENT_RESPONSE, A2A_INTERACTION and three human-in-the-loop request types.
Each row carries timestamp, event_id, event_type, agent, session_id, invocation_id, user_id, OpenTelemetry trace_id/span_id, a JSON content, attributes and latency_ms, plus status, error_message and is_truncated. The defaults from BigQueryLoggerConfig.builder() matter more than the schema:
| Setting | 1.11.0 default | What it means for user analytics |
|---|---|---|
maxContentLength | 512000 | Prompts, responses and tool payloads are logged, up to 512,000 characters per field |
logMultiModalContent | true | Image and file parts are recorded too (optionally offloaded to GCS) |
clusteringFields | event_type, agent, user_id | Per-user queries are cheap, which is good; the column is also an identifier |
batchSize / batchFlushInterval | 1 / 1 s | Row-at-a-time appends; raise both for throughput |
queueMaxSize | 10000 | When full, rows are dropped, not blocked |
eventAllowlist / eventDenylist | empty | Every event type is logged |
logSessionMetadata | true | Session state is copied into attributes on rows |
createViews | false | Set true to get one typed view per event type |
In short, out of the box the table is a transcript store keyed by user. The formatter carries a list of credential-like key names (password, api_key, access_token and similar) whose values it replaces with [REDACTED]; it has no notion of names, emails or health details. For product analytics you want the opposite trade: fewer event types, no content, and a user column that is not a raw account ID. The broader BigQuery setup, including datasets and IAM, is in ADK Java and BigQuery.
Configuring it for behaviour, not transcripts
The configuration below keeps the event types that analytics needs, replaces content with a size summary, turns off multimodal capture and session metadata, batches writes and creates the typed views. The contentFormatter is called as formatter.apply(content, eventType) for each row, so it can treat event types differently.
BigQueryLoggerConfig cfg = BigQueryLoggerConfig.builder()
.projectId("acme-assist")
.datasetId("agent_product") // not the ops dataset: different readers, retention
.tableName("agent_events")
.eventAllowlist(List.of(
"USER_MESSAGE_RECEIVED", "INVOCATION_COMPLETED", "AGENT_COMPLETED",
"TOOL_COMPLETED", "TOOL_ERROR", "LLM_ERROR", "STATE_DELTA"))
.logMultiModalContent(false)
.logSessionMetadata(false) // default true copies ALL session state into attributes
.maxContentLength(2000) // belt and braces on logged payload size
.contentFormatter((content, eventType) ->
Map.of("omitted", true, "chars", String.valueOf(content).length()))
.batchSize(200)
.batchFlushInterval(Duration.ofSeconds(5))
.createViews(true) // views named v_<event_type>, e.g. v_tool_error
.build();
BigQueryAgentAnalyticsPlugin analytics = new BigQueryAgentAnalyticsPlugin(cfg);
Runner runner = Runner.builder()
.agent(rootAgent).appName("assist")
.sessionService(sessionService)
.plugins(analytics, new UxSignalPlugin())
.build();The formatter only sees the content column. In 1.11.0, logSessionMetadata copies the whole session state into attributes and STATE_DELTA rows carry the delta in attributes.state_delta, both bypassing it, so keep personal data out of session state entirely. Then the deltas hold only signals such as the ux:* keys below. The plugin exposes getDropStats(); the 1.11.0 code counts reasons such as queue_full, append_error and shutdown_timeout. Those names are what the current build uses, not a documented contract, so export the whole map as a gauge rather than alerting on one key. Call close() on shutdown so buffered rows flush.
Pseudonymous user IDs
The userId you pass to runner.runAsync(userId, sessionId, message) is the value that ends up in user_id. The formatter rewrites content, not that column, so pseudonymise before the runner. A keyed hash gives a stable identifier per user that cannot be reversed without the key, and rotating or destroying the key unlinks the history in one step.
final class UserIds {
private final SecretKeySpec key; // from Secret Manager, never from config files
UserIds(byte[] secret) { this.key = new SecretKeySpec(secret, "HmacSHA256"); }
String pseudonym(String accountId) {
try {
Mac mac = Mac.getInstance("HmacSHA256");
mac.init(key);
byte[] h = mac.doFinal(accountId.getBytes(StandardCharsets.UTF_8));
return "u_" + HexFormat.of().formatHex(h, 0, 16); // 128 bits is plenty
} catch (GeneralSecurityException e) {
throw new IllegalStateException(e);
}
}
}
// at the gateway
String uid = userIds.pseudonym(principal.accountId());
runner.runAsync(uid, sessionId, userMessage).subscribe(sink::send, sink::fail);Using the pseudonym as the runner's user ID means the session store is keyed by it as well, which is usually what you want. If support staff must look up a user's sessions, they compute the pseudonym from the account ID through an audited service; the analytics dataset never needs the mapping. The wider question of what personal data reaches which sink is covered in ADK Java PII redaction.
Derived signals: friction, outcome, intent
With content gone, meaning has to come from signals computed in the application while the content is still in memory. Three kinds carry most of the value:
- Friction: the user rephrased the previous request, said the answer was wrong, or repeated a question after a tool error.
- Outcome: an explicit thumbs up or down, or an inferred completion such as a booking tool returning a confirmation number.
- Intent: a coarse label from a fixed list, so success can be compared within a task type rather than across very different tasks.
A small plugin computes friction on each user message and writes it to session state under a ux: prefix. When before-agent callbacks all return empty but changed state, the 1.11.0 BaseAgent emits an event carrying the delta, which the analytics plugin logs as STATE_DELTA. A rephrase is approximated by token overlap with the previous user message; crude, but cheap and explainable, and you can swap in embeddings later.
final class UxSignalPlugin extends BasePlugin {
UxSignalPlugin() { super("ux_signals"); }
@Override
public Maybe<Content> beforeAgentCallback(BaseAgent agent, CallbackContext ctx) {
if (!agent.name().equals("root")) return Maybe.empty(); // once per invocation
String now = ctx.userContent().map(UxSignalPlugin::text).orElse("");
String prev = previousUserText(ctx.events(), ctx.invocationId());
double overlap = jaccard(tokens(now), tokens(prev));
ctx.state().put("ux:rephrase", overlap >= 0.6 || now.matches("(?is)^(no|that's wrong|i meant).*"));
ctx.state().put("ux:turn", ctx.events().stream()
.filter(e -> e.author().equals("user")).count());
return Maybe.empty(); // never alters the turn
}
// text(), tokens(), jaccard(), previousUserText() are plain helpers
}Explicit feedback does not belong in the agent at all. The client posts it to a small feedback endpoint that appends rows of (invocation_id, user_id, rating, reason_code, ts) to a separate user_feedback table. Joining on invocation_id connects a rating to the turn that earned it.
From events to retention and success
Build a daily fact table per user, then answer questions from it. The first query rolls events up to one row per user per day, using only columns and JSON paths the 1.11.0 schema defines.
CREATE OR REPLACE TABLE agent_product.user_day AS
SELECT
user_id,
DATE(timestamp) AS day,
COUNT(DISTINCT session_id) AS sessions,
COUNTIF(event_type = 'USER_MESSAGE_RECEIVED') AS turns,
COUNTIF(event_type IN ('TOOL_ERROR', 'LLM_ERROR')) AS errors,
COUNTIF(event_type = 'STATE_DELTA'
AND JSON_VALUE(attributes, '$.state_delta."ux:rephrase"') = 'true') AS rephrases
FROM agent_product.agent_events
WHERE DATE(timestamp) >= DATE_SUB(CURRENT_DATE(), INTERVAL 90 DAY)
GROUP BY user_id, day;Weekly cohort retention then needs only the fact table. A user's cohort is the week of their first active day; week N retention is the share of the cohort active in week N.
WITH first_seen AS ( -- full history, not the 90-day window, or old users look new
SELECT user_id, DATE_TRUNC(MIN(DATE(timestamp)), WEEK(MONDAY)) AS cohort
FROM agent_product.agent_events WHERE event_type = 'USER_MESSAGE_RECEIVED' GROUP BY user_id),
active AS (
SELECT DISTINCT user_id, DATE_TRUNC(day, WEEK(MONDAY)) AS wk FROM agent_product.user_day)
SELECT cohort,
DATE_DIFF(wk, cohort, WEEK(MONDAY)) AS week_n,
COUNT(DISTINCT a.user_id) / ANY_VALUE(cohort_size) AS retained
FROM active a JOIN first_seen f USING (user_id)
JOIN (SELECT cohort, COUNT(*) AS cohort_size FROM first_seen GROUP BY cohort) USING (cohort)
GROUP BY cohort, week_n
ORDER BY cohort, week_n;Check the JSON path for the rephrase flag against a real STATE_DELTA row first; quoting rules for keys containing a colon are easy to get wrong, and a wrong path returns NULL silently, which reads as zero friction. Track the fact table's row count per day in the dashboards described in metrics and dashboards so a broken pipeline shows up as a gap rather than as a quiet improvement.
Worked example: what predicts not coming back
A travel assistant takes 2,000 new users in one week. The retention query shows 760 of them (38%) active again in week 1. Product wants to know what predicts not returning.
Splitting the cohort by whether the first session contained a TOOL_ERROR row: 400 users hit an error, and 88 of them returned (22%); the other 1,600 had clean first sessions and 672 returned (42%). Together that is the 760. The gap is large, but it is a correlation. Users who hit errors may have been attempting harder tasks, such as multi-city bookings, that churn more anyway.
So the analyst repeats the split within one intent label. For intent=flight_change the gap holds at a similar size, and the error rows cluster on one tool, the fare rules lookup, timing out. That is now an engineering ticket with a measured user cost attached: fixing it could move week-1 retention for affected users from 22% toward 42%, worth roughly 80 retained users per 400 affected if the gap is causal. After the fix, the same query on the next two cohorts is the test. If you can, ship the fix behind a percentage rollout and compare the two arms, which removes the confounding entirely.
Failure modes
- Transcripts in the analytics table. Leaving defaults on puts full prompts and responses in a widely read dataset. Use the allowlist and formatter, and test them by querying for a known phrase after a smoke test.
- Raw account IDs.
user_idis a clustering column, so it gets used in every query and copied into every extract. Pseudonymise at the gateway. - Silent drops.
queueMaxSizeoverflow and append failures drop rows rather than slowing the agent; retention computed over a week with an outage looks like churn. ExportgetDropStats()and mark affected days. - Double counting. Counting
AGENT_COMPLETEDrows as turns multiplies by the number of sub-agents. CountUSER_MESSAGE_RECEIVEDor distinctinvocation_id. - Proxy drift. A rephrase heuristic tuned on English misfires on other languages, and a prompt change can alter how users phrase follow-ups. Re-validate signals against a labelled sample each quarter.
- Shared identities. Kiosks, shared logins and API keys used by scripts inflate one user's numbers. Flag accounts with abnormal turn rates and exclude them from cohorts.
Trade-offs
The shipped plugin saves you from writing an event pipeline, a schema and batching logic, and its trace and span IDs let you join product rows to traces from OpenTelemetry. The price is that it is a general-purpose logger whose defaults favour debugging over privacy, so you spend configuration effort removing things. A custom plugin that emits only derived signals is smaller and safer but is yours to maintain.
Inferred signals scale to every session; explicit feedback is sparse and biased toward unhappy users but is unambiguous. Use both, and treat disagreements between them as a prompt to improve the inference. Finally, keep the product dataset separate from the operations dataset: different people read it, it needs longer retention for cohort analysis, and it should be deletable per user on request.
What to do next
- Write down five product questions and the event or signal that answers each.
- Attach
BigQueryAgentAnalyticsPluginwith an event allowlist, a content formatter that omits text,logMultiModalContent(false)and batching. - Pseudonymise user IDs with a keyed hash before
runAsync, with the key in a secret store. - Add a signal plugin that writes
ux:*keys, and a feedback endpoint with its own table. - Build the
user_dayfact table as a scheduled query and the cohort retention query on top of it. - Export
getDropStats()and annotate days with drops before reading trends. - Pick one friction signal, segment it by intent, and turn the largest gap into a ticket with a measured user cost.