Metrics tell you the error rate. Logs tell you what went wrong in one conversation. Product questions need a third thing. Which intents end in escalation? Did the new release lower resolution on refunds? How long do users wait for a human approval? Answering these needs telemetry events: one row per thing that happened, with ids that join rows into invocations and sessions, kept in a warehouse so that you can ask questions you have not thought of yet.
ADK Java 1.11.0 ships a producer for this, BigQueryAgentAnalyticsPlugin. The BigQuery integration article shows how to switch it on. This page treats it as an analytics system. It covers the event model it writes, every configuration default as read from the 1.11.0 bytecode, the delivery guarantees and how rows are dropped, the privacy boundary, how to add business outcomes, and queries that answer product questions. Where the layout of keys inside its JSON columns is not documented, this page says so. Inspect a row before you build on a key.
Telemetry events versus metrics, logs and traces
| Signal | Unit | Good for | Poor for |
|---|---|---|---|
| Metrics | Pre-aggregated counter or histogram | Alerts, SLOs, dashboards | New questions about past behaviour |
| Logs | Line per step, short retention | Debugging one invocation | Joins and cohort analysis |
| Traces | Span tree, usually sampled | Latency inside one request | Counting outcomes across all users |
| Telemetry events | Row per lifecycle step, unsampled | Funnels, cohorts, release comparisons | Real-time paging |
The defining property of telemetry events is that they are complete and joinable. Every invocation is recorded, not a sample, and every row carries invocation_id, session_id and user_id. That is what lets you compute a rate per user cohort a month later. Metric dashboards, covered in metrics and dashboards for ADK Java, stay the place for alerting.
The event model ADK Java writes
The plugin implements every plugin callback and writes one row per callback, with event_type naming the step. The diagram above shows the core lifecycle. The 1.11.0 bytecode also contains these type names: AGENT_RESPONSE, TRANSFER_AGENT, A2A_INTERACTION, and HITL_CONFIRMATION_REQUEST, HITL_CREDENTIAL_REQUEST and HITL_INPUT_REQUEST for human-in-the-loop pauses. Which ones your agents emit depends on what they do, so run SELECT event_type, COUNT(*) ... GROUP BY 1 on a day of data before writing dashboards.
The typed columns are timestamp, event_type, agent, session_id, invocation_id, user_id, trace_id, span_id, parent_span_id, status, error_message and is_truncated. The payload goes in the JSON columns content, attributes and latency_ms, plus content_parts for multimodal parts. Two columns matter most for analytics. event_id is described in the schema as assigned before enqueue and preserved across Storage Write API retries, so it is your deduplication key. And trace_id and span_id let you jump from a suspicious row to its trace.
Configuring the plugin for analytics
Every default below was read from BigQueryLoggerConfig.builder() in 1.11.0. Several of them are wrong for an analytics workload:
| Setting | Default | Analytics guidance |
|---|---|---|
batchSize | 1 | Raise to 50-500; one append per row is wasteful at volume |
batchFlushInterval | 1 s | Keep around 1-5 s; it bounds freshness and the loss window |
queueMaxSize | 10,000 | Size it to peak rows/s multiplied by the worst append stall you tolerate |
shutdownTimeout | 10 s | Must fit inside your platform's termination grace period |
maxContentLength | 512,000 | Lower it unless you analyse full prompts; rows set is_truncated |
logSessionMetadata | true | Copies session state into rows; turn it off or audit the state |
logMultiModalContent | true | Off unless you need parts; large payloads use gcsBucketName offload |
clusteringFields | event_type, agent, user_id | Fine for most queries; the table name defaults to agent_events |
autoSchemaUpgrade / createViews | true / false | Turn views on; view names use viewPrefix, default v |
retryConfig | 3 retries, 1 s, x2, max 10 s | Leave as is; retries are why duplicates exist |
BigQueryLoggerConfig cfg = BigQueryLoggerConfig.builder()
.projectId("acme-prod").datasetId("agent_analytics").location("EU")
.batchSize(200).batchFlushInterval(Duration.ofSeconds(2))
.queueMaxSize(50_000).shutdownTimeout(Duration.ofSeconds(8))
.maxContentLength(4_096)
.logSessionMetadata(false)
.logMultiModalContent(false)
.customTags(Map.of("app_version", BuildInfo.VERSION, "deployment", "eu-west-canary"))
.contentFormatter((content, eventType) -> Masking.mask(content)) // apply(content, eventType)
.createViews(true)
.build();
BigQueryAgentAnalyticsPlugin analytics = new BigQueryAgentAnalyticsPlugin(cfg);
Runner runner = new InMemoryRunner(rootAgent, "support_app", List.of(analytics));customTags is the cheap way to get release comparisons. The map is written into each row's attributes under a custom_tags key, so app version, region and experiment arm become dimensions without code in the agent. The location must match the dataset's location. eventAllowlist and eventDenylist take event type names, if you want to stop, for example, LLM_REQUEST rows, which are the largest.
Delivery, drops and duplicates
The plugin queues rows in memory and a background batch processor appends them through the BigQuery Storage Write API's default stream. That gives you at-least-once delivery while the process lives, with known ways to lose rows. The 1.11.0 classes count drops by reason: queue_full, shutdown_timeout, append_error, serialization_error, writer_create_error, writer_permit_exhausted, late_after_finalize and after_close. getDropStats() returns them as an ImmutableMap<String, Long>. Export it, because an analytics table that silently lost 5% of rows at peak is worse than one with a known gap.
Meter meter = GlobalOpenTelemetry.getMeter("agent.analytics");
meter.gaugeBuilder("agent_analytics_dropped_rows").ofLongs().buildWithCallback(m ->
analytics.getDropStats().forEach((reason, n) ->
m.record(n, Attributes.of(AttributeKey.stringKey("reason"), reason))));Retries mean duplicates are possible, so read through a deduplicating view. Call analytics.close() during graceful shutdown before the JVM exits. The plugin also registers a shutdown hook, but an orderly close inside your own grace period is easier to reason about. Its warnings go through java.util.logging, so bridge JUL into your logging backend or they appear only on stderr.
CREATE OR REPLACE VIEW agent_analytics.events_dedup AS
SELECT * FROM agent_analytics.agent_events
WHERE timestamp >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 90 DAY)
QUALIFY ROW_NUMBER() OVER (PARTITION BY event_id ORDER BY timestamp) = 1;
The privacy boundary
By default this table is a copy of your conversations: user messages, model output and tool arguments, each truncated at maxContentLength. The built-in redaction is by key name only. JSON keys named password, api_key, access_token, refresh_token, id_token or client_secret (any case), or starting with temp:, are replaced with [REDACTED]. An email address inside a user's sentence is not. With logSessionMetadata on, session state is also copied into every row, so anything your agent keeps in state goes into the warehouse.
Use three layers. First, contentFormatter masks or drops content before it is written. It receives the content and the event type, so it can keep tool names and drop free text. Second, BigQuery controls: partition expiration for retention, and column-level access so that analysts query typed columns and only a few people can read content. Third, a pseudonymous user_id, because that column is also a clustering key.
Adding business outcomes
Lifecycle rows describe what the agent did, not whether the user got what they needed. Outcomes such as resolved, escalated, refund issued or thumbs down need explicit events. There are three ways to add them:
- State changes. When an event carries a non-empty state delta, the plugin writes a
STATE_DELTArow with the delta in its attributes. A tool that setscase_statustoescalatedtherefore produces an analytics row for free. This is convenient, but it ties analytics to session state, and every such key persists in the session. - Static tags.
customTagsfor dimensions fixed at deploy time. - An outcome table you own. Your application writes a small row keyed by
session_idandinvocation_idwhen the UI records feedback or the ticketing system closes a case. Many outcomes arrive after the invocation, sometimes days later, from systems that never see ADK.
public record OutcomeEvent(String eventId, String sessionId, String invocationId,
String outcome, // resolved | escalated | abandoned | thumbs_down
Instant at, String source) {}
// Written with the same event_id discipline: generate once, retry with the same id.
Queries that answer product questions
With deduplicated rows and an outcome table, the common product questions become short queries. The tool success rate, per release:
SELECT JSON_VALUE(attributes, '$.custom_tags.app_version') AS version,
JSON_VALUE(content, '$.tool') AS tool,
COUNTIF(event_type = 'TOOL_ERROR') / COUNT(*) AS error_rate,
COUNT(*) AS calls
FROM agent_analytics.events_dedup
WHERE event_type IN ('TOOL_COMPLETED', 'TOOL_ERROR')
AND timestamp >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
GROUP BY version, tool ORDER BY calls DESC;The JSON paths here are the ones to check first, not guarantees. The 1.11.0 bytecode builds tool content with tool and args keys, and latency with keys such as total_ms and time_to_first_token_ms. Run TO_JSON_STRING over a few rows of each event type and fix the paths before saving dashboards. A resolution funnel joins sessions to outcomes: count sessions with an INVOCATION_STARTING, then those whose final outcome is resolved, grouped by the first agent the root transferred to. For human-in-the-loop waits, pair TOOL_PAUSED rows with the later TOOL_COMPLETED for the same function call id, which the plugin records for this pairing. Inspect where the key lives in your rows before joining on it. Cost per outcome follows from token counts in LLM_RESPONSE rows. See agent cost tracking.
Worked example: a resolution drop that was half missing data
Release 4.2 goes out to a canary tagged deployment=eu-west-canary. A week later, the resolution rate for canary sessions is 61% against 68% for the stable version. Before blaming the release, check the drop gauge. queue_full counted 41,000 rows on the canary during a 20-minute BigQuery append stall on Tuesday, all on the canary's three pods. Excluding that window, resolution is 63% against 68%, so the gap is smaller but real.
The tool query shows lookup_policy errors at 9% on 4.2 against 1% on 4.1. Filtering TOOL_ERROR rows by error_message shows a schema validation failure on a new optional argument the model sometimes fills with null. The fix is one line in the tool's schema. After the fix, the canary's resolution rate matches stable within a day. The queue is raised to 50,000 and the batch size to 200, so the next stall drops nothing. Without the drop gauge, the team would have spent the week chasing a 4-point gap that was partly missing data.
Failure modes
- Default batch size in production. One append per row costs throughput, and queue drops follow under load.
- No drop accounting. Rates computed over an incomplete table look like product changes.
- Counting without deduplication. Retries inflate counts slightly and unevenly.
- Session state in every row. With
logSessionMetadataon, a customer profile kept in state is copied thousands of times. - Dashboards on guessed JSON keys. They return nulls silently after an upgrade changes the layout.
- Outcomes inferred from lifecycle rows. 'The agent replied' is not 'the user was helped'. Record outcomes explicitly.
Trade-offs
The built-in plugin gives you a complete event model in a day, but it ties you to BigQuery and to its schema. A plugin of your own that publishes to Pub/Sub or Kafka takes more work, and in return gives you a portable schema and durable buffering outside the JVM. Larger batches cost less but lose more rows on a crash. Capturing content makes failure analysis possible, and it makes the table sensitive data. Most teams capture typed columns and tool names for everyone, and full content only for a short-retention, restricted table. Continuous quality checks on sampled turns, described in continuous evaluation, fit on top of the same rows.
What to do next
- Enable the plugin in staging and list the event types and JSON keys your agents actually produce.
- Set batch size, queue size, shutdown timeout and content length explicitly, and turn off
logSessionMetadataunless your state is clean. - Add a
contentFormatterthat masks free text, and restrict access to thecontentcolumn. - Export
getDropStats()as a gauge, alert on any non-zero rate, and bridge JUL into your logs. - Create the deduplicating view and point every query at it.
- Add
customTagsfor app version and deployment, then build the per-release tool-error query. - Define an outcome event and write it from the systems that know the outcome. Then build the resolution funnel.