A multi-tenant agent product holds data its customers' account teams want and seldom get. Is this customer's staff actually using the agent? Does it finish their tasks or hand them back to a human? Is usage rising before the renewal or quietly falling? Tenant analytics answers those questions with one tenant-level record per day and a health score that says why it moved.

This article builds that pipeline for an ADK Java agent. It covers the event schema, a plugin that emits events without ever copying conversation content, the nightly roll-up into tenant facts, an explainable health score with a fully worked example, cohort benchmarks that cannot reveal another customer's numbers, and the failure modes that make these dashboards lie.

What tenant analytics is for

Four systems look at the same traffic and answer different questions. Keep them apart and each stays simple.

SystemQuestionUnitTolerates loss?
Observabilityis the service healthy now?span, metricyes, sampled
SLA reportingdid we meet the contract?classified runno
Billingwhat does the tenant owe?rated usage eventno
Tenant analyticsis the tenant getting value?tenant-daya little, if counted

Analytics differs in two ways. It is about trends, so a dropped event shifts nothing that matters as long as you count the drops. It is also read by people outside engineering: customer success, sales and the customers themselves. That second point decides the privacy design. Anything that lands in this warehouse should be safe to show an account manager, so conversation text never goes in. The SLA's strict outcome classification lives in agent tenant SLA, and metered usage in agent tenant billing.

Architecture

Tenant analytics pipeline: content-free events in, explained health scores outADK Java RunnerTenantAnalyticsPluginBounded queuebatched, drops countedturn_eventsraw, append-onlyOther sourcesfeedback, seats, outcomestenant_dailydedupe, tenant-local daynightlyHealth scorecomponents + reasonsweekly roll-upCohort benchmarksonly cohorts of 10 or moreCS alertsscore drop + reasonsTenant dashboardown data onlyNo prompt or response text crosses the plugin boundary; user ids are keyed hashes.
The plugin emits counts and names only. Nightly jobs turn raw events into tenant-days, and weekly roll-ups feed an explained score, customer-success alerts and cohort quartiles.

The event schema

Choose events by the decisions they support, then strip each one to counts, names and ids. The turn event carries most of the weight:

FieldTypeNotes
event_idUUIDderived from the invocation id, so retries deduplicate
tenant_idstringfrom session state, set when the session is created
user_hashstringHMAC of the user id with a per-tenant key
session_idstringgroups turns into conversations
agent, agent_releasestringwhich agent and which release served the turn
started_at, ended_attimestampUTC; the tenant-local day is computed later
statusstringcompleted, or error plus an exception class
tool_calls, tool_errorsintcounts only
tool_namesstring[]names, never arguments or results
escalatedboolthe hand-off tool was called

Three more streams join it from outside the agent: thumbs-up and thumbs-down feedback from your API, task outcomes from the system of record (ticket closed, refund issued, order placed), and seat counts from your admin console. Outcomes from the system of record are worth the integration effort. An agent's own claim that it resolved a task measures its confidence, not the result.

Hashing user ids with a per-tenant HMAC key keeps active-user counts exact while making ids useless outside the warehouse, and it stops one user from being linked across tenants. Deleting a tenant's key at offboarding makes its historical hashes unlinkable.

Emitting events from a plugin

A plugin registered on the Runner sees every invocation, so it is the right place to emit turn events. It accumulates per invocation and emits once at the end. It never reads Content, only names and counts. AnalyticsSink and UserIdHasher are your own classes: a bounded queue that batches writes and counts what it drops, and an HMAC wrapper.

public final class TenantAnalyticsPlugin extends BasePlugin {
  private final ConcurrentHashMap<String, TurnAcc> open = new ConcurrentHashMap<>();
  private final AnalyticsSink sink;
  private final UserIdHasher hasher;
  private final String rootApp;

  public TenantAnalyticsPlugin(AnalyticsSink sink, UserIdHasher hasher, String rootApp) {
    super("tenant_analytics"); this.sink = sink; this.hasher = hasher; this.rootApp = rootApp;
  }

  @Override public Maybe<Content> beforeRunCallback(InvocationContext inv) {
    // Nested AgentTool runs share our plugins but are not user turns: skip them.
    if (rootApp.equals(inv.appName())) open.put(inv.invocationId(), new TurnAcc(Instant.now()));
    return Maybe.empty();
  }

  @Override public Maybe<Map<String, Object>> afterToolCallback(BaseTool tool,
      Map<String, Object> args, ToolContext ctx, Map<String, Object> result) {
    TurnAcc acc = open.get(ctx.invocationId());
    if (acc != null) acc.tool(tool.name(), !result.containsKey("error"));
    return Maybe.empty();
  }

  @Override public Maybe<Map<String, Object>> onToolErrorCallback(BaseTool tool,
      Map<String, Object> args, ToolContext ctx, Throwable error) {
    TurnAcc acc = open.get(ctx.invocationId());
    if (acc != null) acc.tool(tool.name(), false);
    return Maybe.empty();
  }

  @Override public Completable afterRunCallback(InvocationContext inv) {
    return Completable.fromAction(() -> emit(inv, "completed"));
  }

  @Override public Completable onRunErrorCallback(InvocationContext inv, Throwable error) {
    return Completable.fromAction(() -> emit(inv, "error:" + error.getClass().getSimpleName()));
  }

  private void emit(InvocationContext inv, String status) {
    TurnAcc acc = open.remove(inv.invocationId());     // remove() makes a second call a no-op
    if (acc == null) return;
    Map<String, Object> st = inv.session().state();
    if (Boolean.TRUE.equals(st.get("synthetic"))) return;   // verifiers, load tests, evals
    Object tenant = st.get("tenant_id");
    if (tenant == null) { sink.countMissingTenant(); return; }
    sink.offer(new TurnEvent(
        UUID.nameUUIDFromBytes(inv.invocationId().getBytes(StandardCharsets.UTF_8)),
        tenant.toString(), hasher.hash(tenant.toString(), inv.userId()), inv.session().id(),
        inv.agent().name(), String.valueOf(st.getOrDefault("agent_release", "unknown")),
        acc.started(), Instant.now(), status,
        acc.calls(), acc.errors(), acc.names(), acc.names().contains("escalate_to_human")));
  }
}

A few details carry weight here. The event id is derived from the invocation id, so a retried write deduplicates downstream. remove() makes emission happen at most once, whichever end-of-run callback fires. The app-name test matters once any AgentTool is created with includePlugins: its nested runner then calls this plugin too, under the calling agent's name as app name, and without the test every delegation would count as an extra turn by a phantom user. Sessions created by verifiers, load tests and evaluations carry synthetic in state and are skipped at the source. Otherwise a busy deploy week looks like a usage spike. The tool-error rule assumes your tools report failure with an error key. Adjust it to your own convention. The tenant id comes from session state written when the session was created, the same rule tenant data separation uses, not from parsing user ids.

From events to daily tenant facts

Raw events are for drilling down. Dashboards read a daily fact table with one row per tenant per tenant-local day, so a Sydney customer's Monday is not split across two UTC dates. The roll-up below is BigQuery SQL. It deduplicates on the event id, then counts tool breadth in a separate step, because unnesting tool names in the main query would multiply every row.

CREATE OR REPLACE TABLE analytics.tenant_daily AS
WITH turns AS (
  SELECT t.*, DATE(t.ended_at, tz.timezone) AS day
  FROM analytics.turn_events AS t
  JOIN analytics.tenants AS tz USING (tenant_id)
  WHERE TRUE
  QUALIFY ROW_NUMBER() OVER (PARTITION BY t.event_id ORDER BY t.ingested_at) = 1
),
tools AS (
  SELECT tenant_id, day, COUNT(DISTINCT tool) AS distinct_tools
  FROM turns, UNNEST(tool_names) AS tool
  GROUP BY tenant_id, day
)
SELECT
  tenant_id, day,
  COUNT(*)                        AS turns,
  COUNT(DISTINCT user_hash)       AS active_users,
  COUNT(DISTINCT session_id)      AS sessions,
  COUNTIF(status != 'completed')  AS failed_turns,
  COUNTIF(escalated)              AS escalated_turns,
  ANY_VALUE(tools.distinct_tools) AS distinct_tools
FROM turns LEFT JOIN tools USING (tenant_id, day)
GROUP BY tenant_id, day;

Do not sum daily active users into weekly ones. A person active on five days counts five times. Compute weekly active users from the raw events over the seven-day window. Rebuild the last few days every night instead of appending, so late events, such as outcomes that arrive when a ticket closes two days later, are counted.

A health score that explains itself

A score that customer success acts on has to come with reasons. A single number from a trained model invites the question "why?" and gives no answer. Start with a few components, each mapped linearly between a bad anchor (0) and a good anchor (1), combined with weights everyone has agreed on. Any component without enough data is dropped and the remaining weights are renormalised.

public record TenantWeek(int seats, int wau, int priorWau, int turns, int failedTurns,
    int resolved, int withOutcome, int distinctTools, int toolsEnabled) {}
public record Health(int score, List<String> reasons) {}

static double band(double x, double bad, double good) {
  return Math.max(0, Math.min(1, (x - bad) / (good - bad)));
}

public static Health score(TenantWeek w) {
  String[] names = {"adoption", "trend", "task success", "reliability", "breadth"};
  double[] weight = {0.30, 0.20, 0.25, 0.15, 0.10};
  double[] v = {
    band((double) w.wau() / Math.max(1, w.seats()), 0.10, 0.60),
    w.priorWau() == 0 ? Double.NaN : band((double) w.wau() / w.priorWau(), 0.70, 1.05),
    w.withOutcome() < 30 ? Double.NaN : band((double) w.resolved() / w.withOutcome(), 0.50, 0.85),
    band(1.0 - (double) w.failedTurns() / Math.max(1, w.turns()), 0.95, 0.995),
    band((double) w.distinctTools() / Math.max(1, w.toolsEnabled()), 0.20, 0.70)};
  List<String> why = new ArrayList<>();
  double sum = 0, wsum = 0;
  for (int i = 0; i < v.length; i++) {
    if (Double.isNaN(v[i])) { why.add(names[i] + ": not enough data"); continue; }
    sum += v[i] * weight[i]; wsum += weight[i];
    if (v[i] < 0.4) why.add(names[i] + " low (" + Math.round(v[i] * 100) + "/100)");
  }
  return new Health((int) Math.round(100 * sum / wsum), why);
}

The anchors are product judgements, not constants of nature. Set them from your own renewal history once you have some. Until then, write them down, review them each quarter, and version the score, so that a change in the formula cannot be mistaken for a change in the customer.

Worked example: reading a score of 32

Acme has 200 seats. This week 46 people used the agent, against 70 the week before. There were 3,100 turns, of which 40 failed. 410 tasks have an outcome from the ticketing system, and 260 of them were resolved. Acme has 12 tools enabled and used 4.

ComponentRaw valueMappedWeightContribution
adoption46 / 200 = 0.23(0.23 - 0.10) / 0.50 = 0.260.300.078
trend46 / 70 = 0.657below 0.70, so 00.200.000
task success260 / 410 = 0.634(0.634 - 0.50) / 0.35 = 0.3830.250.096
reliability1 - 40 / 3,100 = 0.987(0.987 - 0.95) / 0.045 = 0.8240.150.124
breadth4 / 12 = 0.333(0.333 - 0.20) / 0.50 = 0.2670.100.027

Unrounded, the contributions sum to 0.324, so the score is 32, with four reasons attached: adoption 26, trend 0, task success 38 and breadth 27. Reliability scores 82 and is not the problem, so this is not an outage to apologise for. The story is a 34% drop in weekly users with low tool breadth. The account manager's next question is answered by a drill-down on raw events: which agents and tools did the users who stopped coming rely on? Here, the drill-down shows the drop came almost entirely from one team whose main tool started failing its authentication after an IT change on their side. That is a phone call, not a product roadmap item. Without the reasons, the number 32 would have said only that something was wrong.

Cohort benchmarks without exposing another tenant

Customers ask how they compare with similar companies. Answer with your own tenant's position inside a cohort distribution, never with another tenant's values, and only when the cohort is large enough that no single tenant can be inferred.

SELECT industry, size_band,
       COUNT(*) AS tenants,
       APPROX_QUANTILES(adoption, 4) AS adoption_quartiles
FROM analytics.tenant_week
WHERE week = @week
GROUP BY industry, size_band
HAVING COUNT(*) >= 10;   -- smaller cohorts are not published

Ten is a floor, not a guarantee. Also suppress a cohort when one tenant contributes most of its volume, and never let a customer narrow the cohort with filters until it falls below the floor. Each tenant sees only its own data and the published quartiles. That boundary needs the same review as any other cross-tenant access, as described in agent tenant isolation.

Failure modes

  • Synthetic traffic counted as usage. Golden replays and load tests inflate adoption in deploy weeks. Mark them in session state and drop them at the source.
  • Nested runs counted as turns. Sub-agent runs started by an AgentTool look like extra sessions and users. Accept only runs whose app name is yours.
  • Missing tenant id. Sessions created by an old code path have no tenant. Count them, alert when the count is above zero, and do not guess.
  • Summed daily actives. Weekly users computed as a sum of daily users overstate adoption several times over. Compute distinct counts over the window.
  • UTC days. Customers far from UTC get two half-days for each working day, and their trend lines turn ragged.
  • Self-reported success. Treating the agent's own claim of resolution as the outcome inflates task success. Use the system of record.
  • Goodhart pressure. Once the score drives incentives, teams will push turns up rather than value. Keep turn volume out of the score.
  • Content leakage. A well-meant "top questions" feature copies prompts into the warehouse. Build it as a separate, consented pipeline with redaction, never as a column.

Trade-offs

A hand-weighted score is transparent and easy to argue with, but it is not tuned to predict churn. A model trained on renewal outcomes may predict better once you have hundreds of renewals, but it explains itself poorly. A reasonable path is to ship the weighted score now, log its components, and later train a model on those same components, so that the explanation survives. Daily facts cost a nightly job and lose intra-day detail. Live dashboards from raw events cost more and seldom change a customer-success decision. Content-free events give up the topic analysis product managers ask for, in exchange for a warehouse you can open to customers.

What to do next

  1. Write down three decisions tenant analytics should support, and drop every metric that does not feed one of them.
  2. Set tenant_id in session state at creation and add the synthetic flag to every verifier, load-test and eval session.
  3. Ship the plugin with a bounded sink, and alert on drops and missing tenants.
  4. Build tenant_daily on tenant-local days with event-id deduplication, and rebuild a trailing window nightly.
  5. Join task outcomes from the system of record before trusting any success rate.
  6. Ship the five-component score with reasons, recompute the Acme example as a unit test, and version the formula.
  7. Publish cohort quartiles only for cohorts of ten or more, and pair the dashboards with the operational dashboards your on-call team already uses.
Key takeaway: Tenant analytics turns agent traffic into one honest record per tenant per day and a health score that says why it moved. Emit content-free turn events from a plugin, take the tenant from session state, drop synthetic traffic at the source, roll up on tenant-local days, and take outcomes from the system of record. Attach reasons to every score, and publish benchmarks only for cohorts too large to reveal any one customer.