A customer writes in: the assistant said their refund was issued, and it was not. The support engineer needs to see that conversation as the agent saw it: what the user said, which tools ran with which arguments, what came back, what state changed and what the agent finally answered. The on-call engineer needs a different view of the same data: which sessions in the last hour hit tool errors, and which agent in a multi-agent app handled them.

ADK Java already records almost all of this. Every turn is a sequence of events in a session, each event carries its tool calls, state delta and errors, and the runtime emits OpenTelemetry spans keyed by session and invocation. What it does not give you is a safe way to show it to people. This article builds that: a read-only inspection API over BaseSessionService and BaseArtifactService with scoped access, redaction, audit and links into traces, plus a worked example of the refund ticket.

The questions an inspection API must answer

Design the API from the questions, not the storage.

QuestionWho asksData source
What did the agent say and do in this conversation?SupportSession events in order
Which tool calls failed, with what error?Support, opsFunction responses and event error fields
What did the agent know at this point?Support, engineeringState replayed from state deltas
Which file did the user upload, which version did the agent read?SupportArtifact keys and versions
Why was it slow or expensive?OpsSpans and token usage per event
Which sessions hit errors in the last hour?OpsAn index you build from events or spans

The last row matters: the session service is keyed by app, user and session, so it answers questions about one conversation well and questions across users badly. Fleet-wide queries belong in your trace or log backend, not in a loop over listSessions.

What ADK Java already gives you

The session service interface in adk-java has the reads you need. getSession(appName, userId, sessionId, Optional<GetSessionConfig>) returns a Maybe<Session>, empty if the session does not exist; GetSessionConfig.builder() offers numRecentEvents and afterTimestamp to bound how many events come back. listSessions(appName, userId) returns a Single<ListSessionsResponse>, and listEvents returns events with an optional next-page token. On the artifact side, listArtifactKeys, listVersions and loadArtifact with a nullable version cover what was stored.

Each Event exposes id(), invocationId(), author(), content(), functionCalls(), functionResponses(), partial(), errorCode(), errorMessage(), usageMetadata() and timestamp(), which is epoch milliseconds. Its actions() carry stateDelta(), artifactDelta(), transferToAgent() and escalate(). The events article explains the full contract.

The ADK dev server, used by the dev UI, already exposes routes such as GET /apps/{appName}/users/{userId}/sessions/{sessionId}, artifact routes under the same path, and GET /debug/trace/session/{sessionId}. Copy the URL shape, not the server. It lives in the dev module, has no authentication, ships create, delete and run endpoints next to the reads, and its trace routes read spans held in memory by the server process, which a production deployment does not have.

Architecture: a separate read-only service

A read-only inspection service in front of ADK's storesSupport consoleticket in handOps dashboardsmetadata onlyInspection API1. authenticate caller2. check scope and ticket3. read through services4. redact, then shape views5. write audit recordGET onlyBaseSessionServiceevents, stateBaseArtifactServicekeys, versionsTrace backendby session_idAudit logwho read what, under which ticketAgent runtimeRunner writes eventswritesThe runtime writes; the inspection API only reads. No run, create or delete endpoints.
The inspection API is a separate, read-only service. It reads the same stores the runtime writes, and every request passes authentication, scope checks, redaction and audit.

Run it as its own deployment with its own service account that has read-only access to the session database and artifact bucket. If the agent runtime and the inspection API share a process, a bug in the inspection code can affect live conversations, and the read-only guarantee becomes a code review promise instead of a database permission.

Event views, not raw events

Never return raw Event objects. They contain full model content, tool arguments and state values, and their shape changes between ADK versions. Map each event to a view record that you version and redact.

public record ToolCallView(String name, Map<String, Object> args) {}

public record EventView(
    String id, String invocationId, String author, Instant at, String kind,
    String text, List<ToolCallView> toolCalls, List<String> toolErrors,
    Set<String> stateKeysChanged, Map<String, Integer> artifactVersions,
    String transferTo, boolean escalate, String error) {

  static EventView of(Event e, Redactor r) {
    String text = e.content().flatMap(Content::parts).orElse(List.of()).stream()
        .map(p -> p.text().orElse("")).filter(t -> !t.isEmpty())
        .map(r::text).collect(Collectors.joining("\n"));
    List<ToolCallView> calls = e.functionCalls().stream()
        .map(fc -> new ToolCallView(fc.name().orElse("?"), r.args(fc.args().orElse(Map.of()))))
        .toList();
    List<String> toolErrors = e.functionResponses().stream()
        .filter(fr -> fr.response().map(m -> m.containsKey("error")).orElse(false))
        .map(fr -> fr.name().orElse("?")).toList();
    String kind = !calls.isEmpty() ? "tool_call"
        : !e.functionResponses().isEmpty() ? "tool_result"
        : e.finalResponse() ? "final" : "message";
    return new EventView(e.id(), e.invocationId(), e.author(),
        Instant.ofEpochMilli(e.timestamp()), kind, text, calls, toolErrors,
        e.actions().stateDelta().keySet(), e.actions().artifactDelta(),
        e.actions().transferToAgent().orElse(null),
        e.actions().escalate().orElse(false),
        e.errorCode().map(Object::toString).orElse(e.errorMessage().orElse(null)));
  }
}

The toolErrors rule assumes your tools report failure with an error key in the response map, a common convention; adapt it to whatever your tools return. State values are deliberately absent from the timeline: only the keys that changed appear, and values come from a separate, more tightly scoped endpoint.

Timeline and state endpoints

Three endpoints answer most tickets: a session list for one user, a timeline, and state at a given event. A Spring MVC controller keeps the blocking RxJava calls simple, since blockingGet() on a Maybe returns null when the session is absent.

@RestController
@RequestMapping("/inspect/v1/apps/{app}/users/{user}")
public final class InspectionController {
  private static final int MAX_EVENTS = 2_000;
  private final BaseSessionService sessions;
  private final BaseArtifactService artifactService;
  private final AccessPolicy policy;
  private final Redactor redactor;
  private final AuditLog audit;

  // constructor injection omitted

  @GetMapping("/sessions/{sid}/timeline")
  public List<EventView> timeline(@PathVariable String app, @PathVariable String user,
      @PathVariable String sid, @RequestParam(defaultValue = "200") int limit,
      @RequestHeader("X-Ticket-Id") String ticket, Principal who) {
    policy.requireSessionRead(who, app, user, ticket);       // 403 if out of scope
    audit.record(who, ticket, "timeline", app, user, sid);
    var cfg = GetSessionConfig.builder().numRecentEvents(Math.min(limit, MAX_EVENTS)).build();
    Session s = sessions.getSession(app, user, sid, Optional.of(cfg)).blockingGet();
    if (s == null) throw new ResponseStatusException(HttpStatus.NOT_FOUND);
    return s.events().stream()
        .filter(e -> !e.partial().orElse(false))             // drop streaming fragments
        .map(e -> EventView.of(e, redactor)).toList();
  }

  @GetMapping("/sessions/{sid}/state")
  public Map<String, Object> stateAt(@PathVariable String app, @PathVariable String user,
      @PathVariable String sid, @RequestParam String atEventId,
      @RequestHeader("X-Ticket-Id") String ticket, Principal who) {
    policy.requireStateRead(who, app, user, ticket);         // narrower than timeline
    audit.record(who, ticket, "state", app, user, sid + "@" + atEventId);
    Session s = sessions.getSession(app, user, sid, Optional.empty()).blockingGet();
    if (s == null) throw new ResponseStatusException(HttpStatus.NOT_FOUND);
    if (s.events().size() > MAX_EVENTS) throw new ResponseStatusException(HttpStatus.PAYLOAD_TOO_LARGE);
    Map<String, Object> state = new TreeMap<>();
    for (Event e : s.events()) {                             // replay deltas in order
      e.actions().stateDelta().forEach((k, v) -> state.put(k, redactor.value(k, v)));
      if (e.id().equals(atEventId)) return state;
    }
    throw new ResponseStatusException(HttpStatus.NOT_FOUND, "event not in session");
  }
}

Replaying deltas reconstructs what this session recorded, which is not always the whole picture. Keys with the user: and app: prefixes are shared across sessions, so another conversation may have changed them in between, and temp: keys are not meant to outlive an invocation. The session context article explains the prefixes. Label the response as a replay, and show the current shared values separately.

The session list and artifact endpoints follow the same pattern: check scope, audit, read, shape. For artifacts, return names and version numbers by default and serve content only through a separate download endpoint with its own permission, because uploaded files are often the most sensitive data in a session.

  @GetMapping("/sessions/{sid}/artifacts")
  public Map<String, List<Integer>> artifacts(@PathVariable String app,
      @PathVariable String user, @PathVariable String sid,
      @RequestHeader("X-Ticket-Id") String ticket, Principal who) {
    policy.requireSessionRead(who, app, user, ticket);
    audit.record(who, ticket, "artifacts", app, user, sid);
    Map<String, List<Integer>> out = new TreeMap<>();
    for (String name : artifactService.listArtifactKeys(app, user, sid).blockingGet().filenames()) {
      out.put(name, artifactService.listVersions(app, user, sid, name).blockingGet());
    }
    return out;     // names and versions only; content needs a separate permission
  }

For the session list, call listSessions(app, user) and return IDs and lastUpdateTime() only. Implementations may omit events and state from listed sessions, so never build a list view that depends on them.

Access, redaction and audit

An inspection API is a data export path into every customer conversation, so its controls matter more than its features.

  • Scope by ticket. Support may read one user's sessions only while a ticket for that user is open; the policy checks the ticket system, not a header the caller can invent. Ops roles see metadata, such as kinds, timings, tool names and error codes, but no text or arguments.
  • Redact at the edge of the service. The redactor masks known sensitive argument names, account numbers, email addresses and anything your privacy review lists, before data leaves the process. Redact state values by key allow-list, not deny-list.
  • Audit every read. Who, which ticket, which session, which endpoint, when. Ship audit records to a store the inspection service cannot modify.
  • No write paths. No run, create or delete, and no replay of a turn against a live model. Re-running a conversation belongs in an evaluation environment with copied, consented data.
  • Break-glass is separate. If engineers occasionally need raw content, issue time-limited elevated access that pages a second person.

Linking timelines to traces

Timelines tell you what happened; traces tell you how long it took and what the model was sent. ADK Java's tracing sets attributes including gcp.vertex.agent.session_id, gcp.vertex.agent.invocation_id and gcp.vertex.agent.event_id on spans named invoke_agent, call_llm and execute_tool. Each timeline entry can therefore carry a deep link into your trace backend filtered by invocation ID, so the support view and the latency view refer to the same turn. The observability article covers wiring the exporter. Note that some of these span attributes carry full LLM requests, responses and tool arguments, so the trace backend needs the same access controls as the inspection API.

Worked example: the refund that never happened

Back to the refund ticket. The engineer opens the timeline for the user's session and sees, after redaction:

[
 {"author":"user","kind":"message","text":"Please refund order ****4417"},
 {"author":"billing_agent","kind":"tool_call",
  "toolCalls":[{"name":"issue_refund","args":{"order_id":"****4417","amount":"[redacted]"}}]},
 {"author":"billing_agent","kind":"tool_result","toolErrors":["issue_refund"]},
 {"author":"billing_agent","kind":"final",
  "text":"Your refund has been issued and should arrive in 3-5 days.",
  "stateKeysChanged":["refund_status"]}
]

Three facts are visible in seconds. The tool failed. The agent answered as if it had succeeded. And it wrote refund_status, so the state endpoint at the final event shows what it recorded. The trace link for that invocation shows the tool response the model received. The fix is in the agent, not the billing system: an after-tool callback that turns tool errors into an explicit failure the instructions must report, as described in the callbacks article, plus an evaluation case built from this conversation.

Failure modes

  • In-memory session service in production. InMemorySessionService loses everything on restart, so the inspection API finds nothing. Use a persistent session service before building inspection.
  • Unbounded reads. A long-running session can hold thousands of events. Bound timeline reads with numRecentEvents and refuse full replays above a limit.
  • Fleet queries through listSessions. Iterating users to find errors hammers the session database. Index error events or spans in a log or trace store instead.
  • Streaming fragments. Partial events duplicate text and confuse readers; filter them, and keep a toggle for engineers who need them.
  • Timestamp units. timestamp() is epoch milliseconds; treating it as seconds puts every event tens of thousands of years in the future.
  • Redaction drift. A new tool adds an argument containing personal data and the deny-list misses it. Allow-list argument names per tool and fail closed.

Trade-offs

ChoiceGainCost
Separate read-only serviceDatabase-enforced safety, independent scalingOne more deployment to run
Views instead of raw eventsStable contract, redaction in one placeMapping code to maintain per ADK upgrade
State replay from deltasState at any point in the turnShared prefixed keys can mislead without labelling
Trace links instead of copied spansNo second copy of sensitive payloadsTwo systems for one investigation

What to do next

  1. Confirm production uses a persistent session service and artifact service.
  2. Create a read-only database role and service account for a separate inspection deployment.
  3. Implement the event view, the redactor with per-tool argument allow-lists, and the timeline and state endpoints above.
  4. Wire the access policy to your ticket system and send audit records to an append-only store.
  5. Add trace deep links by invocation ID and apply the same access rules to the trace backend.
  6. Take the last five escalated tickets, walk them through the API, and turn each finding into an evaluation case. For memory-related tickets, use the memory inspection guide.
Key takeaway: ADK Java records everything support and operations need in session events, state deltas, artifacts and spans. Expose it through a separate read-only service that maps events to redacted views, scopes access to open tickets, audits every read and links each turn to its trace, and never ship the dev server's mutation endpoints to production.