A customer writes in: the assistant said their refund was issued, and it was not. The support engineer needs to see that conversation as the agent saw it: what the user said, which tools ran with which arguments, what came back, what state changed and what the agent finally answered. The on-call engineer needs a different view of the same data: which sessions in the last hour hit tool errors, and which agent in a multi-agent app handled them.
ADK Java already records almost all of this. Every turn is a sequence of events in a session, each event carries its tool calls, state delta and errors, and the runtime emits OpenTelemetry spans keyed by session and invocation. What it does not give you is a safe way to show it to people. This article builds that: a read-only inspection API over BaseSessionService and BaseArtifactService with scoped access, redaction, audit and links into traces, plus a worked example of the refund ticket.
The questions an inspection API must answer
Design the API from the questions, not the storage.
| Question | Who asks | Data source |
|---|---|---|
| What did the agent say and do in this conversation? | Support | Session events in order |
| Which tool calls failed, with what error? | Support, ops | Function responses and event error fields |
| What did the agent know at this point? | Support, engineering | State replayed from state deltas |
| Which file did the user upload, which version did the agent read? | Support | Artifact keys and versions |
| Why was it slow or expensive? | Ops | Spans and token usage per event |
| Which sessions hit errors in the last hour? | Ops | An index you build from events or spans |
The last row matters: the session service is keyed by app, user and session, so it answers questions about one conversation well and questions across users badly. Fleet-wide queries belong in your trace or log backend, not in a loop over listSessions.
What ADK Java already gives you
The session service interface in adk-java has the reads you need. getSession(appName, userId, sessionId, Optional<GetSessionConfig>) returns a Maybe<Session>, empty if the session does not exist; GetSessionConfig.builder() offers numRecentEvents and afterTimestamp to bound how many events come back. listSessions(appName, userId) returns a Single<ListSessionsResponse>, and listEvents returns events with an optional next-page token. On the artifact side, listArtifactKeys, listVersions and loadArtifact with a nullable version cover what was stored.
Each Event exposes id(), invocationId(), author(), content(), functionCalls(), functionResponses(), partial(), errorCode(), errorMessage(), usageMetadata() and timestamp(), which is epoch milliseconds. Its actions() carry stateDelta(), artifactDelta(), transferToAgent() and escalate(). The events article explains the full contract.
The ADK dev server, used by the dev UI, already exposes routes such as GET /apps/{appName}/users/{userId}/sessions/{sessionId}, artifact routes under the same path, and GET /debug/trace/session/{sessionId}. Copy the URL shape, not the server. It lives in the dev module, has no authentication, ships create, delete and run endpoints next to the reads, and its trace routes read spans held in memory by the server process, which a production deployment does not have.
Architecture: a separate read-only service
Run it as its own deployment with its own service account that has read-only access to the session database and artifact bucket. If the agent runtime and the inspection API share a process, a bug in the inspection code can affect live conversations, and the read-only guarantee becomes a code review promise instead of a database permission.
Event views, not raw events
Never return raw Event objects. They contain full model content, tool arguments and state values, and their shape changes between ADK versions. Map each event to a view record that you version and redact.
public record ToolCallView(String name, Map<String, Object> args) {}
public record EventView(
String id, String invocationId, String author, Instant at, String kind,
String text, List<ToolCallView> toolCalls, List<String> toolErrors,
Set<String> stateKeysChanged, Map<String, Integer> artifactVersions,
String transferTo, boolean escalate, String error) {
static EventView of(Event e, Redactor r) {
String text = e.content().flatMap(Content::parts).orElse(List.of()).stream()
.map(p -> p.text().orElse("")).filter(t -> !t.isEmpty())
.map(r::text).collect(Collectors.joining("\n"));
List<ToolCallView> calls = e.functionCalls().stream()
.map(fc -> new ToolCallView(fc.name().orElse("?"), r.args(fc.args().orElse(Map.of()))))
.toList();
List<String> toolErrors = e.functionResponses().stream()
.filter(fr -> fr.response().map(m -> m.containsKey("error")).orElse(false))
.map(fr -> fr.name().orElse("?")).toList();
String kind = !calls.isEmpty() ? "tool_call"
: !e.functionResponses().isEmpty() ? "tool_result"
: e.finalResponse() ? "final" : "message";
return new EventView(e.id(), e.invocationId(), e.author(),
Instant.ofEpochMilli(e.timestamp()), kind, text, calls, toolErrors,
e.actions().stateDelta().keySet(), e.actions().artifactDelta(),
e.actions().transferToAgent().orElse(null),
e.actions().escalate().orElse(false),
e.errorCode().map(Object::toString).orElse(e.errorMessage().orElse(null)));
}
}The toolErrors rule assumes your tools report failure with an error key in the response map, a common convention; adapt it to whatever your tools return. State values are deliberately absent from the timeline: only the keys that changed appear, and values come from a separate, more tightly scoped endpoint.
Timeline and state endpoints
Three endpoints answer most tickets: a session list for one user, a timeline, and state at a given event. A Spring MVC controller keeps the blocking RxJava calls simple, since blockingGet() on a Maybe returns null when the session is absent.
@RestController
@RequestMapping("/inspect/v1/apps/{app}/users/{user}")
public final class InspectionController {
private static final int MAX_EVENTS = 2_000;
private final BaseSessionService sessions;
private final BaseArtifactService artifactService;
private final AccessPolicy policy;
private final Redactor redactor;
private final AuditLog audit;
// constructor injection omitted
@GetMapping("/sessions/{sid}/timeline")
public List<EventView> timeline(@PathVariable String app, @PathVariable String user,
@PathVariable String sid, @RequestParam(defaultValue = "200") int limit,
@RequestHeader("X-Ticket-Id") String ticket, Principal who) {
policy.requireSessionRead(who, app, user, ticket); // 403 if out of scope
audit.record(who, ticket, "timeline", app, user, sid);
var cfg = GetSessionConfig.builder().numRecentEvents(Math.min(limit, MAX_EVENTS)).build();
Session s = sessions.getSession(app, user, sid, Optional.of(cfg)).blockingGet();
if (s == null) throw new ResponseStatusException(HttpStatus.NOT_FOUND);
return s.events().stream()
.filter(e -> !e.partial().orElse(false)) // drop streaming fragments
.map(e -> EventView.of(e, redactor)).toList();
}
@GetMapping("/sessions/{sid}/state")
public Map<String, Object> stateAt(@PathVariable String app, @PathVariable String user,
@PathVariable String sid, @RequestParam String atEventId,
@RequestHeader("X-Ticket-Id") String ticket, Principal who) {
policy.requireStateRead(who, app, user, ticket); // narrower than timeline
audit.record(who, ticket, "state", app, user, sid + "@" + atEventId);
Session s = sessions.getSession(app, user, sid, Optional.empty()).blockingGet();
if (s == null) throw new ResponseStatusException(HttpStatus.NOT_FOUND);
if (s.events().size() > MAX_EVENTS) throw new ResponseStatusException(HttpStatus.PAYLOAD_TOO_LARGE);
Map<String, Object> state = new TreeMap<>();
for (Event e : s.events()) { // replay deltas in order
e.actions().stateDelta().forEach((k, v) -> state.put(k, redactor.value(k, v)));
if (e.id().equals(atEventId)) return state;
}
throw new ResponseStatusException(HttpStatus.NOT_FOUND, "event not in session");
}
}Replaying deltas reconstructs what this session recorded, which is not always the whole picture. Keys with the user: and app: prefixes are shared across sessions, so another conversation may have changed them in between, and temp: keys are not meant to outlive an invocation. The session context article explains the prefixes. Label the response as a replay, and show the current shared values separately.
The session list and artifact endpoints follow the same pattern: check scope, audit, read, shape. For artifacts, return names and version numbers by default and serve content only through a separate download endpoint with its own permission, because uploaded files are often the most sensitive data in a session.
@GetMapping("/sessions/{sid}/artifacts")
public Map<String, List<Integer>> artifacts(@PathVariable String app,
@PathVariable String user, @PathVariable String sid,
@RequestHeader("X-Ticket-Id") String ticket, Principal who) {
policy.requireSessionRead(who, app, user, ticket);
audit.record(who, ticket, "artifacts", app, user, sid);
Map<String, List<Integer>> out = new TreeMap<>();
for (String name : artifactService.listArtifactKeys(app, user, sid).blockingGet().filenames()) {
out.put(name, artifactService.listVersions(app, user, sid, name).blockingGet());
}
return out; // names and versions only; content needs a separate permission
}For the session list, call listSessions(app, user) and return IDs and lastUpdateTime() only. Implementations may omit events and state from listed sessions, so never build a list view that depends on them.
Access, redaction and audit
An inspection API is a data export path into every customer conversation, so its controls matter more than its features.
- Scope by ticket. Support may read one user's sessions only while a ticket for that user is open; the policy checks the ticket system, not a header the caller can invent. Ops roles see metadata, such as kinds, timings, tool names and error codes, but no text or arguments.
- Redact at the edge of the service. The redactor masks known sensitive argument names, account numbers, email addresses and anything your privacy review lists, before data leaves the process. Redact state values by key allow-list, not deny-list.
- Audit every read. Who, which ticket, which session, which endpoint, when. Ship audit records to a store the inspection service cannot modify.
- No write paths. No run, create or delete, and no replay of a turn against a live model. Re-running a conversation belongs in an evaluation environment with copied, consented data.
- Break-glass is separate. If engineers occasionally need raw content, issue time-limited elevated access that pages a second person.
Linking timelines to traces
Timelines tell you what happened; traces tell you how long it took and what the model was sent. ADK Java's tracing sets attributes including gcp.vertex.agent.session_id, gcp.vertex.agent.invocation_id and gcp.vertex.agent.event_id on spans named invoke_agent, call_llm and execute_tool. Each timeline entry can therefore carry a deep link into your trace backend filtered by invocation ID, so the support view and the latency view refer to the same turn. The observability article covers wiring the exporter. Note that some of these span attributes carry full LLM requests, responses and tool arguments, so the trace backend needs the same access controls as the inspection API.
Worked example: the refund that never happened
Back to the refund ticket. The engineer opens the timeline for the user's session and sees, after redaction:
[
{"author":"user","kind":"message","text":"Please refund order ****4417"},
{"author":"billing_agent","kind":"tool_call",
"toolCalls":[{"name":"issue_refund","args":{"order_id":"****4417","amount":"[redacted]"}}]},
{"author":"billing_agent","kind":"tool_result","toolErrors":["issue_refund"]},
{"author":"billing_agent","kind":"final",
"text":"Your refund has been issued and should arrive in 3-5 days.",
"stateKeysChanged":["refund_status"]}
]Three facts are visible in seconds. The tool failed. The agent answered as if it had succeeded. And it wrote refund_status, so the state endpoint at the final event shows what it recorded. The trace link for that invocation shows the tool response the model received. The fix is in the agent, not the billing system: an after-tool callback that turns tool errors into an explicit failure the instructions must report, as described in the callbacks article, plus an evaluation case built from this conversation.
Failure modes
- In-memory session service in production.
InMemorySessionServiceloses everything on restart, so the inspection API finds nothing. Use a persistent session service before building inspection. - Unbounded reads. A long-running session can hold thousands of events. Bound timeline reads with
numRecentEventsand refuse full replays above a limit. - Fleet queries through listSessions. Iterating users to find errors hammers the session database. Index error events or spans in a log or trace store instead.
- Streaming fragments. Partial events duplicate text and confuse readers; filter them, and keep a toggle for engineers who need them.
- Timestamp units.
timestamp()is epoch milliseconds; treating it as seconds puts every event tens of thousands of years in the future. - Redaction drift. A new tool adds an argument containing personal data and the deny-list misses it. Allow-list argument names per tool and fail closed.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Separate read-only service | Database-enforced safety, independent scaling | One more deployment to run |
| Views instead of raw events | Stable contract, redaction in one place | Mapping code to maintain per ADK upgrade |
| State replay from deltas | State at any point in the turn | Shared prefixed keys can mislead without labelling |
| Trace links instead of copied spans | No second copy of sensitive payloads | Two systems for one investigation |
What to do next
- Confirm production uses a persistent session service and artifact service.
- Create a read-only database role and service account for a separate inspection deployment.
- Implement the event view, the redactor with per-tool argument allow-lists, and the timeline and state endpoints above.
- Wire the access policy to your ticket system and send audit records to an append-only store.
- Add trace deep links by invocation ID and apply the same access rules to the trace backend.
- Take the last five escalated tickets, walk them through the API, and turn each finding into an evaluation case. For memory-related tickets, use the memory inspection guide.