When an agent gives a wrong answer, the first question is always the same: what exactly did the model see, and what exactly did it say back? Google's Agent Development Kit for Java answers that question out of the box. Every model call is recorded as an OpenTelemetry span carrying the full request and response as JSON, and every tool call carries its arguments and result. That is enormously useful for debugging, and it is also a stream of customer text, system prompts and tool payloads flowing into whatever trace backend you connect.
This article describes what ADK Java records, read from the source at release v1.11.0 (published 2 October 2026), how to export it, how to keep the useful parts while redacting the sensitive ones, and how to join traces to a separate prompt ledger. Where the project's own README disagrees with the code, the code wins, and the article says so. For the wider observability design, metrics and dashboards, see ADK Java observability; this page is only about prompts and completions.
What ADK records, span by span
ADK obtains its tracer once, from GlobalOpenTelemetry.getTracer("gcp.vertex.agent"), and creates spans at four levels. The Runner wraps each run in an invocation span. Each agent run is invoke_agent <agent name>. Each model call is call_llm. Each tool execution is execute_tool <tool name>, and when several tools run in parallel the merged result gets an execute_tool (merged) span. A send_data span records content sent in live sessions.
One documentation trap: the telemetry README in the repository still describes tool spans as tool_call [name] and tool_response [name]. The code at v1.11.0 creates execute_tool <name> spans instead. Build dashboards and alerts on what the code emits, and verify span names against a real trace after every upgrade.
| Attribute | Span | Content |
|---|---|---|
gcp.vertex.agent.llm_request | call_llm | JSON: model, the full generation config, and every content turn |
gcp.vertex.agent.llm_response | call_llm | JSON of the whole response object, including the generated text |
gcp.vertex.agent.tool_call_args | execute_tool | JSON of the arguments the model produced |
gcp.vertex.agent.tool_response | execute_tool | JSON of what the tool returned |
gcp.vertex.agent.invocation_id, session_id, event_id | call_llm | Identifiers for joining to sessions and events |
gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens | call_llm | Model name and token usage |
gen_ai.response.finish_reasons, gen_ai.request.max_tokens, gen_ai.request.top_p | call_llm | Finish reason and sampling settings |
gen_ai.agent.name, gen_ai.conversation.id, gen_ai.tool.name | agent and tool spans | Who ran and in which session |
What is inside llm_request
The request JSON is bigger than most people expect. ADK builds it from the model name, the entire GenerateContentConfig and the list of content turns. Agent instructions are appended to the config's system instruction, so every call_llm span carries the full system prompt and the declarations of every tool the agent can call. The contents are the whole conversation history the model sees on that call, not just the latest message. Only parts carrying inline binary data, such as images, are dropped.
Three consequences follow. First, a ten-turn conversation records the early turns again on every later call, so storage grows roughly with the square of conversation length. Second, your system prompt, which may contain business rules you consider confidential, is in every trace. Third, whatever the user typed, including personal data they should not have typed, is in every trace from that point in the conversation onwards. Output tokens include reasoning tokens where the model reports them, and ADK also records the reasoning count separately as gen_ai.usage.reasoning.output_tokens.
Exporting the spans
ADK does not configure an exporter for you. You register an OpenTelemetry SDK as the global instance, and you must do it before any ADK class asks for its tracer, because ADK reads GlobalOpenTelemetry once into a static field. If ADK loads first it holds the no-op tracer, and in current OpenTelemetry Java versions a later global registration can fail outright. Make tracing set-up the first line of main.
import io.opentelemetry.api.common.AttributeKey;
import io.opentelemetry.api.common.Attributes;
import io.opentelemetry.exporter.otlp.trace.OtlpGrpcSpanExporter;
import io.opentelemetry.sdk.OpenTelemetrySdk;
import io.opentelemetry.sdk.resources.Resource;
import io.opentelemetry.sdk.trace.SdkTracerProvider;
import io.opentelemetry.sdk.trace.SpanLimits;
import io.opentelemetry.sdk.trace.export.BatchSpanProcessor;
public final class TracingBootstrap {
public static void init() {
var otlp = OtlpGrpcSpanExporter.builder()
.setEndpoint(System.getenv("OTEL_COLLECTOR_ENDPOINT")) // e.g. http://collector:4317
.build();
var provider = SdkTracerProvider.builder()
.setResource(Resource.getDefault().merge(Resource.create(
Attributes.of(AttributeKey.stringKey("service.name"), "refund-agent"))))
// Cap attribute size so one huge conversation cannot blow up a batch.
.setSpanLimits(SpanLimits.builder().setMaxAttributeValueLength(32_768).build())
.addSpanProcessor(BatchSpanProcessor.builder(new RedactingSpanExporter(otlp)).build())
.build();
OpenTelemetrySdk.builder().setTracerProvider(provider).buildAndRegisterGlobal();
Runtime.getRuntime().addShutdownHook(new Thread(provider::close));
}
}
// In main(): TracingBootstrap.init() before building any agent or Runner.The limit truncates each string attribute to its first 32,768 characters. ADK builds the request JSON from an unordered map, so do not count on which part survives; in long conversations the latest turns are the likeliest casualties. If you need the latest turn reliably, record it separately in the ledger described below rather than raising the limit without bound.
The capture switch
ADK has one built-in control: the environment variable ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS. It defaults to true. Set it to false and the four content attributes, plus the live-session data attribute, are written as the literal {}. Model name, token counts, finish reasons, identifiers and span structure all remain, so latency and cost analysis still work.
It is a process-wide, all-or-nothing switch read once at class load, so you cannot change it per agent, per tenant or per request, and you cannot keep the request while dropping the response. That makes it the right default for production services handling regulated data and the wrong tool when you need content for debugging. The usual answer is to leave capture on and redact in the export path, or to turn capture off and keep a separate, access-controlled ledger.
Redacting instead of dropping
A SpanProcessor is the wrong place to edit attributes: its onEnd hook receives a read-only span. The dependable pattern is an exporter decorator that wraps each finished SpanData in a DelegatingSpanData subclass with scrubbed attributes, then hands it to the real exporter. Redaction then runs on the batch thread, off the request path.
import io.opentelemetry.api.common.AttributeKey;
import io.opentelemetry.api.common.Attributes;
import io.opentelemetry.api.common.AttributesBuilder;
import io.opentelemetry.sdk.common.CompletableResultCode;
import io.opentelemetry.sdk.trace.data.DelegatingSpanData;
import io.opentelemetry.sdk.trace.data.SpanData;
import io.opentelemetry.sdk.trace.export.SpanExporter;
import java.util.*;
import java.util.regex.Pattern;
final class RedactingSpanExporter implements SpanExporter {
private static final List<AttributeKey<String>> CONTENT_KEYS = List.of(
AttributeKey.stringKey("gcp.vertex.agent.llm_request"),
AttributeKey.stringKey("gcp.vertex.agent.llm_response"),
AttributeKey.stringKey("gcp.vertex.agent.tool_call_args"),
AttributeKey.stringKey("gcp.vertex.agent.tool_response"),
AttributeKey.stringKey("gcp.vertex.agent.data"));
private static final Pattern EMAIL = Pattern.compile("[\\w.+-]+@[\\w-]+\\.[\\w.]+");
private static final Pattern DIGITS = Pattern.compile("\\b(?:\\d[ -]?){13,19}\\b");
private final SpanExporter delegate;
RedactingSpanExporter(SpanExporter delegate) { this.delegate = delegate; }
@Override public CompletableResultCode export(Collection<SpanData> spans) {
List<SpanData> out = new ArrayList<>(spans.size());
for (SpanData s : spans) out.add(redact(s));
return delegate.export(out);
}
private static SpanData redact(SpanData span) {
Attributes in = span.getAttributes();
AttributesBuilder b = in.toBuilder();
boolean changed = false;
for (AttributeKey<String> key : CONTENT_KEYS) {
String v = in.get(key);
if (v != null && !v.equals("{}")) { b.put(key, scrub(v)); changed = true; }
}
if (!changed) return span;
Attributes scrubbed = b.build();
return new DelegatingSpanData(span) {
@Override public Attributes getAttributes() { return scrubbed; }
};
}
static String scrub(String s) {
s = EMAIL.matcher(s).replaceAll("[email]");
return DIGITS.matcher(s).replaceAll("[number]"); // card, account, phone-like runs
}
@Override public CompletableResultCode flush() { return delegate.flush(); }
@Override public CompletableResultCode shutdown() { return delegate.shutdown(); }
}Two regular expressions are a starting point, not a data-protection programme. Put your real detectors behind scrub, test them against recorded traces, and remember that regex redaction of JSON must not break the JSON: replace values, never quotes or braces. If you run an OpenTelemetry Collector, you can redact there instead, centrally for every service; the trade-off is that raw content then crosses the network to the collector first.
A separate prompt ledger
Traces are sampled, retained for days and readable by every engineer with backend access. Some teams need the opposite for completions: every response kept, for months, for a small audited group. ADK's model callbacks give you that without touching spans. An afterModelCallbackSync receives the CallbackContext and the LlmResponse; returning Optional.empty() leaves the response unchanged.
LlmAgent agent = LlmAgent.builder()
.name("refund_agent")
.model(System.getenv("AGENT_MODEL"))
.instruction(REFUND_INSTRUCTIONS)
.tools(lookupOrder, issueRefund)
.afterModelCallbackSync((ctx, response) -> {
ledger.append(new LedgerRow(
ctx.invocationId(), // same value as gcp.vertex.agent.invocation_id
ctx.sessionId(),
ctx.agentName(),
response.content().map(Content::text).orElse(""),
response.usageMetadata().flatMap(u -> u.candidatesTokenCount()).orElse(0),
Instant.now()));
return Optional.empty(); // never alter the response from a logging hook
})
.build();Make ledger.append non-blocking: enqueue and return, because a slow write here delays the user's response. Store the ledger with its own access controls and retention, and join it to traces through invocation and session identifiers rather than copying trace IDs around.
Worked example: a wrong refund promise
A customer reports that the refund agent promised a full refund on an order outside the 30-day window. Support has the session ID. Search the trace backend for gen_ai.conversation.id equal to it and open the invocation. The tree shows invoke_agent refund_agent, a call_llm, an execute_tool lookup_order and a second call_llm.
The tool span's tool_response shows the order date correctly, 41 days ago. The second call's llm_request contains that tool result in its contents, so the model did see it. Its config shows the system instruction, and the refund rule there reads: refunds within 30 days, except for annual plans. The order was an annual plan bought 41 days ago, and the instruction never said what happens to annual plans after 30 days. The llm_response shows the model's reasonable reading of an ambiguous rule. The fix is the instruction, not the model, and you can prove it by replaying the recorded request with the corrected instruction.
Without content capture, you would have seen two model calls with normal token counts and a tool call that succeeded, and nothing more. With capture off, the ledger row for that invocation would still give you the completion, but not the instruction text. That is the trade-off in one incident.
Failure modes
| Failure | Cause | Fix |
|---|---|---|
| No ADK spans at all | SDK registered after ADK loaded its static tracer | Register the global SDK first in main |
| Dashboards empty after upgrade | Built on README span names (tool_call [...]) | Use execute_tool names; check spans after each upgrade |
| Spans orphaned or wrongly parented | Own thread pools in tools without OpenTelemetry context propagation | Propagate Context into executors; ADK handles its own RxJava hops |
| Customer data in traces for weeks | Capture defaults to true | Redacting exporter, or capture off plus ledger |
| Trace storage cost explodes | Full history on every call_llm span | Attribute length limit, sampling, redaction that drops history |
| Broken JSON in backend | Redactor replaced quotes or truncated mid-escape | Scrub values only; test on recorded spans |
| Ledger slows responses | Synchronous write in the callback | Async queue with bounded size and drop metrics |
Trade-offs and related reading
Full capture gives you the fastest debugging and replayable incidents, and the largest privacy and cost exposure. Capture off gives you safe traces for latency and cost, and nothing to debug content with. Redaction at the exporter keeps most debugging value, but only as good as your detectors. A separate ledger gives audited, long-lived completions, at the cost of a second store to secure. Most production teams combine capture on with redaction in development and staging, capture off in regulated production paths, and a ledger where completions must be kept.
Related reading: ADK Java callbacks for the full callback lattice, tool observability metrics for the numbers that sit alongside these spans, and audit logging in ADK Java for making the ledger tamper-evident.
What to do next
- Register the OpenTelemetry SDK globally as the first statement in main, then confirm a test run produces invocation, invoke_agent, call_llm and execute_tool spans.
- Open one call_llm span and read its llm_request: confirm you are comfortable with the system prompt, tool declarations and history being stored where it is.
- Decide per environment: capture on with redaction, or ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false.
- Wrap your exporter with a redacting decorator and test it against recorded spans, including JSON validity.
- Set an attribute length limit and a retention period on the trace backend.
- If completions must be kept, add an asynchronous afterModelCallback ledger keyed by invocation and session ID, with its own access control.
- After every ADK upgrade, diff span names and attribute keys against the previous version before trusting dashboards.