The instruction string you pass to LlmAgent.builder().instruction(...) is not the prompt your model receives. By the time a request leaves an ADK Java agent, session state has been substituted into placeholders, the conversation history has been assembled (and possibly filtered), every tool's declaration has been attached, plugins such as GlobalInstructionPlugin may have added text, and a callback may have rewritten the whole thing. When an answer goes wrong, the question that matters is: what exactly did the model see?

Prompt logging answers that question by recording the assembled request of every model call. This article builds it as a Runner plugin on the beforeModelCallback hook in ADK Java 1.11.0, stores it content-addressed so a 20-turn session costs a few tens of kilobytes instead of hundreds, fingerprints each request so you can see when a deploy changed the prompt, and rebuilds logged requests so you can replay them against a model. It covers the request side; tracing spans and completion ledgers are covered in a separate article linked below.

Why the source code is not the prompt

Three gaps make the source code an unreliable guide to the prompt. First, templating: an instruction such as You help {user:tier} customers is resolved from session state per call, so two users get different system text from the same code. Second, assembly: the contents list is built from session events, and anything that trims or summarises history changes what the model saw on turn 14 compared with turn 2. Third, tools: the function declarations generated from your Java method signatures and @Schema descriptions travel with every request, and a one-word change in a description can change which tool the model picks.

Debugging without the assembled request means reconstructing it from code at a past commit, session events and state, which is slow and often impossible because state has moved on. With the request logged, you read it, and you can send it again.

What ADK Java already records

Check what ships before building. ADK Java 1.11.0 includes LoggingPlugin (plugin name logging_plugin) in com.google.adk.plugins. Reading its bytecode, the LLM REQUEST entry it writes through SLF4J contains the model name, the agent name, the first system instruction cut at 200 characters, and the names of the available tools. It does not write the contents list or the tool declarations. That makes it a good development aid for following the control flow, and not a prompt log: you cannot reproduce a call from it.

ADK also emits OpenTelemetry spans for model calls whose attributes can carry the request; the tracing article covers what they hold. Spans are usually sampled, retained for days and sized for a trace backend, which is fine for latency work and awkward for a store of complete prompts. A dedicated prompt log is worth it when you need every call (or a deliberate sample) kept in full, queryable by prompt version, with its own access controls.

Where to capture: the before-model hook

Request-side prompt log: capture, split, address, replayRunnerplugins in orderRedaction pluginrewrites the builderPromptLogPluginbuild() snapshotModel callBaseLlmSplitter + hashersha-256 per partbounded queueBlob storeinstruction, tools,config, contentsCall indexrefs + fingerprintOutcome rowsafter / errorafterModelCallbackReplay toolrebuild LlmRequestDrift reportnew fingerprintsBlobs are written once and referenced by every call that sent the same bytes.
Redaction runs first, the logging plugin snapshots the built request, a worker splits and hashes it, and replay rebuilds requests from blobs.

The Plugin interface gives you three hooks around a model call:

Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder request);
Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse response);
Maybe<LlmResponse> onModelErrorCallback(CallbackContext ctx, LlmRequest.Builder request, Throwable error);

The before hook receives the builder after ADK's request processors have run, so instructions, contents and tool declarations are in place. Calling build() gives an immutable LlmRequest snapshot; the built-in LoggingPlugin does exactly this. Returning Maybe.empty() lets the call proceed unchanged. Returning a response would short-circuit the model, which a logger must never do.

Placement matters in two ways. Plugins run before agent-level callbacks, so if an agent's own beforeModelCallback edits the request, a plugin logs the version before that edit; keep request-rewriting logic in plugins if you want the log to match the wire. And plugins run in registration order, so register the redaction plugin before the logging plugin. Then the log holds what was actually sent, already redacted, and you never store a raw copy you would later have to hunt down.

Content-addressed storage

Most of a prompt repeats. The system instruction is the same for every call an agent version makes (unless it is templated per user), the tool declarations are the same, and each turn's contents list is the previous turn's list plus a few new entries. Logging each request as one JSON document stores the same bytes over and over.

Split each request into parts, hash each part with SHA-256, write each blob once, and log a small call row that references the hashes:

PartSourceTypical reuse
instructionconfig.systemInstruction()every call of that agent version
toolsconfig.tools()every call of that agent version
configconfig JSON minus the two fields abovealmost every call
content[i]each entry of contents()every later call in the session

A worked size estimate: an agent with a 12 KB instruction and 4 KB of tool declarations runs a 20-turn session in which each turn adds about 1.5 KB of contents. Logged whole, call k carries 16 KB plus 1.5k KB of history, so 20 calls store 320 KB of repeated headers plus 315 KB of history: about 635 KB. Content-addressed, the session stores the 16 KB header once and each content entry once, about 46 KB, plus twenty call rows of a few hundred bytes. That is roughly a fourteen-fold saving, and it grows when tool loops make several model calls per turn.

One privacy adjustment: hash user-bearing contents with the session ID as a salt, so that identical text from two users produces two blobs. Deleting a user's data then means deleting their sessions' blobs without checking whether anyone else references them. Instruction, tool and config blobs carry no user data and stay shared.

The plugin and the worker

The plugin does the minimum on the request thread: build the snapshot, enqueue it, return. Serialisation, hashing and storage happen on a worker, and a full queue drops the record and counts it rather than slowing the user's response.

public final class PromptLogPlugin extends BasePlugin {
  private final BlockingQueue<Captured> queue = new ArrayBlockingQueue<>(10_000);
  private final Map<String, AtomicInteger> callSeq = new ConcurrentHashMap<>();
  private final Counter dropped;

  public PromptLogPlugin(MeterRegistry meters) {
    super("prompt_log");
    this.dropped = meters.counter("prompt_log.dropped");
  }

  @Override
  public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder b) {
    return Maybe.fromAction(() -> {
      int seq = callSeq.computeIfAbsent(ctx.invocationId(), k -> new AtomicInteger())
                       .incrementAndGet();
      Captured c = new Captured(ctx.sessionId(), ctx.invocationId(), seq,
                                ctx.agentName(), b.build(), Instant.now());
      if (!queue.offer(c)) dropped.increment();
    });
  }

  @Override
  public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse r) {
    return Maybe.fromAction(() -> outcomes.record(ctx.invocationId(),
        callSeq.getOrDefault(ctx.invocationId(), new AtomicInteger()).get(),
        r.usageMetadata().flatMap(u -> u.promptTokenCount()).orElse(-1),
        r.errorCode().map(Object::toString).orElse("OK")));
  }
}

Calls are keyed by invocation ID plus a per-invocation sequence number, because one invocation can make several model calls and the after hook needs to find its request. This pairing assumes calls within an invocation are sequential, which holds for a single agent's tool loop; with parallel sub-agents, add the agent name to the key. Evict callSeq entries in an afterRunCallback so the map does not grow forever. The onModelErrorCallback should record the exception class against the same key: failed calls are the ones you most want to inspect.

On the worker, split the snapshot:

void persist(Captured c) {
  LlmRequest req = c.request();
  ObjectNode cfg = (ObjectNode) mapper.readTree(req.config().map(x -> x.toJson()).orElse("{}"));
  JsonNode instruction = Objects.requireNonNullElse(cfg.remove("systemInstruction"), NullNode.instance);
  JsonNode tools = Objects.requireNonNullElse(cfg.remove("tools"), NullNode.instance);

  String hInstr = blobs.putShared(canonical(instruction));
  String hTools = blobs.putShared(canonical(tools));
  String hCfg   = blobs.putShared(canonical(cfg));
  List<String> hContents = req.contents().stream()
      .map(content -> blobs.putSalted(c.sessionId(), canonical(content.toJson())))
      .toList();

  String fingerprint = sha256(req.model().orElse("") + hInstr + hTools + hCfg);
  index.insert(new CallRow(c.sessionId(), c.invocationId(), c.seq(), c.agent(),
      req.model().orElse(""), fingerprint, hInstr, hTools, hCfg, hContents, c.at()));
}

canonical re-serialises JSON with sorted keys so that equal content always hashes equal. putShared and putSalted are conditional inserts that do nothing when the hash already exists, which also makes the worker safe to retry.

Prompt fingerprints

The fingerprint is a hash of model name, instruction, tools and config: everything about the request except the conversation. Two calls with the same fingerprint were made by the same prompt version. That one column answers questions that are otherwise guesswork.

  • Did the deploy change the prompt? Count distinct fingerprints per agent per hour; a deploy that should not touch prompts and still produces a new fingerprint has changed a tool description, a dependency's default, or a config value.
  • Which version produced this complaint? Look up the call row by session and invocation, read the fingerprint, and diff its blobs against the current version's.
  • Is per-user data leaking into the system prompt? If one agent shows thousands of instruction hashes per day, a state placeholder is putting per-user values into the instruction. That defeats model-side context caching and is worth moving into contents.

Replaying a logged request

A logged request can be rebuilt and sent again. This is how you reproduce a bad answer, check whether a fix changes it, or run the same prompt against a candidate model.

LlmRequest rebuild(CallRow row, String targetModel) {
  ObjectNode cfg = (ObjectNode) mapper.readTree(blobs.get(row.hCfg()));
  cfg.set("systemInstruction", mapper.readTree(blobs.get(row.hInstr())));
  cfg.set("tools", mapper.readTree(blobs.get(row.hTools())));
  List<Content> contents = row.hContents().stream()
      .map(h -> Content.fromJson(blobs.get(h))).toList();
  return LlmRequest.builder()
      .model(targetModel)                 // the request's name wins over the BaseLlm's
      .contents(contents)
      .config(GenerateContentConfig.fromJson(cfg.toString()))
      .build();
}

LlmResponse replay(CallRow row, BaseLlm model) {
  return model.generateContent(rebuild(row, model.model()), false).blockingLast();
}

The function declarations travel inside the config, which is what the model needs; the replayed request does not execute tools, so you see the model's decision (text or a function call) without side effects. Set the model name from the target BaseLlm, not from the log: ADK's Gemini class prefers the name inside the request, so copying the logged name would silently replay against the original model. Expect variation: at a non-zero temperature, run the replay ten times and report how often the bad behaviour recurs instead of judging from one sample. If your build of the builder requires a tools map, set an empty one.

Worked example: a tool description that changed itself

A support agent starts telling customers that express shipping is free. Nothing in the instruction says so, and the change coincides with a release that only upgraded a dependency.

Query the call index for the agent: fingerprint 9f3c... served every call until Tuesday 14:02, and 41ab... from 14:05. The instruction hashes are identical and the tools hashes differ. Diffing the two tools blobs shows that the declaration for get_shipping_quote lost the sentence "Returns the price in cents; 0 means the option is unavailable" because the upgraded library now reads descriptions from a different annotation attribute. The model, seeing a price of 0 with no explanation, says free.

Replaying five logged calls with the old tools blob gives correct answers 50 times out of 50; with the new blob, wrong answers 31 times out of 50. The fix is a one-line annotation change, and the fingerprint alert that fired at 14:05 becomes a release check: any new fingerprint must be approved.

Operating the log

  • Sampling. Log every call for low-volume or regulated agents. For high volume, sample whole sessions, not calls, so replay has complete history; always log calls that end in an error or a safety block.
  • Retention tiers. Keep call rows and shared blobs long, because they are small and describe versions. Keep salted content blobs on the session's retention schedule and delete them with it.
  • Access. The content blobs are transcripts. Store them in their own bucket or table with narrower read access than application logs, and log reads.
  • Size limits. Inline images and files in contents can be megabytes; store large parts by reference to the artifact store instead of copying bytes.

Failure modes

  • Logging before redaction. Registering the logger first stores raw personal data. Assert plugin order at startup.
  • Blocking the request thread. Hashing and uploading inside the callback adds tens of milliseconds to every model call. Enqueue and return.
  • Silent drops. A full queue loses records. Export the drop counter and alert on it, or replay finds holes exactly where load was highest.
  • Non-canonical JSON. Hashing serialiser output with unstable key order gives new hashes for equal content and quietly destroys the saving.
  • Logging the agent callback's input. If an agent-level callback rewrites the request after the plugin ran, the log shows a prompt that was never sent.
  • Treating replay as proof. A replay against today's model endpoint may not match a past answer exactly; compare distributions, not single outputs.

Trade-offs

Full prompt logging costs storage, a worker and a data store holding transcripts that must be governed. Content addressing removes most of the storage cost, and fingerprints turn the log into a version history you did not otherwise have. The alternatives, LoggingPlugin and sampled spans, are cheaper and sufficient for development and latency work; choose the full log when you need to reproduce production behaviour after the fact, prove what a model was told, or evaluate a model change on real prompts.

Related reading: what ADK's spans record for prompts and completions, a redaction plugin and where each hook applies, making a log tamper-evident and the callback model in ADK Java.

What to do next

  1. Register LoggingPlugin in a development environment and read one session's output to see what it records and what it leaves out.
  2. Write PromptLogPlugin with a bounded queue, a drop counter and invocation-plus-sequence keys; register it after your redaction plugin and assert the order.
  3. Split requests into instruction, tools, config and content blobs; canonicalise JSON before hashing; salt content hashes with the session ID.
  4. Add the fingerprint column and a daily report of new fingerprints per agent; make an unexpected one block the release.
  5. Build the replay function and reproduce one real past answer ten times.
  6. Set retention and access for content blobs to match your session data, and test deletion for one user end to end.
Key takeaway: What your agent's model saw is the assembled request, not your instruction string. ADK Java's LoggingPlugin records only a 200-character instruction excerpt and tool names, so capture the built LlmRequest yourself in a plugin's beforeModelCallback, after redaction, off the request thread. Store it as content-addressed blobs, give every call a fingerprint of its non-conversation parts, and keep a replay function: together they turn wrong answers into reproducible, diffable cases.