Sooner or later someone asks a question about an agent that only an audit log can answer. A customer says the assistant refunded the wrong order. A regulator asks who approved a data export. A security reviewer wants to know whether a prompt injection ever caused a tool to run. If all you have are application logs and a session table, the honest answer is usually that you cannot say for sure, because those records were never designed to be complete, ordered or trustworthy.

This article builds an audit trail for ADK for Java from first principles: what makes a log an audit log, which events to capture, how to capture them with a runner-wide plugin, how to make the stored record tamper-evident, and what to do when the audit write itself fails. The plugin contract in general is covered in extending the ADK Java runtime and the callback lattice in ADK Java callbacks. API details below were checked against the google/adk-java main branch in October 2026.

Advertisement

What makes a log an audit log

An audit log is a record you can rely on when you do not trust the system that produced it. That gives it four properties ordinary logs lack. It is complete: every action in scope is recorded, not a sample. It is attributable: each record names who acted, on whose behalf, under which invocation. It is ordered: you can tell what happened before what. And it is tamper-evident: if a record is altered, removed or inserted later, a check will notice.

ADK already produces two streams that look similar, and it is worth being precise about why neither is the audit log. The session event list, stored by your session service, is the conversation the agent reasons over. It is designed to be edited: state deltas, compaction and session deletion all change it, and that is correct for its purpose. Traces from OpenTelemetry, discussed in tool observability in ADK Java, are designed to be sampled, batched and expired, and they capture message content by default, which is the opposite of what you want in a long-retention store.

StreamBuilt forComplete?Mutable?Retention
Session eventsAgent contextYes, per sessionYes, by designLife of the session
OTel tracesLatency and debuggingNo, sampledNot usually, but expiredDays to weeks
Audit logAccountability and forensicsYes, in scopeNo, append-onlyPolicy-driven, often years

Decide the scope: what to record

Audit everything that changes the world or exposes data, and enough context to reconstruct why. In an agent that means: the user message arriving (as a hash, see the privacy section), each model call and which model version answered, each tool call with its arguments and outcome, every policy decision that blocked or altered something, human confirmations, errors, and the run ending.

Each record should be self-describing and stable across versions. A flat JSON schema works well:

{
  "seq": 18842,                         // per-stream sequence, assigned by the writer
  "ts": "2026-10-02T09:14:07.512Z",
  "app": "support",
  "invocation_id": "e-7f3c...",
  "session_id": "s-1029",
  "user_id": "u-5531",                  // the end user the agent acts for
  "agent": "refund_agent",
  "kind": "tool_call",                  // user_msg | model_call | tool_call | tool_result | policy | error | run_end
  "tool": "issue_refund",
  "function_call_id": "fc-03",
  "args_digest": "sha256:9b1e...",      // hash of canonical JSON args
  "args_redacted": {"order_id": "A-77812", "amount": "<redacted:amount>"},
  "outcome": "ok",
  "schema": 1,
  "prev_hash": "sha256:41ad...",
  "hash": "sha256:c07e...",
  "mac": "hmac-sha256:5f2a..."
}

Two fields deserve explanation. The args_digest lets an investigator prove that a specific argument set was used, by hashing the candidate and comparing, even when the stored copy is redacted. The prev_hash, hash and mac fields form the tamper-evidence chain, built by your writer rather than by ADK.

Advertisement

Where the audit plugin sits

The audit plugin observes every hook first, and writes to a sink the agent cannot editRunnerrunAsync(user, session)PluginManagerregistration order1. AuditPluginalways empty2. PolicyPluginmay short-circuitAgent callbacksthen model / toolBounded audit queuedrop counter if fullAuditRecordChained writerseq, prev_hash, HMACSession storemutable conversationappendEventAppend-only storeWORM / no UPDATEVerifier jobre-hash, alert on gapKey serviceHMAC key, per-user DEKskeys
The audit plugin is registered first and always returns empty, so it observes every hook even when a later plugin short-circuits. Records flow through a bounded queue to a chained writer and an append-only store.

Plugins are the right tool because they apply to every agent under a runner, including sub-agents you did not write.

Ordering is the correctness point. The PluginManager calls plugins in registration order and, for hooks that return a Maybe, stops at the first non-empty result. Plugin results are also consulted before the agent's own callbacks. If a policy plugin that blocks a tool is registered before the audit plugin, the audit plugin never sees that beforeToolCallback, and your log silently misses exactly the events an investigator cares about most. So the audit plugin goes first and must always return Maybe.empty().

Two other verified behaviours shape the design. The afterToolCallback still runs when a before-tool hook short-circuited, receiving the substituted result, so outcomes of blocked calls can be recorded there too. And onEventCallback runs after the event is appended to the session, so it observes the stored history; it is not a place to prevent anything.

The audit plugin in code

The plugin below records the main hooks. It never blocks, never throws and never does I/O on the calling thread: it builds a record and offers it to a sink.

import com.google.adk.agents.CallbackContext;
import com.google.adk.agents.InvocationContext;
import com.google.adk.models.LlmResponse;
import com.google.adk.plugins.BasePlugin;
import com.google.adk.tools.BaseTool;
import com.google.adk.tools.ToolContext;
import com.google.genai.types.Content;
import io.reactivex.rxjava3.core.Completable;
import io.reactivex.rxjava3.core.Maybe;
import java.util.Map;

/** Observes every hook; always returns empty so later plugins and callbacks still run. */
public final class AuditPlugin extends BasePlugin {
  private final AuditSink sink;          // your interface: offer() must not block or throw
  private final ArgRedactor redactor;    // per-tool allowlist of argument fields to keep

  public AuditPlugin(AuditSink sink, ArgRedactor redactor) {
    super("audit");
    this.sink = sink;
    this.redactor = redactor;
  }

  @Override
  public Maybe<Content> onUserMessageCallback(InvocationContext ctx, Content msg) {
    return observe(() -> sink.offer(AuditRecord.of(ctx, "user_msg")
        .digest(Canonical.sha256(msg))));
  }

  @Override
  public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
    return observe(() -> sink.offer(AuditRecord.of(ctx, "model_call")
        .put("model_version", resp.modelVersion().orElse("unknown"))
        .put("error_code", resp.errorCode().map(Object::toString).orElse(null))));
  }

  @Override
  public Maybe<Map<String, Object>> beforeToolCallback(
      BaseTool tool, Map<String, Object> args, ToolContext ctx) {
    return observe(() -> sink.offer(AuditRecord.of(ctx, "tool_call")
        .put("tool", tool.name())
        .put("function_call_id", ctx.functionCallId().orElse(null))
        .digest(Canonical.sha256(args))
        .put("args_redacted", redactor.redact(tool.name(), args))));
  }

  @Override
  public Maybe<Map<String, Object>> afterToolCallback(
      BaseTool tool, Map<String, Object> args, ToolContext ctx, Map<String, Object> result) {
    return observe(() -> sink.offer(AuditRecord.of(ctx, "tool_result")
        .put("tool", tool.name())
        .put("function_call_id", ctx.functionCallId().orElse(null))
        .put("outcome", result.containsKey("error") ? "error" : "ok")));
  }

  @Override
  public Maybe<Map<String, Object>> onToolErrorCallback(
      BaseTool tool, Map<String, Object> args, ToolContext ctx, Throwable error) {
    return observe(() -> sink.offer(AuditRecord.of(ctx, "error")
        .put("tool", tool.name())
        .put("error_class", error.getClass().getName())));
  }

  @Override
  public Completable afterRunCallback(InvocationContext ctx) {
    return Completable.fromAction(() -> safely(() -> sink.offer(AuditRecord.of(ctx, "run_end"))));
  }

  private static <T> Maybe<T> observe(Runnable r) {
    return Maybe.fromAction(() -> safely(r));   // completes empty: never changes behaviour
  }

  private static void safely(Runnable r) {
    try { r.run(); } catch (RuntimeException e) { AuditMetrics.pluginFailure(e); }
  }
}

AuditRecord.of pulls invocationId(), sessionId(), userId() and agentName() from the context; InvocationContext exposes the session rather than a session id accessor, so read it via ctx.session().id() there. Policy plugins should also call the same sink with a policy record when they block, so the reason is captured and not just the outcome. Register it first: Runner.builder()...plugins(new AuditPlugin(sink, redactor), new PolicyPlugin(...)).

Making the stored log tamper-evident

Append-only storage stops honest mistakes; it does not stop someone with database access from rewriting history. Tamper evidence comes from chaining: each record includes the hash of the previous one, so changing or deleting a record breaks every hash after it. A keyed MAC on top means an attacker who can write rows but does not hold the key cannot forge a valid chain from scratch. This writer runs on a single thread per stream, which is what makes the sequence and chain well defined.

final class ChainedWriter implements Runnable {
  private final BlockingQueue<AuditRecord> queue;
  private final AppendOnlyStore store;        // INSERT only; no UPDATE/DELETE grant
  private final Mac mac;                      // HmacSHA256, key from your key service
  private long seq;
  private byte[] prevHash;

  ChainedWriter(BlockingQueue<AuditRecord> q, AppendOnlyStore s, Mac mac, Checkpoint last) {
    this.queue = q; this.store = s; this.mac = mac;
    this.seq = last.seq(); this.prevHash = last.hash();   // resume the chain after restart
  }

  public void run() {
    while (!Thread.currentThread().isInterrupted()) {
      List<AuditRecord> batch = drainUpTo(queue, 500);
      List<SealedRecord> sealed = new ArrayList<>(batch.size());
      for (AuditRecord r : batch) {
        byte[] body = Canonical.json(r.withSeq(++seq).withPrev(prevHash));
        byte[] hash = sha256(body);
        sealed.add(new SealedRecord(seq, body, hash, mac.doFinal(hash)));
        prevHash = hash;
      }
      store.appendAll(sealed);   // one transaction; on failure, retry the same batch
    }
  }
}

Canonical serialisation matters, or verification produces false alarms. Sort keys, fix number formats and use UTF-8. Periodically publish the latest hash somewhere outside the database operator's control, such as a separate account's object storage with retention lock, so that even truncating the tail of the chain is detectable.

Fail open or fail closed

In the code checked, the plugin manager logs an error raised by a plugin and then re-emits it; it does not swallow it. For a before-tool hook that error takes the same path as a tool failure: the tool does not run, onToolErrorCallback handlers get a chance to substitute a result, and otherwise the error fails the call. So an audit plugin that throws turns an audit outage into a user outage, failing closed by accident on every tool, read-only ones included. That is why the observing plugin above catches everything in safely and treats audit as fail-open.

For tools in a high-risk list, such as payments, deletions and data exports, choose fail-closed explicitly and cleanly. Make the audit write for those tools synchronous and durable before the tool runs, and if it fails, return a non-empty result from beforeToolCallback that tells the model the action is unavailable. That changes behaviour, so it belongs in a small separate plugin registered right after the observing one, keeping the main audit plugin purely passive. Record an intent record before execution and an outcome record after; an intent without an outcome is itself a signal that something crashed mid-action.

  • Fail-open (the observer): bounded queue, drop and count on overflow, alert on any drop.
  • Fail-closed (high-risk tools): synchronous intent write, block on failure, alert immediately.
  • Never: an unbounded in-memory buffer, which turns an audit outage into an out-of-memory outage.

Privacy without losing traceability

Audit stores are kept for years and read by many reviewers, so keep personal data out of them. Store digests instead of content for messages. Keep an allowlist per tool of argument fields that may be stored in clear, redact the rest, and keep their digest. For content that must be retained, encrypt it with a per-user data key held in a key service, and store only ciphertext in the log.

That last step resolves the tension between an immutable log and a deletion request. You cannot delete a record from a hash chain without breaking it, but you can destroy the user's key, which renders their encrypted fields unreadable while the chain, its hashes and the non-personal fields stay verifiable. This is usually called crypto-shredding. Agree with your privacy team which fields are personal, and see PII redaction in ADK Java for detection patterns.

Worked example: reconstructing a disputed refund

A customer claims the agent refunded order A-77812 when they asked about A-77821. The investigator queries the audit store by session and walks it in sequence order. Record 18840 is a user_msg digest; hashing the transcript the customer supplied matches it, so the input is established. Record 18841 is a model_call with a model version. Record 18842 is a tool_call to issue_refund with order_id A-77812, and its args_digest matches the payments system's request log. Record 18843 is the tool_result, outcome ok.

The chain verifies from the last published checkpoint, so nobody edited these records afterwards. The conclusion: the model transposed two digits when filling the argument, and no policy check compared the order id with the orders listed earlier in the session. The fix is a grounding check in a policy plugin, not a prompt tweak.

Verification and querying

An unverified chain is decoration. Run a verifier on a schedule that recomputes each hash and MAC in sequence, checks that sequence numbers have no gaps, and compares the tail with the last externally published hash. Alert on any mismatch, and treat the alert as a security incident, not a bug ticket.

-- Gaps in the sequence: missing or deleted records
SELECT seq + 1 AS missing_from
FROM audit_log a
WHERE NOT EXISTS (SELECT 1 FROM audit_log b WHERE b.seq = a.seq + 1)
  AND seq < (SELECT max(seq) FROM audit_log);

-- Tool calls with an intent but no outcome in the last day
SELECT invocation_id, function_call_id, tool
FROM audit_log
WHERE kind = 'tool_call' AND ts > now() - interval '1 day'
EXCEPT
SELECT invocation_id, function_call_id, tool FROM audit_log WHERE kind = 'tool_result';

Keep the query role read-only and separate from the writer role.

Failure modes

  • Audit plugin registered after a blocking plugin. Blocked actions vanish from the log. Assert registration order in a startup test.
  • Blocking by throwing. A plugin error surfaces as a failed call, not a readable refusal. Block by returning a value, and catch everything in observers.
  • Raw content in the audit store. The store becomes the largest personal data repository you own. Digest, allowlist, encrypt.
  • Multiple writers on one chain. Interleaved sequence numbers and broken hashes. Use one writer per stream, or one chain per partition.
  • No external checkpoint. Truncating the newest records is undetectable from inside the database.
  • Silent drops. A full queue that drops without a metric hides an outage for weeks.

What to do next

  1. Write down the audit scope: which tools are high-risk, which fields are personal, and the retention period required.
  2. Implement a passive audit plugin on the hooks above and register it first; add a startup test that fails if it is not first.
  3. Add a separate fail-closed plugin for high-risk tools that writes a durable intent record before execution.
  4. Build the chained writer with canonical JSON, HMAC and a store whose role has INSERT but not UPDATE or DELETE.
  5. Publish the latest hash to an independent, retention-locked location on a schedule.
  6. Schedule the verifier and the intent-without-outcome query, with alerts routed to security.
  7. Rehearse one investigation end to end, from a user complaint to a verified sequence of records.
Key takeaway: An audit log is complete, attributable, ordered and tamper-evident, which neither ADK session events nor traces are. Build it as the first-registered, always-empty plugin so it sees every hook, write through a single chained HMAC writer to append-only storage, and publish checkpoints externally. Because a throwing plugin fails the call, catch everything in the observer, decide fail-open versus fail-closed per tool, and block by returning a value. Keep personal data out with digests, allowlists and per-user keys, and verify the chain on a schedule.