Sooner or later someone asks a question about an agent that only an audit log can answer. A customer says the assistant refunded the wrong order. A regulator asks who approved a data export. A security reviewer wants to know whether a prompt injection ever caused a tool to run. If all you have are application logs and a session table, the honest answer is usually that you cannot say for sure, because those records were never designed to be complete, ordered or trustworthy.
This article builds an audit trail for ADK for Java from first principles: what makes a log an audit log, which events to capture, how to capture them with a runner-wide plugin, how to make the stored record tamper-evident, and what to do when the audit write itself fails. The plugin contract in general is covered in extending the ADK Java runtime and the callback lattice in ADK Java callbacks. API details below were checked against the google/adk-java main branch in October 2026.
What makes a log an audit log
An audit log is a record you can rely on when you do not trust the system that produced it. That gives it four properties ordinary logs lack. It is complete: every action in scope is recorded, not a sample. It is attributable: each record names who acted, on whose behalf, under which invocation. It is ordered: you can tell what happened before what. And it is tamper-evident: if a record is altered, removed or inserted later, a check will notice.
ADK already produces two streams that look similar, and it is worth being precise about why neither is the audit log. The session event list, stored by your session service, is the conversation the agent reasons over. It is designed to be edited: state deltas, compaction and session deletion all change it, and that is correct for its purpose. Traces from OpenTelemetry, discussed in tool observability in ADK Java, are designed to be sampled, batched and expired, and they capture message content by default, which is the opposite of what you want in a long-retention store.
| Stream | Built for | Complete? | Mutable? | Retention |
|---|---|---|---|---|
| Session events | Agent context | Yes, per session | Yes, by design | Life of the session |
| OTel traces | Latency and debugging | No, sampled | Not usually, but expired | Days to weeks |
| Audit log | Accountability and forensics | Yes, in scope | No, append-only | Policy-driven, often years |
Decide the scope: what to record
Audit everything that changes the world or exposes data, and enough context to reconstruct why. In an agent that means: the user message arriving (as a hash, see the privacy section), each model call and which model version answered, each tool call with its arguments and outcome, every policy decision that blocked or altered something, human confirmations, errors, and the run ending.
Each record should be self-describing and stable across versions. A flat JSON schema works well:
{
"seq": 18842, // per-stream sequence, assigned by the writer
"ts": "2026-10-02T09:14:07.512Z",
"app": "support",
"invocation_id": "e-7f3c...",
"session_id": "s-1029",
"user_id": "u-5531", // the end user the agent acts for
"agent": "refund_agent",
"kind": "tool_call", // user_msg | model_call | tool_call | tool_result | policy | error | run_end
"tool": "issue_refund",
"function_call_id": "fc-03",
"args_digest": "sha256:9b1e...", // hash of canonical JSON args
"args_redacted": {"order_id": "A-77812", "amount": "<redacted:amount>"},
"outcome": "ok",
"schema": 1,
"prev_hash": "sha256:41ad...",
"hash": "sha256:c07e...",
"mac": "hmac-sha256:5f2a..."
}Two fields deserve explanation. The args_digest lets an investigator prove that a specific argument set was used, by hashing the candidate and comparing, even when the stored copy is redacted. The prev_hash, hash and mac fields form the tamper-evidence chain, built by your writer rather than by ADK.
Where the audit plugin sits
Plugins are the right tool because they apply to every agent under a runner, including sub-agents you did not write.
Ordering is the correctness point. The PluginManager calls plugins in registration order and, for hooks that return a Maybe, stops at the first non-empty result. Plugin results are also consulted before the agent's own callbacks. If a policy plugin that blocks a tool is registered before the audit plugin, the audit plugin never sees that beforeToolCallback, and your log silently misses exactly the events an investigator cares about most. So the audit plugin goes first and must always return Maybe.empty().
Two other verified behaviours shape the design. The afterToolCallback still runs when a before-tool hook short-circuited, receiving the substituted result, so outcomes of blocked calls can be recorded there too. And onEventCallback runs after the event is appended to the session, so it observes the stored history; it is not a place to prevent anything.
The audit plugin in code
The plugin below records the main hooks. It never blocks, never throws and never does I/O on the calling thread: it builds a record and offers it to a sink.
import com.google.adk.agents.CallbackContext;
import com.google.adk.agents.InvocationContext;
import com.google.adk.models.LlmResponse;
import com.google.adk.plugins.BasePlugin;
import com.google.adk.tools.BaseTool;
import com.google.adk.tools.ToolContext;
import com.google.genai.types.Content;
import io.reactivex.rxjava3.core.Completable;
import io.reactivex.rxjava3.core.Maybe;
import java.util.Map;
/** Observes every hook; always returns empty so later plugins and callbacks still run. */
public final class AuditPlugin extends BasePlugin {
private final AuditSink sink; // your interface: offer() must not block or throw
private final ArgRedactor redactor; // per-tool allowlist of argument fields to keep
public AuditPlugin(AuditSink sink, ArgRedactor redactor) {
super("audit");
this.sink = sink;
this.redactor = redactor;
}
@Override
public Maybe<Content> onUserMessageCallback(InvocationContext ctx, Content msg) {
return observe(() -> sink.offer(AuditRecord.of(ctx, "user_msg")
.digest(Canonical.sha256(msg))));
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
return observe(() -> sink.offer(AuditRecord.of(ctx, "model_call")
.put("model_version", resp.modelVersion().orElse("unknown"))
.put("error_code", resp.errorCode().map(Object::toString).orElse(null))));
}
@Override
public Maybe<Map<String, Object>> beforeToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx) {
return observe(() -> sink.offer(AuditRecord.of(ctx, "tool_call")
.put("tool", tool.name())
.put("function_call_id", ctx.functionCallId().orElse(null))
.digest(Canonical.sha256(args))
.put("args_redacted", redactor.redact(tool.name(), args))));
}
@Override
public Maybe<Map<String, Object>> afterToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx, Map<String, Object> result) {
return observe(() -> sink.offer(AuditRecord.of(ctx, "tool_result")
.put("tool", tool.name())
.put("function_call_id", ctx.functionCallId().orElse(null))
.put("outcome", result.containsKey("error") ? "error" : "ok")));
}
@Override
public Maybe<Map<String, Object>> onToolErrorCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx, Throwable error) {
return observe(() -> sink.offer(AuditRecord.of(ctx, "error")
.put("tool", tool.name())
.put("error_class", error.getClass().getName())));
}
@Override
public Completable afterRunCallback(InvocationContext ctx) {
return Completable.fromAction(() -> safely(() -> sink.offer(AuditRecord.of(ctx, "run_end"))));
}
private static <T> Maybe<T> observe(Runnable r) {
return Maybe.fromAction(() -> safely(r)); // completes empty: never changes behaviour
}
private static void safely(Runnable r) {
try { r.run(); } catch (RuntimeException e) { AuditMetrics.pluginFailure(e); }
}
}AuditRecord.of pulls invocationId(), sessionId(), userId() and agentName() from the context; InvocationContext exposes the session rather than a session id accessor, so read it via ctx.session().id() there. Policy plugins should also call the same sink with a policy record when they block, so the reason is captured and not just the outcome. Register it first: Runner.builder()...plugins(new AuditPlugin(sink, redactor), new PolicyPlugin(...)).
Making the stored log tamper-evident
Append-only storage stops honest mistakes; it does not stop someone with database access from rewriting history. Tamper evidence comes from chaining: each record includes the hash of the previous one, so changing or deleting a record breaks every hash after it. A keyed MAC on top means an attacker who can write rows but does not hold the key cannot forge a valid chain from scratch. This writer runs on a single thread per stream, which is what makes the sequence and chain well defined.
final class ChainedWriter implements Runnable {
private final BlockingQueue<AuditRecord> queue;
private final AppendOnlyStore store; // INSERT only; no UPDATE/DELETE grant
private final Mac mac; // HmacSHA256, key from your key service
private long seq;
private byte[] prevHash;
ChainedWriter(BlockingQueue<AuditRecord> q, AppendOnlyStore s, Mac mac, Checkpoint last) {
this.queue = q; this.store = s; this.mac = mac;
this.seq = last.seq(); this.prevHash = last.hash(); // resume the chain after restart
}
public void run() {
while (!Thread.currentThread().isInterrupted()) {
List<AuditRecord> batch = drainUpTo(queue, 500);
List<SealedRecord> sealed = new ArrayList<>(batch.size());
for (AuditRecord r : batch) {
byte[] body = Canonical.json(r.withSeq(++seq).withPrev(prevHash));
byte[] hash = sha256(body);
sealed.add(new SealedRecord(seq, body, hash, mac.doFinal(hash)));
prevHash = hash;
}
store.appendAll(sealed); // one transaction; on failure, retry the same batch
}
}
}Canonical serialisation matters, or verification produces false alarms. Sort keys, fix number formats and use UTF-8. Periodically publish the latest hash somewhere outside the database operator's control, such as a separate account's object storage with retention lock, so that even truncating the tail of the chain is detectable.
Fail open or fail closed
In the code checked, the plugin manager logs an error raised by a plugin and then re-emits it; it does not swallow it. For a before-tool hook that error takes the same path as a tool failure: the tool does not run, onToolErrorCallback handlers get a chance to substitute a result, and otherwise the error fails the call. So an audit plugin that throws turns an audit outage into a user outage, failing closed by accident on every tool, read-only ones included. That is why the observing plugin above catches everything in safely and treats audit as fail-open.
For tools in a high-risk list, such as payments, deletions and data exports, choose fail-closed explicitly and cleanly. Make the audit write for those tools synchronous and durable before the tool runs, and if it fails, return a non-empty result from beforeToolCallback that tells the model the action is unavailable. That changes behaviour, so it belongs in a small separate plugin registered right after the observing one, keeping the main audit plugin purely passive. Record an intent record before execution and an outcome record after; an intent without an outcome is itself a signal that something crashed mid-action.
- Fail-open (the observer): bounded queue, drop and count on overflow, alert on any drop.
- Fail-closed (high-risk tools): synchronous intent write, block on failure, alert immediately.
- Never: an unbounded in-memory buffer, which turns an audit outage into an out-of-memory outage.
Privacy without losing traceability
Audit stores are kept for years and read by many reviewers, so keep personal data out of them. Store digests instead of content for messages. Keep an allowlist per tool of argument fields that may be stored in clear, redact the rest, and keep their digest. For content that must be retained, encrypt it with a per-user data key held in a key service, and store only ciphertext in the log.
That last step resolves the tension between an immutable log and a deletion request. You cannot delete a record from a hash chain without breaking it, but you can destroy the user's key, which renders their encrypted fields unreadable while the chain, its hashes and the non-personal fields stay verifiable. This is usually called crypto-shredding. Agree with your privacy team which fields are personal, and see PII redaction in ADK Java for detection patterns.
Worked example: reconstructing a disputed refund
A customer claims the agent refunded order A-77812 when they asked about A-77821. The investigator queries the audit store by session and walks it in sequence order. Record 18840 is a user_msg digest; hashing the transcript the customer supplied matches it, so the input is established. Record 18841 is a model_call with a model version. Record 18842 is a tool_call to issue_refund with order_id A-77812, and its args_digest matches the payments system's request log. Record 18843 is the tool_result, outcome ok.
The chain verifies from the last published checkpoint, so nobody edited these records afterwards. The conclusion: the model transposed two digits when filling the argument, and no policy check compared the order id with the orders listed earlier in the session. The fix is a grounding check in a policy plugin, not a prompt tweak.
Verification and querying
An unverified chain is decoration. Run a verifier on a schedule that recomputes each hash and MAC in sequence, checks that sequence numbers have no gaps, and compares the tail with the last externally published hash. Alert on any mismatch, and treat the alert as a security incident, not a bug ticket.
-- Gaps in the sequence: missing or deleted records
SELECT seq + 1 AS missing_from
FROM audit_log a
WHERE NOT EXISTS (SELECT 1 FROM audit_log b WHERE b.seq = a.seq + 1)
AND seq < (SELECT max(seq) FROM audit_log);
-- Tool calls with an intent but no outcome in the last day
SELECT invocation_id, function_call_id, tool
FROM audit_log
WHERE kind = 'tool_call' AND ts > now() - interval '1 day'
EXCEPT
SELECT invocation_id, function_call_id, tool FROM audit_log WHERE kind = 'tool_result';Keep the query role read-only and separate from the writer role.
Failure modes
- Audit plugin registered after a blocking plugin. Blocked actions vanish from the log. Assert registration order in a startup test.
- Blocking by throwing. A plugin error surfaces as a failed call, not a readable refusal. Block by returning a value, and catch everything in observers.
- Raw content in the audit store. The store becomes the largest personal data repository you own. Digest, allowlist, encrypt.
- Multiple writers on one chain. Interleaved sequence numbers and broken hashes. Use one writer per stream, or one chain per partition.
- No external checkpoint. Truncating the newest records is undetectable from inside the database.
- Silent drops. A full queue that drops without a metric hides an outage for weeks.
What to do next
- Write down the audit scope: which tools are high-risk, which fields are personal, and the retention period required.
- Implement a passive audit plugin on the hooks above and register it first; add a startup test that fails if it is not first.
- Add a separate fail-closed plugin for high-risk tools that writes a durable intent record before execution.
- Build the chained writer with canonical JSON, HMAC and a store whose role has INSERT but not UPDATE or DELETE.
- Publish the latest hash to an independent, retention-locked location on a schedule.
- Schedule the verifier and the intent-without-outcome query, with alerts routed to security.
- Rehearse one investigation end to end, from a user complaint to a verified sequence of records.