An audit log for an LLM application has one job: months later, someone asks what the system did and why, and the log answers without guesswork. Who asked, which model answered with which instructions and which retrieved documents, what text the user actually saw, and which actions the application took as a result. Most teams discover their log cannot answer that during the first serious incident, because it was designed as debug telemetry: sampled, missing streamed output, blind to retries, and unable to tie a tool call back to the prompt that caused it.

This page is about the capture side, the part that decides whether the record is complete. It covers where to capture, how to correlate events across agent steps, delivery guarantees, streaming and retries, and how to prove nothing is missing. The event schema, content vault and crypto-shredding are covered in LLM audit logs and tamper-evident storage in LLM audit logging architecture.

Where audit events come from

User or callerauthenticatedApplicationidentity, intent, decisionsLLM gatewaymodel I/O, attempts, streamrequest idModel providerreturns model idRetrieverdoc ids, versionsTool executorintent, then outcometool callLocal durable spool per producerevent id, producer sequence number, fsync before ackapp eventsgateway eventsCollector and audit storededupe by event idReconcilersequence gaps, counts vs provider usageEach producer records what only it can see; correlation ids tie the records into one decision.
Producers write to local durable spools; a collector deduplicates by event id and a reconciler checks sequence gaps and provider usage counts.

Start from the questions, not the schema

Design backwards from the questions. A complaint, a regulator, a security review and an internal incident ask variations of the same five: who initiated this, under what authority, what did the model receive, what did it produce and what was delivered, and what happened in the world as a consequence. Telemetry answers how the system is performing in aggregate, so it can be sampled and dropped under load. An audit log answers what happened in one specific case, so it cannot be sampled, and a missing record is a defect rather than noise.

That difference drives everything below. Each requirement for an audit trail is a completeness requirement: every attempt, every streamed token that reached a user, every tool side effect, every model version, all linked. Storage format matters less than whether the capture points can see the facts at all.

Choosing capture points

No single component sees everything, so audit events come from several producers, and each is authoritative for what it alone observes.

Capture pointSees reliablyMisses
Application codeauthenticated user, tenant, feature, business decision taken on the outputretries inside SDKs, the exact bytes sent, provider-side model resolution
LLM gateway or proxyexact request and response, every attempt, fallbacks, token usage, model id returnedwho the end user is unless propagated, what the app did with the answer
Tool executorthe side effect actually performed, its arguments and resultwhy the model asked for it
Retrieverdocument ids, versions and scores placed in contexthow the model used them
Provider-side logsbilling-grade request countsyour identities; retention and access are the provider's terms

The usual mistake is to log only in the application, after the SDK call returns. That captures the final answer but not the two failed attempts before it, the fallback to a different model, or the fact that the stream was cut halfway. Route all model traffic through one gateway (a proxy service or a single client library that every team must use), make the gateway the authoritative recorder of model I/O, and have the application record identity and decisions, joined by a request id the gateway refuses to proceed without.

Correlating an agent's steps

An agentic request is a tree, not a line: the user turn spawns a planning call, which requests two tools, one of which runs a retrieval and a second model call. To reconstruct it you need a request_id for the whole interaction and a step_id with a parent_step_id for every node, propagated through headers on every internal hop and into tool executors and sub-agents. If you already run distributed tracing, reuse its trace and span ids, but write audit events to a separate pipeline: trace backends sample and expire, and an audit store must do neither.

from dataclasses import dataclass, field, asdict
import hashlib, json, time, uuid

@dataclass
class AuditEvent:
    kind: str                     # "request", "model_attempt", "tool_intent", "tool_outcome", "delivery"
    request_id: str               # one per user-visible interaction, minted at the edge
    step_id: str                  # this step; agent steps form a tree
    parent_step_id: str | None
    producer: str                 # "app", "gateway", "tool-executor"
    producer_seq: int             # strictly increasing per producer instance, for gap detection
    actor: dict                   # {"user": ..., "tenant": ..., "on_behalf_of": ...}
    payload: dict                 # kind-specific; large content stored by reference
    event_id: str = field(default_factory=lambda: uuid.uuid4().hex)
    ts_ms: int = field(default_factory=lambda: int(time.time() * 1000))

def content_ref(vault, text: str) -> dict:
    """Store content in the access-controlled vault; the audit event keeps a hash and pointer."""
    digest = hashlib.sha256(text.encode("utf-8")).hexdigest()
    vault.put(digest, text)               # idempotent: same text, same key
    return {"sha256": digest, "chars": len(text)}

Note what the envelope does not contain: raw prompt and output text. Content goes to a separate vault with tighter access and its own deletion rules, and the event carries a hash and pointer. That keeps the audit store safe to query broadly and makes erasure requests tractable without breaking the event chain; the mechanics are on the audit logs page.

Delivery guarantees and the order of writes

Audit delivery must be at-least-once with deduplication, and the ordering of writes relative to actions matters. Each producer appends events to a local durable spool (a write-ahead file or an outbox table in the same database transaction as the business change) and acknowledges only after the write is durable. A shipper forwards spooled events to the collector, retries until acknowledged, and the store deduplicates on event_id. If the collector is down, events accumulate locally rather than vanishing.

For tool calls with side effects, record intent before acting and outcome after: write a tool_intent event with the arguments, perform the action, then write tool_outcome. An intent without an outcome is itself evidence: the process crashed mid-action and someone must check whether the refund went through. If the intent cannot be recorded, a high-impact action should not run; read-only chat can reasonably fail open. That policy choice is discussed with a working plugin in ADK Java audit logging.

Streaming, retries and what the user actually saw

Streaming breaks the naive log-the-response approach because there may never be a complete response. The user may close the tab after forty tokens, an output filter may cut the stream mid-sentence, or the connection may drop. What matters for an audit is what the user actually saw, so the gateway must accumulate the delivered text and write it in a finally block that runs on success, error and cancellation, together with a finish reason that distinguishes a natural stop, a length cut-off, a filter block, a disconnect and an error.

async def audited_stream(gateway, audit, req, attempt_no):
    delivered, finish = [], "unknown"
    audit.emit("model_attempt_start", req, attempt=attempt_no, model_requested=req.model,
               params={"temperature": req.temperature, "max_tokens": req.max_tokens},
               prompt=content_ref(audit.vault, req.rendered_prompt),
               prompt_template=req.template_id, template_version=req.template_version)
    try:
        async for chunk in gateway.stream(req):
            if chunk.blocked_by_filter:
                finish = "filtered"
                break
            delivered.append(chunk.text)
            yield chunk.text                        # forwarded to the user
        else:
            finish = chunk.finish_reason            # e.g. "stop" or "length"
    except (asyncio.CancelledError, GeneratorExit):   # task cancelled, or consumer stopped reading
        finish = "client_disconnected"
        raise
    except Exception as e:
        finish = f"error:{type(e).__name__}"
        raise
    finally:                                        # runs on success, error and cancellation
        audit.emit("model_attempt_end", req, attempt=attempt_no, finish=finish,
                   model_returned=gateway.last_model_id,   # the provider's resolved version
                   delivered=content_ref(audit.vault, "".join(delivered)),
                   usage=gateway.last_usage)

Retries and fallbacks get the same treatment: every attempt is its own pair of events with an attempt number, and a final delivery event says which attempt's output reached the user. Without that, a retried request with a partial first stream looks identical to a clean one, and as the worked example below shows, the partial stream can be exactly the part that matters.

The decision record

To explain a decision you need the inputs that shaped it, not only the text in and out. The decision record attached to each model attempt should include:

  • The model id the provider returned, not the alias you requested. Aliases move to new versions; the returned identifier is what ran.
  • Sampling parameters: temperature, top-p, max tokens, seed if used, tool-choice settings.
  • The prompt template id and version plus a hash of the fully rendered prompt, so you can tell whether two answers came from the same instructions.
  • Retrieved context: document ids, versions or content hashes, and scores, so you can show which policy document the answer was grounded in on that date.
  • Tool definitions version: the schema the model was offered determines what it could ask for.
  • Guardrail verdicts: which input and output checks ran, their versions and outcomes.

This is enough to explain an outcome and usually to re-run it approximately. Exact replay is not guaranteed even with a seed, because providers change serving infrastructure; how to reason about that during an investigation is covered in AI forensics.

Proving nothing is missing

A log you cannot prove complete is a weak witness. Run three reconciliations daily and alert on any non-zero result. First, sequence gaps: each producer instance numbers its events, and a hole in the sequence means lost events. Second, orphans: an attempt-start without an attempt-end, or a tool intent without an outcome, flags either a crash or a pipeline loss. Third, external counts: compare the number of model attempts recorded against the request counts on the provider's usage report for the same hour and API key; a persistent shortfall means some traffic is bypassing the gateway, typically a team calling the provider directly with its own key.

-- 1. Gaps: producer sequence numbers must be contiguous per producer instance
SELECT producer, instance, prev_seq + 1 AS missing_from, producer_seq - 1 AS missing_to
FROM (
  SELECT producer, instance, producer_seq,
         LAG(producer_seq) OVER (PARTITION BY producer, instance ORDER BY producer_seq) AS prev_seq
  FROM audit_events WHERE ts_ms >= :day_start AND ts_ms < :day_end
) t
WHERE producer_seq <> prev_seq + 1;

-- 2. Orphans: tool intents with no outcome after 15 minutes
SELECT i.request_id, i.step_id, i.payload->>'tool' AS tool
FROM audit_events i
LEFT JOIN audit_events o ON o.kind = 'tool_outcome' AND o.step_id = i.step_id
WHERE i.kind = 'tool_intent' AND o.event_id IS NULL
  AND i.ts_ms < :now_ms - 15 * 60 * 1000;

Worked example: the refund nobody can find

A customer says the support assistant promised them a full refund, and support staff cannot find the promise in the conversation transcript. Querying by the customer's id and date returns one request_id with these events:

SeqProducerEventDetail
1apprequestuser u-8812, tenant retail-eu, feature support-chat
2gatewaymodel_attempt_startattempt 1, template refund-v14, 3 retrieved docs incl. refund-policy v7
3gatewaymodel_attempt_endattempt 1, finish client_disconnected after 212 delivered chars
4gatewaymodel_attempt_startattempt 2, same prompt hash (client auto-retry on reconnect)
5gatewaymodel_attempt_endattempt 2, finish stop, 640 delivered chars
6appdeliveryattempt 2 stored in transcript

The transcript held only attempt 2, which correctly said refunds over 30 days need approval. The delivered text of attempt 1 in the vault reads: 'You are entitled to a full refund, and I can' before the connection dropped. The customer saw that partial sentence before the client retried. The audit log shows the claim was true, identifies the template version that produced the over-confident opening, and shows the bug is in the client, which replaced the visible partial answer on retry without telling the user. Without per-attempt delivered text the investigation would have concluded the customer was mistaken.

Failure modes

  • Logging after the SDK returns. Retries, fallbacks and partial streams are invisible. Capture in the gateway.
  • Shadow traffic. A team calls the provider with its own key and none of its traffic is audited. Reconcile against provider usage per key.
  • Sampled audit events. Someone applies the tracing sample rate to the audit pipeline. Keep the pipelines separate in code and configuration.
  • Raw content in the event store. Prompts full of personal data become searchable by every analyst. Store references and hashes; see PII leakage for what leaks where.
  • Alias instead of version. The log says the model was the alias, which pointed at three different versions over the quarter.
  • Unbounded spool. The collector is down for a day and local disks fill. Alert on spool age and size and decide in advance whether high-impact actions stop.

Trade-offs

Full-content capture costs storage and creates a sensitive data store; hash-and-reference designs are cheaper and safer but need the vault to be reliable. Gateway capture adds a hop of latency, typically small, and a dependency that must be highly available. Fail-closed policies protect integrity at the price of availability, so apply them only to actions with external effects. Per-attempt streamed content roughly doubles writes for retried requests, which is a small cost against being able to answer what the user saw.

What to do next

  1. Write down the five questions your audit log must answer and test today's log against one real past incident.
  2. Route all model traffic through one gateway and reject requests that carry no request id or actor.
  3. Record every attempt with start and end events, delivered text and finish reason, written in a finally block.
  4. Record tool intent before and outcome after every side-effecting call, and block high-impact tools when intent cannot be recorded.
  5. Capture the returned model id, sampling parameters, template version, retrieved document versions and guardrail verdicts.
  6. Add per-producer sequence numbers and run gap, orphan and provider-usage reconciliations daily.
  7. Move raw content to a vault and keep only hashes and references in the event store.
Key takeaway: An LLM audit log is only as good as its capture points. Route model traffic through one gateway that records every attempt, its delivered text and finish reason, and the returned model version; record identity and decisions in the application and intent and outcome around every tool call; link everything with request and step ids; deliver at least once from durable spools; and prove completeness daily with sequence-gap, orphan and provider-usage reconciliation.