Most teams first meet an audit of their LLM system as a spreadsheet of questions from a customer, a certification body or an internal risk function, with a deadline. The instinctive response is to answer each question in prose. That is the wrong unit of work. Auditors do not grade descriptions; they test whether a stated control exists and whether it operated, by asking for evidence and sampling from records. A team that can produce the right record for any sampled change, incident or access grant passes. A team that has to reconstruct it from chat history and memory does not, however good its engineering is.

This article treats audit preparation as an engineering problem. It explains what an auditor tests, how to scope an LLM system so that prompts, models, retrieval corpora and tools are visible as assets, how to build a control matrix, how to collect evidence automatically so it is dated and tamper-evident, and how to produce the populations an auditor will sample from. A worked readiness review shows what usually goes wrong. SOC 2 specifics are covered separately in SOC 2 for LLM applications; this page is framework-agnostic.

Advertisement

What an auditor actually tests

Whatever the framework, an assessment asks two questions about each control. Is it designed to address the risk it claims to, and did it operate as designed throughout the period? Design is checked by reading policies, procedures and configuration and by walking through one example. Operation is checked by sampling: the auditor obtains the full list of relevant events in the period, such as every production prompt change, picks a sample, and asks for the evidence that the control applied to each one.

Three consequences follow. You need complete populations, because an auditor who suspects the list is incomplete will widen the testing. You need evidence that existed at the time, because a document created last week to describe a change made in March is not evidence that the control operated in March. And you need consistency, because one sampled change without approval is an exception that must be explained, and several become a finding.

Know which assessment you are preparing for

Frameworks differ in what they certify and what they expect to see, so identify the target before collecting anything.

AssessmentWhat it isWhat it emphasises for an LLM system
SOC 2An attestation report on controls relevant to the Trust Services Criteria, over a point in time or a periodChange management, access, monitoring and vendor oversight, applied to prompts, models and providers
ISO/IEC 42001:2023A certifiable AI management system standard with reference controls in Annex AA functioning management system: AI policy, risk and impact assessment, lifecycle controls, internal audit and management review
EU AI Act, high-risk systemsLegal obligations, including technical documentation (Annex IV) and automatic record-keeping (Article 12)Documentation of design, data and testing; logs that allow events to be traced; post-market monitoring
NIST AI RMF and AI 600-1Voluntary framework with Govern, Map, Measure and Manage functions, and a Generative AI profileNot certifiable, but often the vocabulary customers use in questionnaires
Customer due diligenceSecurity questionnaires and contract clausesData use, retention, model training on customer data, subprocessors, incident notification

Application dates and scope under the EU AI Act have been subject to amendment proposals, so confirm the current position with counsel rather than relying on any article's dates, including this one. The engineering below serves all of these: they ask for different subsets of the same evidence.

Advertisement

Scope the system as assets, not as a model

Auditors scope systems, and an LLM application is several assets that change independently. Inventory each one with an owner, a version identifier and a change process.

  • Models. Provider and exact model version or snapshot, fine-tuned checkpoints, embedding models. A floating alias is a change process you do not control; record it as such.
  • Prompts and instructions. System prompts, templates, few-shot examples and tool descriptions. They change behaviour as much as code does and need the same change control.
  • Retrieval corpora. Sources, ingestion jobs, access rules and deletion handling. A corpus refresh is a change to the system's knowledge.
  • Tools and permissions. What each tool can read and write, under whose identity, and which actions need human confirmation.
  • Guardrails and filters. Input and output classifiers, their thresholds and their versions.
  • Evaluation assets. Test sets, red-team corpora and the thresholds that gate releases.
  • Third parties. Model providers, vector databases and observability vendors that see prompts or outputs, with their own assurance reports.
Audit readiness as a pipeline: systems emit records, collectors turn them into dated evidence, controls cite evidencePrompt + config repoPRs, approvalsModel registryversions, cardsEval + red teamreports, gatesGateway logsrequests, blocksIdentity + ticketsaccess, incidentsEvidence collectorsscheduled jobsEvidence storewrite-once, hash + manifestControl matrixcontrol -> owner -> evidencePopulationsfor samplingGap reportmissing evidenceThe auditor samples from populations and asks for the linked evidence. Readiness means every sampleresolves to a dated, unaltered artifact without anyone reconstructing it from memory.Illustrative architecture; framework-specific control names are deliberately omitted
Evidence comes from the systems that already record changes. Collectors copy it on a schedule into a write-once store, and the control matrix points at it.

Build the control matrix

A control matrix is the index of the audit. Each row is one control: the risk it addresses, a testable statement of what happens, an owner, a frequency, the population it applies to and where its evidence lives. Write control statements so that a stranger could test them; 'prompts are reviewed' is not testable, 'every merge to the prompts directory on main requires one approving review from a code owner other than the author' is.

ControlPopulationEvidence
Prompt and tool-description changes require peer review and a passing eval run before releaseMerged changes under the prompts directoryPR record with approver, linked eval report with commit hash
Model version changes follow a documented evaluation and approvalModel registry promotionsRegistry entry, eval comparison, approval ticket
Release is blocked if safety and quality evals fall below thresholdsRelease pipeline runsGate logs showing scores and thresholds
Production access to prompts, logs and corpora is limited and reviewed quarterlyAccess grants and review recordsIAM export, signed review
Interactions are logged with model version and guardrail decisions, retained for a defined periodLog retention configurationConfiguration export, sample records, log integrity checks
Red-team testing occurs before major releases and findings are tracked to closureReleases and findingsRed-team reports, ticket history
AI incidents are triaged, investigated and reviewedIncident ticketsTickets, timelines, post-incident reviews

Keep the matrix small enough to operate. Twenty well-evidenced controls survive an audit; eighty aspirational ones produce findings. For each control, decide whether its evidence is automatic, a by-product of the tooling, or manual, such as a quarterly sign-off, and prefer automatic.

Collect evidence as code

Evidence is strongest when a machine collects it on a schedule, stores it unmodified and can show when it was collected. A small collector framework does this: each job produces bytes, the framework stores them under the control id and date, and writes a manifest with a SHA-256 hash, the collection time, the period covered and the commit of the collector itself. Store the output in object storage with retention lock or versioning so it cannot be quietly replaced.

import datetime, hashlib, json, pathlib, subprocess

def collect(control_id, name, produce, period, out_dir="evidence"):
    started = datetime.datetime.now(datetime.timezone.utc)
    payload = produce()                                  # bytes: query result, export, report
    folder = pathlib.Path(out_dir) / control_id / started.strftime("%Y-%m-%d")
    folder.mkdir(parents=True, exist_ok=True)
    (folder / name).write_bytes(payload)
    manifest = {
        "control": control_id,
        "artifact": name,
        "sha256": hashlib.sha256(payload).hexdigest(),
        "collected_at": started.isoformat(),
        "period": period,                                # e.g. {"from": "2026-07-01", "to": "2026-09-30"}
        "collector_commit": subprocess.run(["git", "rev-parse", "HEAD"],
                                           capture_output=True, text=True).stdout.strip(),
    }
    (folder / (name + ".manifest.json")).write_text(json.dumps(manifest, indent=2))
    return manifest

Write one collector per evidence type: the IAM export for access reviews, the guardrail configuration, the model registry state, the retention policy of the log bucket. Run them daily or weekly. A configuration snapshot taken every week for a quarter is far stronger evidence that a setting stayed on than one screenshot taken the day before fieldwork.

Produce populations and check them before the auditor does

For change controls, the population is the list of every change in the period, and you should be able to generate it from source systems in one command. Then join it against the evidence the control requires and report the gaps yourself.

def prompt_changes(since, until):
    out = subprocess.run(
        ["git", "log", "--first-parent", "main", f"--since={since}", f"--until={until}",
         "--format=%H|%an|%cI|%s", "--", "prompts/"],
        capture_output=True, text=True, check=True).stdout
    return [dict(zip(["sha", "author", "date", "subject"], l.split("|", 3)))
            for l in out.splitlines() if l]

def gaps(changes, eval_dir, approvals):
    report = []
    for ch in changes:
        problems = []
        if not (pathlib.Path(eval_dir) / f"{ch['sha']}.json").exists():
            problems.append("no eval report for this commit")
        if approvals.get(ch["sha"], set()) - {ch["author"]} == set():
            problems.append("no approval by someone other than the author")
        if problems:
            report.append({**ch, "problems": problems})
    return report

Run the gap report monthly, not just before the audit. Every gap found early is a process fix; every gap found by the auditor is an exception. Keep the generating query with the population, because auditors often ask how the list was produced and whether it could have missed anything, such as changes pushed directly to main or prompts edited in a vendor console that bypasses the repository.

Evidence that is specific to LLM systems

Several evidence types have no equivalent in a conventional application and are where LLM audits usually stall.

  • Eval reports tied to versions. Each report should record the model version, prompt hash, dataset version, metrics and threshold, so it proves which configuration was tested. How to make those numbers meaningful is covered in LLM safety evals.
  • Red-team records. Scope, method, findings and closure status for each exercise, as described in red teaming LLM systems. Findings left open without a risk acceptance become audit findings.
  • Data-use evidence. The provider contract terms on training with your data, the retention settings in use, and proof that deletion requests reach logs, caches and retrieval indexes.
  • Traceability of outputs. For a sampled production interaction, the ability to show which model version, prompt version, retrieved documents and guardrail decisions produced it. This is the same capability investigators need, covered in AI forensics.
  • Human oversight. Where a control says a person approves a high-impact action, logs showing who approved what and when.

Worked example: readiness review for a claims-summarisation assistant

An insurer runs an assistant that summarises claim files for adjusters and drafts letters. It uses a hosted model, a retrieval index over policy documents, and two tools: one reads the claim file, one creates a draft letter that an adjuster must approve. A certification audit is ten weeks away.

Week one builds the asset inventory and finds that the model is referenced by a floating alias, so the team cannot say which model version handled any given claim in the quarter. They pin a dated version and record the alias period as a known limitation with a compensating control: weekly evals against the alias that were already running. Week two drafts eighteen controls. Week three writes collectors and generates the first populations.

The prompt-change population for the quarter has 41 merges. The gap report finds three with no eval report, all from one week when the eval job was disabled during an outage. A diff of live prompts against the repository finds a fourth: a hotfix made in the provider console, never committed, so git could not list it. The team does not backfill evidence. It documents the four exceptions, root cause and remediation: the eval job now fails closed, console write access is removed, and a nightly job diffs live prompts against the repository and alerts on drift.

A mock audit in week seven, run by someone from internal audit who had not seen the system, samples ten prompt changes, five access grants and three incidents. Every sample resolves to evidence within minutes except one incident whose timeline exists only in a chat thread; the incident procedure is changed to require a ticket. By fieldwork, the team hands over the matrix, populations with their generating queries, and an exceptions register. The auditor still samples, but the conversation is about the four documented exceptions rather than about whether controls exist.

Failure modes

  • Evidence created after the fact. Documents written to describe past events are not evidence the control operated; auditors check timestamps.
  • Shadow change paths. Vendor consoles, feature flags and notebook edits that change prompts or models outside the reviewed process.
  • Incomplete populations. A list generated from one system when changes can happen in two.
  • Untestable controls. Policies that say 'appropriate' or 'regularly' cannot be shown to have operated.
  • Logging versus privacy. Retaining every prompt satisfies traceability and violates minimisation. Decide retention, redaction and access deliberately, and evidence both sides.
  • Scope creep at fieldwork. An undisclosed agent tool or data source discovered by the auditor widens the scope at the worst time.

Trade-offs

DecisionOption AOption B
Control countFew, well-evidenced: easier to operate, may leave risks uncoveredMany: broad coverage, more exceptions
Evidence collectionAutomated collectors: durable, needs engineering timeManual screenshots: quick, weak and hard to repeat
Model pinningPinned versions: traceable, needs a deliberate upgrade processFloating alias: newest model, no traceability
Log retentionLong: strong traceability, larger privacy and breach exposureShort: smaller exposure, weaker investigations

What to do next

  1. Name the assessment you face and list the evidence it will ask for, then map it onto the systems that already hold records.
  2. Inventory models, prompts, corpora, tools, guardrails, eval assets and vendors, each with an owner and version identifier.
  3. Write fifteen to twenty testable controls in a matrix with population and evidence location for each.
  4. Build collectors that store hashed, dated evidence in write-once storage, and run them on a schedule from today.
  5. Generate populations for prompt and model changes, run the gap report monthly, and remove any change path that bypasses review.
  6. Run a mock audit with an outsider at least six weeks before fieldwork, and keep an honest exceptions register.
Key takeaway: Passing an audit of an LLM system is less about writing answers than about producing records on demand. Treat models, prompts, corpora, tools, guardrails and eval sets as versioned assets with owners; write a small set of testable controls; collect evidence automatically into write-once storage with hashes and dates; and generate complete populations so you find your own exceptions before an auditor samples them. Pin model versions, close shadow change paths such as vendor consoles, and rehearse with a mock audit while there is still time to fix the process.