Most teams first meet an audit of their LLM system as a spreadsheet of questions from a customer, a certification body or an internal risk function, with a deadline. The instinctive response is to answer each question in prose. That is the wrong unit of work. Auditors do not grade descriptions; they test whether a stated control exists and whether it operated, by asking for evidence and sampling from records. A team that can produce the right record for any sampled change, incident or access grant passes. A team that has to reconstruct it from chat history and memory does not, however good its engineering is.
This article treats audit preparation as an engineering problem. It explains what an auditor tests, how to scope an LLM system so that prompts, models, retrieval corpora and tools are visible as assets, how to build a control matrix, how to collect evidence automatically so it is dated and tamper-evident, and how to produce the populations an auditor will sample from. A worked readiness review shows what usually goes wrong. SOC 2 specifics are covered separately in SOC 2 for LLM applications; this page is framework-agnostic.
What an auditor actually tests
Whatever the framework, an assessment asks two questions about each control. Is it designed to address the risk it claims to, and did it operate as designed throughout the period? Design is checked by reading policies, procedures and configuration and by walking through one example. Operation is checked by sampling: the auditor obtains the full list of relevant events in the period, such as every production prompt change, picks a sample, and asks for the evidence that the control applied to each one.
Three consequences follow. You need complete populations, because an auditor who suspects the list is incomplete will widen the testing. You need evidence that existed at the time, because a document created last week to describe a change made in March is not evidence that the control operated in March. And you need consistency, because one sampled change without approval is an exception that must be explained, and several become a finding.
Know which assessment you are preparing for
Frameworks differ in what they certify and what they expect to see, so identify the target before collecting anything.
| Assessment | What it is | What it emphasises for an LLM system |
|---|---|---|
| SOC 2 | An attestation report on controls relevant to the Trust Services Criteria, over a point in time or a period | Change management, access, monitoring and vendor oversight, applied to prompts, models and providers |
| ISO/IEC 42001:2023 | A certifiable AI management system standard with reference controls in Annex A | A functioning management system: AI policy, risk and impact assessment, lifecycle controls, internal audit and management review |
| EU AI Act, high-risk systems | Legal obligations, including technical documentation (Annex IV) and automatic record-keeping (Article 12) | Documentation of design, data and testing; logs that allow events to be traced; post-market monitoring |
| NIST AI RMF and AI 600-1 | Voluntary framework with Govern, Map, Measure and Manage functions, and a Generative AI profile | Not certifiable, but often the vocabulary customers use in questionnaires |
| Customer due diligence | Security questionnaires and contract clauses | Data use, retention, model training on customer data, subprocessors, incident notification |
Application dates and scope under the EU AI Act have been subject to amendment proposals, so confirm the current position with counsel rather than relying on any article's dates, including this one. The engineering below serves all of these: they ask for different subsets of the same evidence.
Scope the system as assets, not as a model
Auditors scope systems, and an LLM application is several assets that change independently. Inventory each one with an owner, a version identifier and a change process.
- Models. Provider and exact model version or snapshot, fine-tuned checkpoints, embedding models. A floating alias is a change process you do not control; record it as such.
- Prompts and instructions. System prompts, templates, few-shot examples and tool descriptions. They change behaviour as much as code does and need the same change control.
- Retrieval corpora. Sources, ingestion jobs, access rules and deletion handling. A corpus refresh is a change to the system's knowledge.
- Tools and permissions. What each tool can read and write, under whose identity, and which actions need human confirmation.
- Guardrails and filters. Input and output classifiers, their thresholds and their versions.
- Evaluation assets. Test sets, red-team corpora and the thresholds that gate releases.
- Third parties. Model providers, vector databases and observability vendors that see prompts or outputs, with their own assurance reports.
Build the control matrix
A control matrix is the index of the audit. Each row is one control: the risk it addresses, a testable statement of what happens, an owner, a frequency, the population it applies to and where its evidence lives. Write control statements so that a stranger could test them; 'prompts are reviewed' is not testable, 'every merge to the prompts directory on main requires one approving review from a code owner other than the author' is.
| Control | Population | Evidence |
|---|---|---|
| Prompt and tool-description changes require peer review and a passing eval run before release | Merged changes under the prompts directory | PR record with approver, linked eval report with commit hash |
| Model version changes follow a documented evaluation and approval | Model registry promotions | Registry entry, eval comparison, approval ticket |
| Release is blocked if safety and quality evals fall below thresholds | Release pipeline runs | Gate logs showing scores and thresholds |
| Production access to prompts, logs and corpora is limited and reviewed quarterly | Access grants and review records | IAM export, signed review |
| Interactions are logged with model version and guardrail decisions, retained for a defined period | Log retention configuration | Configuration export, sample records, log integrity checks |
| Red-team testing occurs before major releases and findings are tracked to closure | Releases and findings | Red-team reports, ticket history |
| AI incidents are triaged, investigated and reviewed | Incident tickets | Tickets, timelines, post-incident reviews |
Keep the matrix small enough to operate. Twenty well-evidenced controls survive an audit; eighty aspirational ones produce findings. For each control, decide whether its evidence is automatic, a by-product of the tooling, or manual, such as a quarterly sign-off, and prefer automatic.
Collect evidence as code
Evidence is strongest when a machine collects it on a schedule, stores it unmodified and can show when it was collected. A small collector framework does this: each job produces bytes, the framework stores them under the control id and date, and writes a manifest with a SHA-256 hash, the collection time, the period covered and the commit of the collector itself. Store the output in object storage with retention lock or versioning so it cannot be quietly replaced.
import datetime, hashlib, json, pathlib, subprocess
def collect(control_id, name, produce, period, out_dir="evidence"):
started = datetime.datetime.now(datetime.timezone.utc)
payload = produce() # bytes: query result, export, report
folder = pathlib.Path(out_dir) / control_id / started.strftime("%Y-%m-%d")
folder.mkdir(parents=True, exist_ok=True)
(folder / name).write_bytes(payload)
manifest = {
"control": control_id,
"artifact": name,
"sha256": hashlib.sha256(payload).hexdigest(),
"collected_at": started.isoformat(),
"period": period, # e.g. {"from": "2026-07-01", "to": "2026-09-30"}
"collector_commit": subprocess.run(["git", "rev-parse", "HEAD"],
capture_output=True, text=True).stdout.strip(),
}
(folder / (name + ".manifest.json")).write_text(json.dumps(manifest, indent=2))
return manifestWrite one collector per evidence type: the IAM export for access reviews, the guardrail configuration, the model registry state, the retention policy of the log bucket. Run them daily or weekly. A configuration snapshot taken every week for a quarter is far stronger evidence that a setting stayed on than one screenshot taken the day before fieldwork.
Produce populations and check them before the auditor does
For change controls, the population is the list of every change in the period, and you should be able to generate it from source systems in one command. Then join it against the evidence the control requires and report the gaps yourself.
def prompt_changes(since, until):
out = subprocess.run(
["git", "log", "--first-parent", "main", f"--since={since}", f"--until={until}",
"--format=%H|%an|%cI|%s", "--", "prompts/"],
capture_output=True, text=True, check=True).stdout
return [dict(zip(["sha", "author", "date", "subject"], l.split("|", 3)))
for l in out.splitlines() if l]
def gaps(changes, eval_dir, approvals):
report = []
for ch in changes:
problems = []
if not (pathlib.Path(eval_dir) / f"{ch['sha']}.json").exists():
problems.append("no eval report for this commit")
if approvals.get(ch["sha"], set()) - {ch["author"]} == set():
problems.append("no approval by someone other than the author")
if problems:
report.append({**ch, "problems": problems})
return reportRun the gap report monthly, not just before the audit. Every gap found early is a process fix; every gap found by the auditor is an exception. Keep the generating query with the population, because auditors often ask how the list was produced and whether it could have missed anything, such as changes pushed directly to main or prompts edited in a vendor console that bypasses the repository.
Evidence that is specific to LLM systems
Several evidence types have no equivalent in a conventional application and are where LLM audits usually stall.
- Eval reports tied to versions. Each report should record the model version, prompt hash, dataset version, metrics and threshold, so it proves which configuration was tested. How to make those numbers meaningful is covered in LLM safety evals.
- Red-team records. Scope, method, findings and closure status for each exercise, as described in red teaming LLM systems. Findings left open without a risk acceptance become audit findings.
- Data-use evidence. The provider contract terms on training with your data, the retention settings in use, and proof that deletion requests reach logs, caches and retrieval indexes.
- Traceability of outputs. For a sampled production interaction, the ability to show which model version, prompt version, retrieved documents and guardrail decisions produced it. This is the same capability investigators need, covered in AI forensics.
- Human oversight. Where a control says a person approves a high-impact action, logs showing who approved what and when.
Worked example: readiness review for a claims-summarisation assistant
An insurer runs an assistant that summarises claim files for adjusters and drafts letters. It uses a hosted model, a retrieval index over policy documents, and two tools: one reads the claim file, one creates a draft letter that an adjuster must approve. A certification audit is ten weeks away.
Week one builds the asset inventory and finds that the model is referenced by a floating alias, so the team cannot say which model version handled any given claim in the quarter. They pin a dated version and record the alias period as a known limitation with a compensating control: weekly evals against the alias that were already running. Week two drafts eighteen controls. Week three writes collectors and generates the first populations.
The prompt-change population for the quarter has 41 merges. The gap report finds three with no eval report, all from one week when the eval job was disabled during an outage. A diff of live prompts against the repository finds a fourth: a hotfix made in the provider console, never committed, so git could not list it. The team does not backfill evidence. It documents the four exceptions, root cause and remediation: the eval job now fails closed, console write access is removed, and a nightly job diffs live prompts against the repository and alerts on drift.
A mock audit in week seven, run by someone from internal audit who had not seen the system, samples ten prompt changes, five access grants and three incidents. Every sample resolves to evidence within minutes except one incident whose timeline exists only in a chat thread; the incident procedure is changed to require a ticket. By fieldwork, the team hands over the matrix, populations with their generating queries, and an exceptions register. The auditor still samples, but the conversation is about the four documented exceptions rather than about whether controls exist.
Failure modes
- Evidence created after the fact. Documents written to describe past events are not evidence the control operated; auditors check timestamps.
- Shadow change paths. Vendor consoles, feature flags and notebook edits that change prompts or models outside the reviewed process.
- Incomplete populations. A list generated from one system when changes can happen in two.
- Untestable controls. Policies that say 'appropriate' or 'regularly' cannot be shown to have operated.
- Logging versus privacy. Retaining every prompt satisfies traceability and violates minimisation. Decide retention, redaction and access deliberately, and evidence both sides.
- Scope creep at fieldwork. An undisclosed agent tool or data source discovered by the auditor widens the scope at the worst time.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Control count | Few, well-evidenced: easier to operate, may leave risks uncovered | Many: broad coverage, more exceptions |
| Evidence collection | Automated collectors: durable, needs engineering time | Manual screenshots: quick, weak and hard to repeat |
| Model pinning | Pinned versions: traceable, needs a deliberate upgrade process | Floating alias: newest model, no traceability |
| Log retention | Long: strong traceability, larger privacy and breach exposure | Short: smaller exposure, weaker investigations |
What to do next
- Name the assessment you face and list the evidence it will ask for, then map it onto the systems that already hold records.
- Inventory models, prompts, corpora, tools, guardrails, eval assets and vendors, each with an owner and version identifier.
- Write fifteen to twenty testable controls in a matrix with population and evidence location for each.
- Build collectors that store hashed, dated evidence in write-once storage, and run them on a schedule from today.
- Generate populations for prompt and model changes, run the gap report monthly, and remove any change path that bypasses review.
- Run a mock audit with an outsider at least six weeks before fieldwork, and keep an honest exceptions register.