A compliance dashboard can be wrong for an afternoon and nobody is harmed: someone fixes the query and the numbers move. A compliance report is different. It is sent to a board, a customer or a regulator, someone relies on it, and it is attributable to the people who signed it. Under the EU AI Act, supplying incorrect, incomplete or misleading information to authorities is itself a fineable breach, separate from whatever the report was about. The engineering problem is therefore not drawing charts. It is producing a document whose every number can be traced to evidence, reproduced later, and corrected in an orderly way when it turns out to be wrong.

This article treats AI compliance reporting as a build pipeline. It assumes you already compute the underlying measures, which AI compliance metrics covers, and that you have a program with an obligations register, which the AI compliance program covers. Here the subject is what happens between those inputs and a signed PDF: report types and audiences, frozen snapshots, a builder you can re-run, incident clocks, sign-off, restatements, and a worked example.

Which reports, for whom

Start by listing every report you produce, because each has a different audience, trigger and tolerance for error. Most organisations shipping AI end up with some version of this set.

ReportAudienceTriggerWhat it must answer
AI risk reportBoard or risk committeeQuarterlyIs AI risk within appetite, what changed, what needs a decision
Management review inputTop management under an ISO/IEC 42001 management systemPlanned intervals (clause 9.3)Performance of the AI management system, audit results, nonconformities, opportunities
Serious incident reportMarket surveillance authorityEvent, with a legal deadlineWhat happened, causal link, risk assessment, corrective action
Post-market monitoring summaryInternal, available to authorities on requestPlan-defined cadenceIs the system still performing as documented in production
Customer assurance packEnterprise customers and their auditorsContract renewal or requestWhich controls exist and the evidence that they operate
Control attestationInternal audit and external auditorsPeriod endControl owners assert design and operation for the period

Two of these are not optional for high-risk systems under the EU AI Act. Providers must run post-market monitoring under Article 72 and report serious incidents under Article 73. Following the 2026 Omnibus amendment, the high-risk obligations apply from 2 December 2027 for Annex III systems and 2 August 2028 for Annex I product components, as set out in the EU AI Act guide. Building the reporting pipeline now, while the reports are voluntary, is the cheapest way to have it working when they are not.

Anatomy of a defensible report

A compliance report is a build artifact: frozen inputs, versioned code, signed outputEvidence storestests, logs, ticketsfreezeAs-of snapshotmanifest + hashesReport builderpinned git commitDraft reportevery figure has an idReview and sign-offpreparer, reviewer, ownersignedReport registerversion, hash, recipientsDeliveredboard, regulator, clienterror foundRestatementnew version, notify recipientsRe-running the builder on the stored snapshotmust reproduce the signed figures exactly.
The reporting pipeline. The snapshot and builder commit together make every signed figure reproducible; corrections flow through restatement, never silent edits.

A report you can defend has five parts, and only one of them is the document people read.

  1. An as-of snapshot. A frozen copy, or immutable query, of every input as it stood at the cut-off: control test results, evaluation runs, incident tickets, model inventory. Live tables keep changing, so a report that queries them at render time cannot be reproduced next month.
  2. A manifest. A list of every input with its source, row count and content hash, plus the hash of the manifest itself. This is what lets you prove later that a figure came from a specific evidence set.
  3. A pinned builder. The code that turns the snapshot into figures and prose placeholders, identified by a git commit. Spreadsheets with hand-edited cells are the usual thing this replaces.
  4. The rendered report. Every figure carries a stable identifier that maps back to a query and to evidence records, so a reviewer can click from a number to the rows behind it.
  5. A sign-off record. Who prepared, reviewed and approved, when, and against which manifest hash. Signing a hash rather than a file name means a later change to the inputs cannot inherit the approval.

Code: a report builder you can re-run

The builder below is deliberately small. It freezes inputs into a snapshot directory, hashes them, computes figures from the snapshot only, and refuses to record a sign-off unless the approver differs from the preparer and the hash matches what was reviewed.

import hashlib, json, datetime as dt
from pathlib import Path

def sha256(path: Path) -> str:
    return hashlib.sha256(path.read_bytes()).hexdigest()

def freeze(sources: dict, as_of: dt.date, out: Path) -> dict:
    """Export each source as of the cut-off and write a manifest."""
    out.mkdir(parents=True, exist_ok=False)          # never overwrite a snapshot
    entries = {}
    for name, export in sources.items():
        f = out / f"{name}.jsonl"
        rows = export(as_of)                         # query bounded by as_of
        f.write_text("".join(json.dumps(r, sort_keys=True) + "\n" for r in rows))
        entries[name] = {"rows": len(rows), "sha256": sha256(f)}
    manifest = {"as_of": as_of.isoformat(), "inputs": entries,
                "builder_commit": current_git_commit()}
    body = json.dumps(manifest, sort_keys=True).encode()
    manifest["manifest_sha256"] = hashlib.sha256(body).hexdigest()
    (out / "manifest.json").write_text(json.dumps(manifest, indent=2))
    return manifest

def figures(snapshot: Path) -> dict:
    """Every number in the report is computed here, from the snapshot only."""
    tests = [json.loads(l) for l in (snapshot / "control_tests.jsonl").open()]
    key = [t for t in tests if t["key_control"]]
    return {
        "F1_key_controls_tested": len({t["control_id"] for t in key}),
        "F2_key_control_failures": sum(t["result"] == "fail" for t in key),
        "F3_open_incidents": sum(1 for _ in (snapshot / "incidents.jsonl").open()),
    }

def sign(register: Path, manifest: dict, role: str, person: str, preparer: str):
    if role == "approver" and person == preparer:
        raise PermissionError("approver must differ from preparer")
    rec = {"manifest_sha256": manifest["manifest_sha256"], "role": role,
           "person": person, "at": dt.datetime.now(dt.timezone.utc).isoformat()}
    with register.open("a") as fh:                   # append-only register
        fh.write(json.dumps(rec) + "\n")

Reproducibility is then a test you can run: rebuild figures from the stored snapshot with the recorded commit and compare them with the signed values. Run it in CI each time the builder changes, against the last four snapshots. If a refactor changes a historical figure, you have either fixed a bug, which means a restatement, or introduced one. Either way you find out before a regulator does. Store snapshots under the same retention rules as the underlying audit logs; audit logging for LLM systems covers the tamper-evident storage that makes the hashes worth having.

Incident reports and their clocks

Event-driven reports have deadlines measured from awareness, not from the incident. Under Article 73 a provider of a high-risk system must report a serious incident to the market surveillance authority immediately after establishing a causal link, or its reasonable likelihood, and in any case within 15 days of becoming aware. The limit is two days for a widespread infringement or a serious and irreversible disruption of critical infrastructure, and ten days where a person has died. An incomplete initial report is allowed, followed by a complete one. The provider must then investigate, assess the risk, take corrective action, and avoid altering the system in a way that would hamper evaluation of the cause before informing the authorities.

The duty sits mainly with the provider. A deployer that identifies a serious incident must immediately inform the provider first, then the importer or distributor and the relevant authority, under Article 26(5); if it cannot reach the provider, Article 73 applies to it directly. The Commission has consulted on guidance and a reporting template; use the final published version when it exists rather than inventing a format. Providers of GPAI models with systemic risk report serious incidents to the AI Office under Article 55, a separate channel.

Encode the clocks, because nobody computes deadlines reliably at 2 a.m.:

from datetime import datetime, timedelta

LIMITS = {"death": 10, "widespread_or_critical_infra": 2, "other_serious": 15}

def deadline(aware_at: datetime, category: str) -> datetime:
    return aware_at + timedelta(days=LIMITS[category])

def on_classified(incident):
    due = deadline(incident.aware_at, incident.category)
    page(incident.owner, f"Art 73 report due {due:%Y-%m-%d %H:%M} UTC")
    schedule(due - timedelta(days=1), escalate_to="head_of_compliance")
    incident.freeze_evidence()      # snapshot logs, model version, configs now

Record aware_at as a field set by a human at triage, not the ticket creation time, and treat any change to it as an audited event. The rule worth teaching every on-call engineer is that the clock starts when the organisation knows, so a ticket that sits in an unread queue for a week has used half the allowance.

Writing the narrative

Figures do not make a report. Readers need to know what changed, what is outside tolerance and what decision is wanted. Three habits keep the narrative honest.

  • Every claim cites a figure id. A sentence such as "all key controls operated effectively" must point at F1 and F2. Reviewers reject prose with no reference.
  • Exceptions lead. Put failures, overdue remediations and expired exceptions on page one, with owner and date. A report that leads with green tiles trains its readers to skim.
  • State the limits. Say what was not tested this period, which systems are out of scope, and where a figure rests on a sample. Omissions discovered later damage trust more than a candid amber.

Agree materiality thresholds with the audience in advance: for example, any failed key control, any serious incident, or any production model without a current evaluation is always reported, while minor control deviations are aggregated. Written thresholds stop each quarter's report becoming a negotiation about what to leave out.

Sign-off and segregation of duties

Segregation of duties applies to reports as much as to code. The preparer builds the snapshot and draft. A reviewer, normally from the second line (risk or compliance), re-performs a sample of figures from the evidence and challenges the narrative. The accountable owner, such as the head of AI or the business owner of the system, approves. Nobody signs off figures they produced. For regulator submissions, legal review sits between reviewer and approver.

Keep sign-offs in an append-only register keyed by manifest hash, and record recipients and delivery time for every version. When an auditor asks what the board was told about a model in the second quarter, the register answers in one query, which is exactly the evidence an AI compliance audit will request.

Corrections and restatements

Reports will be wrong sometimes: a query double counts, a late ticket reclassifies an incident, a control test is found to have been performed badly. Never edit a delivered report in place. Publish a new version against a new snapshot, mark exactly which figures changed and why, send it to every recipient of the original, and record whether the error would have changed any decision. For regulator submissions, a corrected or follow-up report through the same channel is the route; quietly fixing the next quarter's figures is how an error becomes a misleading-information finding.

Worked example: a lender in the first quarter of 2028

A lender built its own credit-scoring model, so it is both provider and deployer of an Annex III high-risk system. Its first-quarter 2028 risk report is built from a snapshot frozen on 31 March: 42 key controls tested, 2 failures (a fairness re-evaluation missed its date, and a model change shipped without a second approver), and one serious incident.

The incident: on 2 January a feature-pipeline change mis-coded postcodes for one region, and applicants there were declined at three times the normal rate. An analyst saw the anomaly on 8 January but filed a low-priority ticket; triage on 10 January classified it as a potential infringement of fundamental-rights obligations and set aware_at to that date. The builder flagged that the earlier ticket existed, so legal reviewed whether awareness began on 8 January. To be safe, the team worked to the earlier date: a 15-day deadline of 23 January rather than 25 January. An initial report went on 12 January, the complete report with root cause, affected population and corrective action on 20 January.

In the quarterly report the incident leads page one with figure ids for affected applications and remediation status, the two failed controls follow with owners and dates, and the limits section states that the fairness metrics for one product line rest on a sample. In April, a late ticket shows four more affected applications. Version 2 is issued, figure F7 changes from 1,184 to 1,188, every recipient is notified, and the register shows both versions with their hashes.

Failure modes

FailureHow it shows upPrevention
Report queries live tablesFigures cannot be reproduced; auditors get different numbersFreeze an as-of snapshot with a manifest
Hand-edited spreadsheet cellsA figure has no lineage; nobody knows who changed itAll figures computed by the pinned builder
Awareness time taken from ticket creationDeadline computed late or earlyHuman-set aware_at at triage, audited
Approver also prepared the figuresSelf-review; audit findingBuilder enforces role separation
Silent correction next quarterMisleading-information exposureFormal restatement to all recipients
Green-first layoutExceptions missed by readersExceptions lead, with owners and dates

Trade-offs

The pipeline costs engineering time: snapshot storage, a builder to maintain and a register to operate. For a team with one model and an annual customer questionnaire, a disciplined folder of exported evidence with a checksum file may be enough. The full pipeline pays off once you report to more than one audience from the same evidence, face legal deadlines, or expect audits that ask you to reproduce past statements. Automation also has a limit: the builder can guarantee that numbers are traceable, not that the narrative is candid. Keep the human review, and spend the time it saves on judgement rather than arithmetic.

What to do next

  1. List every compliance report you produce with its audience, trigger and owner, and mark which are legally required and from when.
  2. Freeze the next quarterly report from an as-of snapshot with a hashed manifest, and give every figure an id.
  3. Move figure calculations out of spreadsheets into a builder pinned to a git commit, and add a CI test that rebuilds the last signed report.
  4. Add an aware_at field to incident triage and wire the Article 73 clocks to paging, even before the obligation applies.
  5. Create an append-only sign-off register keyed by manifest hash, enforcing preparer and approver separation.
  6. Write a restatement procedure and materiality thresholds, and get the board or customer to agree them before you need them.
Key takeaway: Treat every compliance report as a build artifact. Freeze inputs into an as-of snapshot with a hashed manifest, compute every figure with pinned code, give each figure an id that traces to evidence, sign the manifest hash with separate preparer and approver, and correct errors by formal restatement. Encode incident clocks such as the Article 73 limits of 15, 10 and 2 days from awareness, and lead every report with its exceptions.