Public bodies buy most of the AI they use, so the purchase contract has become one of the main places AI rules bite. Procurement rules decide what a vendor must disclose, what testing the buyer runs before award, what the vendor may do with the buyer's data, and how the buyer leaves. For engineers on either side, those clauses translate into concrete systems: evaluation harnesses, data lineage, documentation pipelines and exit tooling.

This article maps the main procurement regimes in force in late 2026 (US federal, EU, UK and Canada), then focuses on the three obligations that are hardest to evidence: buyer-side acceptance testing, proof that buyer data was not used for training, and a workable exit. It complements the US executive order guide, which tracks the federal instruments in code, and AI contracts in depth, which covers commercial terms.

Why procurement became an AI control

Three forces made procurement an AI control surface. Governments rarely build foundation models, so the vendor relationship is where risk can be shaped. Public decisions carry rights and appeal duties that the buyer cannot meet without records from the vendor. And AI-specific statutes often place duties on the deployer, the public body, which then pushes them down to suppliers through contract terms.

One correction to a common claim: OMB memorandum M-24-10 was a use and governance memo, not a procurement one; its acquisition counterpart was M-24-18. Both were rescinded and replaced in April 2025 by M-25-21 (use) and M-25-22 (acquisition).

The regimes that apply

JurisdictionInstrumentWhat it asks of AI purchasesStatus
US federalOMB M-25-22Lifecycle requirements from market research to closeout; testing of proposed solutions; lock-in protection; contracts must permanently prohibit using non-public inputted agency data and outputted results to further train publicly or commercially available AI, absent explicit agency consentApplies to solicitations issued 180 days or more after 3 April 2025, and to options exercised after that point
US federalOMB M-26-04LLM contracts must address the two Unbiased AI Principles; a minimum transparency package (acceptable use policy, model or system card, end-user resources, feedback channel)Agency policies due 11 March 2026; memo expires 11 December 2027
EUAI Act deployer dutiesPublic bodies deploying high-risk systems must follow instructions, ensure human oversight, keep logs, perform a fundamental rights impact assessment and register the useAnnex III high-risk obligations now apply from 2 December 2027
EUMCC-AI model clausesNon-binding templates, updated March 2025, in a high-risk version and a lighter version for non-high-risk systemsVoluntary, widely reused in tenders
UKProcurement Act 2023 plus guidanceGeneral procurement regime; AI-specific expectations come from government AI guidance and the AI PlaybookAct in force since 24 February 2025
CanadaDirective on Automated Decision-MakingAlgorithmic Impact Assessment sets an impact level from I to IV, which scales testing, explanation and human-review dutiesIn force for federal institutions

The detail of each instrument changes, so treat the table as a map and re-read the source before relying on a date. The engineering pattern underneath is stable: classify the use, publish evidence, let the buyer test, prove data handling, and support exit. For the EU obligations themselves, see the EU AI Act guide.

Lifecycle and evidence architecture

The procurement lifecycle, and the engineering evidence each stage consumesRequirementsuse case, impact tierSolicitationclauses, test planSelectionbuyer runs testsAdministrationmonitoring, changesCloseoutdata return, exitImpact assessmentAIA, FRIA inputsEvidence packcards, AUP, docsAcceptance harnessbuyer data, scoresNo-train proofcanaries, lineageExit drillexport and restoreVendor evaluation and data platformone pipeline produces every artefact, pinned to model and prompt versionsBuyers increasingly test before award and audit after it; vendors who generate evidence from CI answer in days, not months.
Each procurement stage consumes a specific artefact. Vendors who generate them from one versioned pipeline can answer tenders quickly and keep them true after award.

Read the diagram from the buyer's side. At requirements, the buyer classifies the use: high-impact under M-25-21, high-risk under the AI Act, or an impact level under the Canadian directive. That classification decides which clauses go into the solicitation. At selection, M-25-22 tells agencies to test proposed solutions to the greatest extent practicable, which means the buyer needs a harness and representative data. During administration, the contract governs model changes, monitoring and data use. At closeout, the buyer exercises its rights to data and derived products.

From the vendor's side, every green box is a build artefact. If a model card, test report or data-flow description is written by hand for each bid, it will be out of date by the time the contract is signed.

Buyer-side acceptance testing

Performance-based acquisition means the buyer states outcomes and measures them, rather than prescribing a technology. For an LLM service that requires a harness the buyer controls, run against the buyer's own representative cases, with fixed scoring rules published in the solicitation so every bidder is measured the same way.

import json, statistics, time
from dataclasses import dataclass

@dataclass
class Case:
    id: str
    prompt: str
    must_include: list[str]       # facts the answer must state
    must_refuse: bool = False     # out-of-policy request

def score(case, answer):
    if case.must_refuse:
        return 1.0 if looks_like_refusal(answer) else 0.0
    hits = sum(1 for f in case.must_include if f.lower() in answer.lower())
    return hits / len(case.must_include)

def evaluate(bidder, cases, runs=3):
    rows = []
    for case in cases:
        for r in range(runs):                         # repeat: outputs are stochastic
            t0 = time.perf_counter()
            ans = bidder.complete(case.prompt)
            rows.append({"case": case.id, "run": r, "score": score(case, ans),
                         "latency_s": time.perf_counter() - t0,
                         "model_version": bidder.reported_version()})
    scores = [x["score"] for x in rows]
    return {"mean": statistics.mean(scores),
            "p10": sorted(scores)[len(scores) // 10],
            "versions": sorted({x["model_version"] for x in rows}),
            "rows": rows}

# publish the scoring rule with the solicitation; keep the case set private until testing

Three rules keep the comparison fair. Repeat each case, because a single run of a stochastic system is noise. Record the model version the bidder reports on every call, and reject a run where it changes mid-test. And keep the case set confidential until testing, publishing only the categories and the scoring method, so bidders cannot tune to it. Simple string checks are shown for clarity; for open-ended answers, use rubric scoring by trained staff on a sample, with agreement measured between raters.

Proving buyer data was not used for training

The M-25-22 data clause is easy to sign and hard to prove. A vendor's assurance is stronger when backed by three engineering controls. First, a tenant-level training flag enforced at the point where data enters any training or fine-tuning corpus, not only in a settings page. Second, lineage: every training dataset build records which tenants' data it drew from, so the vendor can answer an audit query mechanically. Third, canaries: unique, meaningless strings planted in the buyer's traffic that should never appear in any training corpus or model output.

import secrets

def make_canaries(n=20):
    return [f"zq-{secrets.token_hex(8)}" for _ in range(n)]   # store these privately

def corpus_contains(corpus_shards, canaries):
    hits = set()
    for shard in corpus_shards:                 # run inside every dataset build job
        for line in shard:
            for c in canaries:
                if c in line:
                    hits.add(c)
    return hits

def audit_build(build, tenant_flags, canaries):
    banned = {t for t, allowed in tenant_flags.items() if not allowed}
    leaked_tenants = banned & set(build.source_tenants)       # lineage check
    leaked_canaries = corpus_contains(build.shards, canaries)   # content check
    if leaked_tenants or leaked_canaries:
        raise RuntimeError(f"build {build.id} blocked: {leaked_tenants or leaked_canaries}")

Canaries are a tripwire, not a proof: their absence shows only that those strings did not leak through the paths you scanned. The lineage check is the primary control, and the canary check catches the cases where lineage is wrong, such as logs copied into a research bucket outside the build system. Remember the clause covers outputs too: generated results for the agency must be excluded from training, not just prompts.

Model changes during the contract

Hosted models change underneath a contract. A provider retires a model version, updates safety filters or swaps the default, and the service the buyer tested is no longer the service it runs. Procurement terms usually address this with advance notice of material changes, a right to stay on a tested version for a period, and a right to rerun acceptance tests before switching. M-25-22 also points agencies towards ongoing performance monitoring and sunset criteria, so a degraded system can be retired rather than tolerated.

Engineer for it on both sides. Vendors should report the exact model, prompt and retrieval versions on every response or in a status endpoint, so the buyer can detect drift without trusting a changelog. Buyers should schedule the acceptance harness to run on every reported version change and on a fixed cadence, and compare against the award baseline with an agreed tolerance. A drop beyond tolerance becomes a contract event with a defined remedy, not an argument.

Exit and portability drills

M-25-22 asks agencies to prevent vendor lock-in through knowledge transfer, data and model portability, clear licensing and pricing transparency, and at closeout to exercise ongoing rights to data and derived products. An exit clause that was never tested is a hope. Run an exit drill before go-live and annually:

  1. Export everything the contract says the buyer owns: prompts and system prompts, retrieval corpora, fine-tuning datasets, evaluation sets, logs and decision records, in documented open formats.
  2. Restore the export into a neutral environment and confirm it is complete and readable without vendor tooling.
  3. Rerun the acceptance harness against an alternative model using the exported prompts and corpora, and record the quality gap.
  4. Confirm deletion at the vendor after export, with a certificate that names systems and backups.

The quality gap in step 3 is the real measure of lock-in. If prompts were tuned for one model's quirks, the gap will be large, and the buyer should know that before renewal.

Worked example: one product, two public buyers

Worked example: a vendor sells a document summarisation service built on a hosted LLM to a US federal agency and to an EU city council in the same quarter.

  • US agency. The solicitation post-dates the M-25-22 applicability date, so it includes the data-use prohibition, portability terms and a pre-award test. Because the product is an LLM, M-26-04 adds the minimum transparency package. The vendor's pipeline emits the acceptable use policy, a system card pinned to the deployed model version, user guides and a feedback address.
  • EU council. Summarising case files is not automatically high-risk, so the council uses the lighter MCC-AI template. If the same summaries fed eligibility decisions for essential public services, the use could fall under Annex III and the council would need a fundamental rights impact assessment, with high-risk clauses from December 2027.
  • Shared engineering. One evaluation pipeline produces both evidence packs; one lineage system enforces no-training for both tenants; one exit drill covers both export formats.

Failure modes

  • Stale evidence: the model card describes last quarter's model. Generate cards from the release pipeline and fail the build when versions differ.
  • Settings-page compliance: the no-training toggle exists but a log export job copies prompts into a research dataset. Enforce at dataset build.
  • Untestable bids: the buyer has no representative data, so selection falls back to demos. Build the case set during requirements.
  • Unexercised exit: export works, but nothing can read it.
  • Wrong instrument: teams cite rescinded memos. Re-check sources each bid cycle; see AI in government for the operational controls these clauses support.

Trade-offs

ChoiceBenefitCost
Buyer-run acceptance testsEvidence instead of claimsStaff time and test data preparation
Strict no-training clausesProtects agency dataVendor cannot improve models on agency usage
Annual exit drillsReal portabilityEngineering effort on both sides
Model clause templatesFast, consistent tendersCan be generic for the specific use

What to do next

  1. List every public-sector buyer you serve and the instrument that governs each.
  2. Classify each use case by impact tier in every relevant regime.
  3. Generate model or system cards, acceptable use policies and test reports from CI, pinned to versions.
  4. Enforce tenant no-training flags at dataset build, with lineage and canary checks.
  5. Buyers: build a confidential acceptance case set and publish the scoring rule with the solicitation.
  6. Run an exit drill before go-live and record the quality gap on an alternative model.
  7. Review the instruments' status every bid cycle; dates in this area have moved repeatedly.
Key takeaway: AI procurement rules turn policy into contract clauses, and contract clauses into engineering artefacts. Classify each use under every regime that applies, generate evidence from your release pipeline, let buyers test with their own data, enforce no-training at dataset build with lineage and canaries, and prove the exit works before you need it.