Public bodies buy most of the AI they use, so the purchase contract has become one of the main places AI rules bite. Procurement rules decide what a vendor must disclose, what testing the buyer runs before award, what the vendor may do with the buyer's data, and how the buyer leaves. For engineers on either side, those clauses translate into concrete systems: evaluation harnesses, data lineage, documentation pipelines and exit tooling.
This article maps the main procurement regimes in force in late 2026 (US federal, EU, UK and Canada), then focuses on the three obligations that are hardest to evidence: buyer-side acceptance testing, proof that buyer data was not used for training, and a workable exit. It complements the US executive order guide, which tracks the federal instruments in code, and AI contracts in depth, which covers commercial terms.
Why procurement became an AI control
Three forces made procurement an AI control surface. Governments rarely build foundation models, so the vendor relationship is where risk can be shaped. Public decisions carry rights and appeal duties that the buyer cannot meet without records from the vendor. And AI-specific statutes often place duties on the deployer, the public body, which then pushes them down to suppliers through contract terms.
One correction to a common claim: OMB memorandum M-24-10 was a use and governance memo, not a procurement one; its acquisition counterpart was M-24-18. Both were rescinded and replaced in April 2025 by M-25-21 (use) and M-25-22 (acquisition).
The regimes that apply
| Jurisdiction | Instrument | What it asks of AI purchases | Status |
|---|---|---|---|
| US federal | OMB M-25-22 | Lifecycle requirements from market research to closeout; testing of proposed solutions; lock-in protection; contracts must permanently prohibit using non-public inputted agency data and outputted results to further train publicly or commercially available AI, absent explicit agency consent | Applies to solicitations issued 180 days or more after 3 April 2025, and to options exercised after that point |
| US federal | OMB M-26-04 | LLM contracts must address the two Unbiased AI Principles; a minimum transparency package (acceptable use policy, model or system card, end-user resources, feedback channel) | Agency policies due 11 March 2026; memo expires 11 December 2027 |
| EU | AI Act deployer duties | Public bodies deploying high-risk systems must follow instructions, ensure human oversight, keep logs, perform a fundamental rights impact assessment and register the use | Annex III high-risk obligations now apply from 2 December 2027 |
| EU | MCC-AI model clauses | Non-binding templates, updated March 2025, in a high-risk version and a lighter version for non-high-risk systems | Voluntary, widely reused in tenders |
| UK | Procurement Act 2023 plus guidance | General procurement regime; AI-specific expectations come from government AI guidance and the AI Playbook | Act in force since 24 February 2025 |
| Canada | Directive on Automated Decision-Making | Algorithmic Impact Assessment sets an impact level from I to IV, which scales testing, explanation and human-review duties | In force for federal institutions |
The detail of each instrument changes, so treat the table as a map and re-read the source before relying on a date. The engineering pattern underneath is stable: classify the use, publish evidence, let the buyer test, prove data handling, and support exit. For the EU obligations themselves, see the EU AI Act guide.
Lifecycle and evidence architecture
Read the diagram from the buyer's side. At requirements, the buyer classifies the use: high-impact under M-25-21, high-risk under the AI Act, or an impact level under the Canadian directive. That classification decides which clauses go into the solicitation. At selection, M-25-22 tells agencies to test proposed solutions to the greatest extent practicable, which means the buyer needs a harness and representative data. During administration, the contract governs model changes, monitoring and data use. At closeout, the buyer exercises its rights to data and derived products.
From the vendor's side, every green box is a build artefact. If a model card, test report or data-flow description is written by hand for each bid, it will be out of date by the time the contract is signed.
Buyer-side acceptance testing
Performance-based acquisition means the buyer states outcomes and measures them, rather than prescribing a technology. For an LLM service that requires a harness the buyer controls, run against the buyer's own representative cases, with fixed scoring rules published in the solicitation so every bidder is measured the same way.
import json, statistics, time
from dataclasses import dataclass
@dataclass
class Case:
id: str
prompt: str
must_include: list[str] # facts the answer must state
must_refuse: bool = False # out-of-policy request
def score(case, answer):
if case.must_refuse:
return 1.0 if looks_like_refusal(answer) else 0.0
hits = sum(1 for f in case.must_include if f.lower() in answer.lower())
return hits / len(case.must_include)
def evaluate(bidder, cases, runs=3):
rows = []
for case in cases:
for r in range(runs): # repeat: outputs are stochastic
t0 = time.perf_counter()
ans = bidder.complete(case.prompt)
rows.append({"case": case.id, "run": r, "score": score(case, ans),
"latency_s": time.perf_counter() - t0,
"model_version": bidder.reported_version()})
scores = [x["score"] for x in rows]
return {"mean": statistics.mean(scores),
"p10": sorted(scores)[len(scores) // 10],
"versions": sorted({x["model_version"] for x in rows}),
"rows": rows}
# publish the scoring rule with the solicitation; keep the case set private until testingThree rules keep the comparison fair. Repeat each case, because a single run of a stochastic system is noise. Record the model version the bidder reports on every call, and reject a run where it changes mid-test. And keep the case set confidential until testing, publishing only the categories and the scoring method, so bidders cannot tune to it. Simple string checks are shown for clarity; for open-ended answers, use rubric scoring by trained staff on a sample, with agreement measured between raters.
Proving buyer data was not used for training
The M-25-22 data clause is easy to sign and hard to prove. A vendor's assurance is stronger when backed by three engineering controls. First, a tenant-level training flag enforced at the point where data enters any training or fine-tuning corpus, not only in a settings page. Second, lineage: every training dataset build records which tenants' data it drew from, so the vendor can answer an audit query mechanically. Third, canaries: unique, meaningless strings planted in the buyer's traffic that should never appear in any training corpus or model output.
import secrets
def make_canaries(n=20):
return [f"zq-{secrets.token_hex(8)}" for _ in range(n)] # store these privately
def corpus_contains(corpus_shards, canaries):
hits = set()
for shard in corpus_shards: # run inside every dataset build job
for line in shard:
for c in canaries:
if c in line:
hits.add(c)
return hits
def audit_build(build, tenant_flags, canaries):
banned = {t for t, allowed in tenant_flags.items() if not allowed}
leaked_tenants = banned & set(build.source_tenants) # lineage check
leaked_canaries = corpus_contains(build.shards, canaries) # content check
if leaked_tenants or leaked_canaries:
raise RuntimeError(f"build {build.id} blocked: {leaked_tenants or leaked_canaries}")Canaries are a tripwire, not a proof: their absence shows only that those strings did not leak through the paths you scanned. The lineage check is the primary control, and the canary check catches the cases where lineage is wrong, such as logs copied into a research bucket outside the build system. Remember the clause covers outputs too: generated results for the agency must be excluded from training, not just prompts.
Model changes during the contract
Hosted models change underneath a contract. A provider retires a model version, updates safety filters or swaps the default, and the service the buyer tested is no longer the service it runs. Procurement terms usually address this with advance notice of material changes, a right to stay on a tested version for a period, and a right to rerun acceptance tests before switching. M-25-22 also points agencies towards ongoing performance monitoring and sunset criteria, so a degraded system can be retired rather than tolerated.
Engineer for it on both sides. Vendors should report the exact model, prompt and retrieval versions on every response or in a status endpoint, so the buyer can detect drift without trusting a changelog. Buyers should schedule the acceptance harness to run on every reported version change and on a fixed cadence, and compare against the award baseline with an agreed tolerance. A drop beyond tolerance becomes a contract event with a defined remedy, not an argument.
Exit and portability drills
M-25-22 asks agencies to prevent vendor lock-in through knowledge transfer, data and model portability, clear licensing and pricing transparency, and at closeout to exercise ongoing rights to data and derived products. An exit clause that was never tested is a hope. Run an exit drill before go-live and annually:
- Export everything the contract says the buyer owns: prompts and system prompts, retrieval corpora, fine-tuning datasets, evaluation sets, logs and decision records, in documented open formats.
- Restore the export into a neutral environment and confirm it is complete and readable without vendor tooling.
- Rerun the acceptance harness against an alternative model using the exported prompts and corpora, and record the quality gap.
- Confirm deletion at the vendor after export, with a certificate that names systems and backups.
The quality gap in step 3 is the real measure of lock-in. If prompts were tuned for one model's quirks, the gap will be large, and the buyer should know that before renewal.
Worked example: one product, two public buyers
Worked example: a vendor sells a document summarisation service built on a hosted LLM to a US federal agency and to an EU city council in the same quarter.
- US agency. The solicitation post-dates the M-25-22 applicability date, so it includes the data-use prohibition, portability terms and a pre-award test. Because the product is an LLM, M-26-04 adds the minimum transparency package. The vendor's pipeline emits the acceptable use policy, a system card pinned to the deployed model version, user guides and a feedback address.
- EU council. Summarising case files is not automatically high-risk, so the council uses the lighter MCC-AI template. If the same summaries fed eligibility decisions for essential public services, the use could fall under Annex III and the council would need a fundamental rights impact assessment, with high-risk clauses from December 2027.
- Shared engineering. One evaluation pipeline produces both evidence packs; one lineage system enforces no-training for both tenants; one exit drill covers both export formats.
Failure modes
- Stale evidence: the model card describes last quarter's model. Generate cards from the release pipeline and fail the build when versions differ.
- Settings-page compliance: the no-training toggle exists but a log export job copies prompts into a research dataset. Enforce at dataset build.
- Untestable bids: the buyer has no representative data, so selection falls back to demos. Build the case set during requirements.
- Unexercised exit: export works, but nothing can read it.
- Wrong instrument: teams cite rescinded memos. Re-check sources each bid cycle; see AI in government for the operational controls these clauses support.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| Buyer-run acceptance tests | Evidence instead of claims | Staff time and test data preparation |
| Strict no-training clauses | Protects agency data | Vendor cannot improve models on agency usage |
| Annual exit drills | Real portability | Engineering effort on both sides |
| Model clause templates | Fast, consistent tenders | Can be generic for the specific use |
What to do next
- List every public-sector buyer you serve and the instrument that governs each.
- Classify each use case by impact tier in every relevant regime.
- Generate model or system cards, acceptable use policies and test reports from CI, pinned to versions.
- Enforce tenant no-training flags at dataset build, with lineage and canary checks.
- Buyers: build a confidential acceptance case set and publish the scoring rule with the solicitation.
- Run an exit drill before go-live and record the quality gap on an alternative model.
- Review the instruments' status every bid cycle; dates in this area have moved repeatedly.