When a bank's chatbot gets something wrong, the customer can switch banks. When a benefits agency's model gets something wrong, the resident usually has nowhere else to go. Governments hold coercive power. They run services people cannot opt out of, and they operate under public-records, privacy and administrative-law duties that most companies never face. Those facts change the security model for AI. The asset you protect is not only data and uptime. It is the lawfulness and fairness of decisions made about people, and the agency's ability to explain and defend each of them later.

This article is written for engineers and security leads building or reviewing AI systems inside a public body. It covers how to classify a use case by impact, a reference architecture in which the model supports decisions but never makes adverse ones, the threats specific to public-facing government systems, a worked example from benefits administration, the monitoring that shows whether the system treats people fairly, and the procurement terms that keep the agency in control. Rules in this area change with governments, so check the current text of anything cited.

What makes government AI different

Five properties set government AI apart, and each one turns into an engineering requirement.

PropertyWhat it meansEngineering requirement
No exit optionResidents cannot choose another providerError costs land on people, so tune for harm, not average accuracy
Due processAdverse decisions need reasons, notice and appealEvery output that feeds a decision must be explainable from stored evidence
Public recordsPrompts, outputs and logs may be disclosableTreat logs as records: retention schedules, redaction, search
Adversarial publicSome applicants and outsiders try to game or attack the systemUntrusted input handling on every document a resident uploads
ProcurementModels usually come from vendors under contractContract terms are security controls: change notice, data use, testing access, exit

There is a sixth, quieter property: scale of identical decisions. One flawed prompt template, applied to forty thousand applications a month, is a policy change nobody approved. That is why change control matters as much as model quality here.

Classify the use before the model

Start with classification, not with the model. In the United States, OMB memorandum M-25-21 (April 2025) requires federal agencies to identify high-impact AI: AI whose output serves as a principal basis for decisions or actions with a legal, material, binding or significant effect on rights or safety. High-impact uses must follow minimum practices: pre-deployment testing, an AI impact assessment, ongoing monitoring for performance and adverse impacts, adequate training of the people who use it, human oversight with the ability to intervene, consistent remedies or appeals, and feedback from end users and the public. The EU AI Act takes a similar line and lists AI used to assess eligibility for essential public benefits and services among its high-risk uses. Even if neither applies to you, the structure is sound engineering, so adopt it.

Classify each use case in a registry, with a named owner, before any build starts. A practical tiering:

TierExamplesMinimum controls
LowDrafting internal memos, summarising public meeting transcriptsAcceptable-use policy, approved tools only, no resident data
ModeratePublic information chatbot, routing letters to the right teamGrounded answers with citations, escalation to staff, logging, accuracy testing
HighEligibility triage, fraud flags, inspection targeting, case prioritisationImpact assessment, disparity testing, human decision, notice and appeal, monitoring, sign-off
ProhibitedFully automated adverse decisions without review, emotion inference on residentsNot built

The tier is not set by the technology. A summariser is low impact when it condenses a council meeting. It is high impact when the caseworker reads its summary of a medical file instead of the file. Record the use, and re-classify when the use changes.

Reference architecture

A rights-safe decision pipeline: the model drafts and ranks, a person decides, every step is a recordResident channelportal, mail, phoneIntake gatewayscan, redact, isolateUse-case registryimpact tier, ownertier?Model in boundarypinned versionPolicy corpussigned, versionedcitesDecision supportdraft + citationsCaseworkerdecides, signsNoticereasons + appeal pathAppealhuman reviewDecision record and audit loginputs hash, model version, citations, reviewer, outcome, overturnsMonitoringoverturn rate, group disparity, override rate, drift
Reference architecture for a high-impact government use case. The model drafts and ranks inside an authorised boundary; a named caseworker decides; the notice carries reasons and an appeal path; the decision record ties all of it together.

Each component exists to answer a question an auditor, a court or a resident will eventually ask. The intake gateway treats every uploaded document as hostile: malware scanning, file-type allow lists, text extraction in a sandbox, and marking of resident-supplied text so the model can never confuse it with instructions. The use-case registry is consulted at runtime, so a request for a use that was never approved fails closed. The model runs inside the agency's authorised boundary, which for US federal cloud services usually means a FedRAMP-authorised environment, with the exact model version pinned. The policy corpus is the authoritative rule text, versioned and signed, so every citation points at the rule that was in force on the decision date.

The core design rule is simple: the model may recommend, rank, draft and cite; only a person may deny, reduce or terminate. Approvals can sometimes be automated, because an erroneous approval is recoverable and does not trigger due-process duties in the same way. Denials cannot.

The decision record

The decision record is the single most important artefact. If it is designed well, appeals, records requests, audits and monitoring all become queries. If it is missing, none of them is possible.

from dataclasses import dataclass, field
from datetime import datetime, timezone
import hashlib, json

ADVERSE = {"deny", "reduce", "terminate", "refer_for_investigation"}

@dataclass
class DecisionRecord:
    case_id: str
    use_case_id: str             # registry entry; carries tier and owner
    model_version: str           # pinned identifier, never "latest"
    policy_version: str          # corpus release the citations refer to
    input_sha256: str            # hash of the exact documents the model saw
    recommendation: str
    citations: list              # [{"rule": "4.2(b)", "quote": "..."}]
    reviewer_id: str | None = None
    final_action: str | None = None
    reasons: str | None = None
    created: str = field(default_factory=lambda: datetime.now(timezone.utc).isoformat())

def finalise(rec: DecisionRecord, reviewer_id: str, action: str, reasons: str, corpus) -> DecisionRecord:
    if action in ADVERSE and not reviewer_id:
        raise PermissionError("adverse actions need a named human reviewer")
    for cit in rec.citations:                       # every quote must exist in the cited rule
        if cit["quote"] not in corpus.text(rec.policy_version, cit["rule"]):
            raise ValueError(f"citation {cit['rule']} does not match policy {rec.policy_version}")
    if action in ADVERSE and len(reasons.split()) < 15:
        raise ValueError("notice reasons too thin to support an appeal")
    rec.reviewer_id, rec.final_action, rec.reasons = reviewer_id, action, reasons
    return rec

def seal(rec: DecisionRecord, prev_hash: str) -> str:
    body = json.dumps(rec.__dict__, sort_keys=True)
    return hashlib.sha256((prev_hash + body).encode()).hexdigest()   # hash chain for tamper evidence

Three details matter. Verifying quotes against the versioned corpus catches the most damaging failure in this domain: a fluent reason that cites a rule which does not say that. The minimum-reasons check is crude, but it stops empty notices such as "ineligible per policy". The hash chain makes later tampering detectable, which is what you need when a record is challenged years later. Store records on the agency's retention schedule, not the vendor's. The design of tamper-evident trails is covered in LLM audit logging.

Threats specific to public systems

Public-facing government systems face the usual LLM threats plus some that are specific to them.

  • Injection through applications. A resident uploads a letter containing text addressed to the model, such as an instruction to mark the case as urgent and eligible. Treat all uploaded text as data. Spotlight it, keep it out of the instruction channel, and never let the model take actions from it. The pattern is explained in indirect prompt injection.
  • Gaming the threshold. Once the criteria a triage model rewards become known, applications start to match them. Watch for input shifts after publicity; secret criteria are no defence, because due process needs explainable ones.
  • Cross-case leakage. A model with retrieval over all cases can quote another resident's file. Filter retrieval by case and by the caseworker's permissions before ranking, never after.
  • Corpus tampering. Whoever can edit the policy corpus can change outcomes for everyone. Require two-person review and signing for corpus releases.
  • Records exposure. Prompts and outputs may be released under public-records law. Avoid putting other residents' data or security details into prompts, and design redaction for logs before the first request arrives. Records handling is covered in AI in public records.
  • Silent vendor change. A hosted model updated under you invalidates your testing. Pin versions and require notice by contract.

Worked example: benefit renewals

Consider a state agency that receives about 40,000 benefit renewals a month. Caseworkers are behind, and the proposal is a model that reads each renewal packet, checks it against the rules and sorts it into three queues: complete, likely eligible; needs documents; and possible ineligibility.

Under the tiering above this is high impact, because the queue changes how quickly and how carefully a person's benefits are handled. The design therefore follows these rules. The model never closes a case. Complete, likely eligible cases still get a caseworker's approval, but with a short checklist instead of a full review. Needs documents produces a draft request listing the missing items with rule citations, which a caseworker edits and sends. Possible ineligibility goes to a full manual review, and the model's rationale is shown after the caseworker records an initial view, to reduce automation bias.

Before launch, the team runs the model on 2,000 historical renewals with known outcomes. It reports, per queue, how many eventually-eligible people were put in the possible ineligibility queue, broken down by language of the application, disability status and region. It also compares how long people waited under the old process and the new one. The impact assessment records the expected benefit, the risks and the mitigations, and the accountable official signs it. These figures are an illustration of the method, not results from a real deployment.

Monitoring fairness and accuracy

After launch, three signals tell you whether the system is fair and working. The overturn rate is the share of adverse actions reversed on appeal or reconsideration. The override rate is how often caseworkers disagree with the model. Both very high and near-zero override rates are warnings: the first means the model is wrong, the second may mean people have stopped checking. Group disparity compares outcomes and queue placement across groups.

from collections import defaultdict

def monitor(records, group_key, min_n=200):
    by_group = defaultdict(lambda: {"n": 0, "flagged": 0, "overturned": 0, "adverse": 0, "override": 0})
    for r in records:
        g = by_group[r[group_key]]
        g["n"] += 1
        g["flagged"] += r["queue"] == "possible_ineligibility"
        g["adverse"] += r["final_action"] in ADVERSE
        g["overturned"] += r.get("appeal_outcome") == "reversed"
        g["override"] += r["final_action_differs_from_recommendation"]
    rates = {k: v["flagged"] / v["n"] for k, v in by_group.items() if v["n"] >= min_n}
    base = min(rates.values())
    alerts = [k for k, r in rates.items() if base and r / base > 1.25]   # illustrative threshold, set by policy
    return by_group, alerts

The 1.25 ratio is a placeholder. The acceptable disparity is a policy decision for the accountable official, not an engineering default. Group data is often incomplete in government systems, so record how missing values were handled and treat small groups with care, which is why the function ignores groups below min_n. More on choosing and reading fairness metrics in LLM fairness.

Procurement as a security control

Most agencies buy rather than build, so the contract carries many of the controls. Write these into the statement of work and verify them before signing:

  • Data use. Agency and resident data is not used to train or improve vendor models, and is deleted on a defined schedule, with evidence.
  • Version control. Model versions are pinned, with advance notice of changes and the right to stay on a tested version for a defined period.
  • Testing access. The agency can run its own evaluation and disparity tests, including on representative data, and can publish summary results.
  • Logs and records. The agency receives complete logs in an open format, so it can meet records and appeal duties without the vendor.
  • Incident duties. Notice deadlines for security incidents and for discovered model defects.
  • Exit. Data return, transition support and no lock-in of the decision records.

In the US, OMB issued a companion memorandum on AI acquisition, M-25-22, alongside M-25-21. Read the current versions of both, because federal AI policy has been revised with each administration.

Failure modes

  • Automation bias. Caseworkers accept the recommendation without reading the file. Show rationale after an initial human view and audit a random sample.
  • Hallucinated rules. Notices cite rules that do not exist or do not say that. Verify every quote against the versioned corpus.
  • Language gap. The model performs worse on applications in minority languages, so those residents land in slower queues. Test per language before launch.
  • Unexplainable denials. Records lack the evidence the model used, so appeals cannot be answered. Hash and store the exact inputs.
  • Vendor drift. A silent model update changes outcomes. Pin versions and rerun the evaluation on every change.

Trade-offs

Human review of every adverse action costs staff time, and the savings come from the cases the model helps with, not from removing people. Accepting that is the price of due process. Hiding the model's rationale until after the caseworker's first view reduces anchoring but slows review. Strict citation checking rejects some correct paraphrases, which is the right trade in a legal setting. Hosting inside an authorised boundary narrows the choice of models. Publishing your use-case inventory and impact summaries invites scrutiny, and builds the trust that lets the next project go ahead. Plan model incidents with LLM incident response.

What to do next

  1. Build a use-case registry with tier, owner and approved purpose, and make the runtime check it.
  2. Classify every current AI use against a high-impact definition, and stop any fully automated adverse decisions.
  3. Implement a decision record with pinned model and policy versions, input hashes, verified citations and a named reviewer.
  4. Put resident uploads through an intake gateway that isolates their text from instructions.
  5. Run a backtest on historical cases with disparity reporting by language, disability and region before launch.
  6. Agree monitoring thresholds for overturn, override and disparity rates with the accountable official.
  7. Add data use, version pinning, testing access, logs and exit clauses to every AI contract.
  8. Plan for public-records requests on prompts and logs before the first one arrives.
Key takeaway: Government AI is secured by controlling decisions, not only data. Classify every use by its impact on people, let the model recommend and cite while a named person decides adverse actions, keep a decision record that can answer an appeal years later, treat resident uploads as hostile, watch overturn, override and disparity rates, and make the vendor contract carry version pinning, testing access and exit.