In October 2022 the White House Office of Science and Technology Policy published the Blueprint for an AI Bill of Rights, a short list of protections people should have when automated systems make decisions about them. It was never a law, it said so on its cover, and after the change of administration in January 2025 it disappeared from the White House website and now lives on the archived site. Engineers could reasonably conclude it no longer matters.
That conclusion is wrong in a practical way. The Blueprint's five principles, safe systems, protection from algorithmic discrimination, data privacy, notice and explanation, and human alternatives, are the same five things that binding rules in the EU, in US states and cities, and in data protection law keep requiring. The laws differ in scope and deadlines, and they keep moving, but the engineering underneath them is stable: record decisions, tell people, explain outcomes, let a human reconsider, and measure disparities. This article explains the principles, summarizes where the rules stood at the end of September 2026, and then designs the systems that satisfy them, with a worked example from hiring. It is an engineering guide, not legal advice; your counsel decides which rules apply to you.
What the Blueprint was, and was not
The Blueprint, subtitled Making Automated Systems Work for the American People, set out five principles and a technical companion called From Principles to Practice describing how organizations could implement them. It was explicitly non-binding and did not constitute US government policy. It sat alongside, and was later overtaken by, Executive Order 14110 on AI of October 2023, which was itself revoked on 20 January 2025. The current federal posture is different in direction: an executive order of 11 December 2025, Ensuring a National Policy Framework for Artificial Intelligence, directs agencies to pursue a minimally burdensome national approach and to challenge state AI laws they consider burdensome through litigation and funding conditions. That order does not by itself invalidate any state law; state statutes remain enforceable unless a court or a lawful federal action displaces them.
So the Blueprint is best read as a design specification that happens to have had a government author. Its value for engineers is that it phrases requirements from the affected person's side, which is exactly the side that complaints, audits and lawsuits come from.
The five principles as requirements
| Principle | What the person should get | Engineering requirement |
|---|---|---|
| Safe and effective systems | Systems tested before and monitored after deployment, not used where unsafe | Pre-deployment evaluation, ongoing monitoring, a kill switch |
| Algorithmic discrimination protections | No unjustified different treatment by protected characteristics | Disparity testing before launch and continuously, with documented remediation |
| Data privacy | Data collected only as needed, with meaningful consent | Data inventory, minimization, retention limits, access controls |
| Notice and explanation | To know an automated system is used and why it produced an outcome | Pre-use notices, decision records, reason-based explanations |
| Human alternatives, consideration and fallback | To opt out where appropriate and reach a person who can fix errors | Appeal channel, human review queue with authority to override |
Where the principles are binding today
The European Union's AI Act is the broadest source. Its transparency obligations in Article 50, including telling people when they are interacting with an AI system, apply from 2 August 2026. Its high-risk regime, which covers AI used in employment, credit, education and access to essential services among other Annex III areas, requires risk management, data governance, logging, human oversight and accuracy controls, and Article 86 gives affected persons a right to an explanation of individual decisions taken on the basis of high-risk systems. The Digital Omnibus on AI, formally adopted in mid-2026, moved the application date for stand-alone Annex III high-risk obligations from 2 August 2026 to 2 December 2027. Deferred is not cancelled: the systems described below take longer than sixteen months to build and bed in.
The GDPR already gives individuals in the EU the right, under Article 22, not to be subject to decisions based solely on automated processing that produce legal or similarly significant effects, subject to exceptions, together with safeguards including human intervention and the ability to contest the decision.
In the United States the action is at state and city level. New York City's Local Law 144, enforced since July 2023, requires employers using automated employment decision tools to commission an annual independent bias audit reporting selection rates and impact ratios, publish a summary, and notify candidates. Colorado's original AI Act, SB 24-205, was postponed, stayed by a federal court in April 2026 and then replaced by SB 26-189, signed on 14 May 2026 and effective 1 January 2027. The replacement narrows the focus to transparency for automated decision-making technology used in consequential decisions: a clear notice before use, a plain-language notice within 30 days after an adverse outcome that the technology materially influenced, rights to access and correct personal data and to request meaningful human review where commercially reasonable, and records kept for at least three years. Other states have their own rules and more are pending. The obligations that recur across all of them are the five principles.
Architecture: a rights layer
Treat these obligations as a layer of services around the decision system rather than as properties of the model. The model will change every few months; the layer should not. The layer has six parts. A notice service shows the right disclosure before the system is used. A decision record captures everything needed to reconstruct and explain a decision. An explanation builder turns the record into a plain-language statement. A human review queue receives appeals and routes them to someone with authority to change the outcome. A disparity monitor computes outcome rates by group. A data inventory says what each field is for and when it is deleted.
The decision record
Everything else depends on this record. If you cannot reconstruct why a specific person got a specific outcome on a specific day, you cannot explain it, review it, audit it or defend it. Write the record at decision time, append-only, keyed by a decision id that also appears in every notice the person receives.
@dataclass(frozen=True)
class DecisionRecord:
decision_id: str # printed on every notice the subject receives
subject_ref: str # pseudonymous id, not raw PII
decided_at: datetime
purpose: str # e.g. "screen application for req 4411"
system_version: str # model id + prompt hash + rules version
inputs_ref: str # pointer to the exact features / documents used
outcome: str # "advance" | "reject" | "refer_to_human"
score: float | None
reason_codes: list[str] # e.g. ["MISSING_REQUIRED_CERT", "EXPERIENCE_BELOW_MIN"]
human_involved: bool
notice_shown_id: str # which pre-use notice version the subject saw
retention_until: date
def record(decision: DecisionRecord, store) -> None:
store.append(asdict(decision)) # append-only; corrections are new recordsTwo fields deserve emphasis. system_version must pin everything that affects the outcome: model identifier, prompt template hash, retrieval index version and rules version. An explanation computed against today's model for last month's decision is fiction. And reason_codes are the bridge to explanations: a small, reviewed vocabulary of reasons, each tied to a checkable fact about the input. The logging controls and tamper evidence for such stores are covered in audit logging for LLM systems.
Explanations that are true
With large language models there is a tempting shortcut: ask the model why it decided. Do not ship that as the explanation. A model's stated rationale is generated text, not a trace of its computation, and it can be fluent, plausible and wrong. An explanation that misstates the reason is worse than none, because the person acts on it and the record contradicts it.
Use the model where it is strong and constrain it where it is not. Have the decision step produce structured reason codes against a fixed rubric, validated by code (a certification is missing or not; experience is above the minimum or not). Then generate the explanation text from those codes, either from reviewed templates or by letting an LLM rephrase only the supplied codes and facts, with a check that the output mentions no reason outside the record. Include what the person can do: which information to correct, how to request human review, and the deadline. Keep a copy of the exact text shown, tied to the decision id.
Measuring discrimination: a worked hiring example
Suppose an LLM-assisted screener decides which applicants advance to interview. The standard first measurement is the selection rate per group and the impact ratio: each group's selection rate divided by the rate of the most selected group. The four-fifths guideline from US employment practice treats a ratio below 0.8 as evidence of adverse impact worth investigating. It is a screening heuristic, not a safe harbor, and a ratio above 0.8 does not prove fairness.
def impact_ratios(outcomes, min_n=30):
"""outcomes: {group: (applicants, selected)} -> {group: (rate, ratio, flag)}"""
rates = {g: s / n for g, (n, s) in outcomes.items() if n >= min_n}
best = max(rates.values())
report = {}
for g, (n, s) in outcomes.items():
if n < min_n:
report[g] = (None, None, "too few to measure; aggregate or wait")
continue
ratio = rates[g] / best
report[g] = (round(rates[g], 3), round(ratio, 3),
"INVESTIGATE" if ratio < 0.8 else "ok")
return report
impact_ratios({"group_a": (200, 60), "group_b": (150, 33), "group_c": (12, 2)})
# group_a: rate 0.300, ratio 1.000, ok
# group_b: rate 0.220, ratio 0.733, INVESTIGATE
# group_c: too few to measureGroup B's ratio of 0.733 triggers an investigation, not an automatic conclusion. The next step is to find which reason codes drive the gap. If most group B rejections carry EXPERIENCE_BELOW_MIN and the minimum was set without a job-related justification, the requirement itself is the problem. If the gap appears only when the LLM summarizes free-text résumés, test for proxies: rerun the screener on the same applications with names, addresses and school names redacted and compare. Small groups such as group C need care: rates from twelve people swing wildly, so aggregate over time rather than drawing conclusions or silently dropping them. Protected attributes are often not collected at decision time; bias audits typically use voluntarily provided demographic data held separately under strict access control, which is itself a data privacy decision to document.
Human alternatives that are real
A human review channel only satisfies the principle if the human can and does change outcomes. Automation bias, the tendency to defer to the machine, turns reviewers into a rubber stamp. Design against it: show reviewers the underlying inputs and reason codes before the system's recommendation, give them authority and time to overturn, and track the overturn rate. An overturn rate near zero over hundreds of appeals is a signal to investigate, not a sign of a perfect model. The broader design of approval steps is in human-in-the-loop controls.
Set service levels for appeals, route them away from the team whose metrics depend on the original decision, and feed overturned cases back into evaluation sets so the same error is caught before the next release. Where an opt-out is offered, make the non-automated path genuinely equivalent in speed and outcome, or the opt-out is a penalty.
Data privacy and safe operation
LLM systems create new copies of personal data in prompts, retrieval indexes, traces and evaluation sets. The data inventory should cover all of them, with purpose and retention for each, and the decision record should hold references rather than raw personal data wherever possible. Redact before logging, and apply the controls in PII handling for LLM applications to prompts and outputs alike.
For safe and effective operation, evaluate before launch on data that resembles the real population, including the groups you will monitor, and keep evaluating after launch because inputs drift. Keep a documented way to stop automated decisions and fall back to manual processing within hours. For how these controls map onto NIST AI RMF, ISO/IEC 42001 and the EU AI Act as management-system requirements, see AI safety frameworks.
Failure modes and trade-offs
| Failure | What it looks like | Prevention |
|---|---|---|
| Unreconstructable decision | Cannot say which model or prompt produced an outcome | Pin system_version in every record |
| Confabulated explanation | Stated reason contradicts the record | Explain from reason codes only |
| Rubber-stamp review | Overturn rate near zero | Show inputs first; track overturns |
| Proxy discrimination | Gap persists without protected fields | Redaction reruns; audit features and text |
| Notice drift | Product copy changed, notice versions unknown | Version notices; record which one was shown |
| Compliance by deadline | Controls built after the date, untested | Build the layer once, map it to each rule |
The trade-offs are real. Reason-code explanations are less nuanced than free-form ones. Meaningful review costs reviewer time. Demographic monitoring requires holding sensitive data you would otherwise avoid. Structured decisions constrain how freely an LLM can be used. Each of these is cheaper than rebuilding a decision system under a regulator's timeline, and each makes the system easier to debug, which is its own return.
What to do next
- Inventory every place your products use automated systems to make or materially influence consequential decisions about people.
- For each, write down which jurisdictions' users it touches and have counsel map the applicable rules and dates.
- Implement an append-only decision record with pinned system versions and reason codes.
- Version your pre-use and adverse-outcome notices and record which version each person saw.
- Generate explanations from reason codes, never from the model's self-description.
- Stand up an appeal queue with authority to override, a service level and an overturn-rate metric.
- Compute selection rates and impact ratios per group on a schedule, with small-sample handling and a documented remediation path.
- Map the data each system copies into prompts, logs and indexes, and set retention for each.