In October 2022 the White House Office of Science and Technology Policy published the Blueprint for an AI Bill of Rights, a short list of protections people should have when automated systems make decisions about them. It was never a law, it said so on its cover, and after the change of administration in January 2025 it disappeared from the White House website and now lives on the archived site. Engineers could reasonably conclude it no longer matters.

That conclusion is wrong in a practical way. The Blueprint's five principles, safe systems, protection from algorithmic discrimination, data privacy, notice and explanation, and human alternatives, are the same five things that binding rules in the EU, in US states and cities, and in data protection law keep requiring. The laws differ in scope and deadlines, and they keep moving, but the engineering underneath them is stable: record decisions, tell people, explain outcomes, let a human reconsider, and measure disparities. This article explains the principles, summarizes where the rules stood at the end of September 2026, and then designs the systems that satisfy them, with a worked example from hiring. It is an engineering guide, not legal advice; your counsel decides which rules apply to you.

Advertisement

What the Blueprint was, and was not

The Blueprint, subtitled Making Automated Systems Work for the American People, set out five principles and a technical companion called From Principles to Practice describing how organizations could implement them. It was explicitly non-binding and did not constitute US government policy. It sat alongside, and was later overtaken by, Executive Order 14110 on AI of October 2023, which was itself revoked on 20 January 2025. The current federal posture is different in direction: an executive order of 11 December 2025, Ensuring a National Policy Framework for Artificial Intelligence, directs agencies to pursue a minimally burdensome national approach and to challenge state AI laws they consider burdensome through litigation and funding conditions. That order does not by itself invalidate any state law; state statutes remain enforceable unless a court or a lawful federal action displaces them.

So the Blueprint is best read as a design specification that happens to have had a government author. Its value for engineers is that it phrases requirements from the affected person's side, which is exactly the side that complaints, audits and lawsuits come from.

The five principles as requirements

PrincipleWhat the person should getEngineering requirement
Safe and effective systemsSystems tested before and monitored after deployment, not used where unsafePre-deployment evaluation, ongoing monitoring, a kill switch
Algorithmic discrimination protectionsNo unjustified different treatment by protected characteristicsDisparity testing before launch and continuously, with documented remediation
Data privacyData collected only as needed, with meaningful consentData inventory, minimization, retention limits, access controls
Notice and explanationTo know an automated system is used and why it produced an outcomePre-use notices, decision records, reason-based explanations
Human alternatives, consideration and fallbackTo opt out where appropriate and reach a person who can fix errorsAppeal channel, human review queue with authority to override
Advertisement

Where the principles are binding today

The European Union's AI Act is the broadest source. Its transparency obligations in Article 50, including telling people when they are interacting with an AI system, apply from 2 August 2026. Its high-risk regime, which covers AI used in employment, credit, education and access to essential services among other Annex III areas, requires risk management, data governance, logging, human oversight and accuracy controls, and Article 86 gives affected persons a right to an explanation of individual decisions taken on the basis of high-risk systems. The Digital Omnibus on AI, formally adopted in mid-2026, moved the application date for stand-alone Annex III high-risk obligations from 2 August 2026 to 2 December 2027. Deferred is not cancelled: the systems described below take longer than sixteen months to build and bed in.

The GDPR already gives individuals in the EU the right, under Article 22, not to be subject to decisions based solely on automated processing that produce legal or similarly significant effects, subject to exceptions, together with safeguards including human intervention and the ability to contest the decision.

In the United States the action is at state and city level. New York City's Local Law 144, enforced since July 2023, requires employers using automated employment decision tools to commission an annual independent bias audit reporting selection rates and impact ratios, publish a summary, and notify candidates. Colorado's original AI Act, SB 24-205, was postponed, stayed by a federal court in April 2026 and then replaced by SB 26-189, signed on 14 May 2026 and effective 1 January 2027. The replacement narrows the focus to transparency for automated decision-making technology used in consequential decisions: a clear notice before use, a plain-language notice within 30 days after an adverse outcome that the technology materially influenced, rights to access and correct personal data and to request meaningful human review where commercially reasonable, and records kept for at least three years. Other states have their own rules and more are pending. The obligations that recur across all of them are the five principles.

Architecture: a rights layer

Treat these obligations as a layer of services around the decision system rather than as properties of the model. The model will change every few months; the layer should not. The layer has six parts. A notice service shows the right disclosure before the system is used. A decision record captures everything needed to reconstruct and explain a decision. An explanation builder turns the record into a plain-language statement. A human review queue receives appeals and routes them to someone with authority to change the outcome. A disparity monitor computes outcome rates by group. A data inventory says what each field is for and when it is deleted.

A rights layer around an automated decisionApplicant / userrequestNotice servicebefore use: AI involvedDecision systemmodel + rules + LLMDecision recordinputs, version, reasonsExplanation builderreason codes to textHuman review queueappeal, overrideDisparity monitorimpact ratios by groupData inventorypurpose, retentionlogadverse noticeoutcomesEvery box is a buildable service; together they make notice, explanation, fallback and non-discrimination testable.
The rights layer. The decision system can be a scoring model, rules, an LLM or a mix; the surrounding services stay the same when it changes.

The decision record

Everything else depends on this record. If you cannot reconstruct why a specific person got a specific outcome on a specific day, you cannot explain it, review it, audit it or defend it. Write the record at decision time, append-only, keyed by a decision id that also appears in every notice the person receives.

@dataclass(frozen=True)
class DecisionRecord:
    decision_id: str                 # printed on every notice the subject receives
    subject_ref: str                 # pseudonymous id, not raw PII
    decided_at: datetime
    purpose: str                     # e.g. "screen application for req 4411"
    system_version: str              # model id + prompt hash + rules version
    inputs_ref: str                  # pointer to the exact features / documents used
    outcome: str                     # "advance" | "reject" | "refer_to_human"
    score: float | None
    reason_codes: list[str]          # e.g. ["MISSING_REQUIRED_CERT", "EXPERIENCE_BELOW_MIN"]
    human_involved: bool
    notice_shown_id: str             # which pre-use notice version the subject saw
    retention_until: date

def record(decision: DecisionRecord, store) -> None:
    store.append(asdict(decision))   # append-only; corrections are new records

Two fields deserve emphasis. system_version must pin everything that affects the outcome: model identifier, prompt template hash, retrieval index version and rules version. An explanation computed against today's model for last month's decision is fiction. And reason_codes are the bridge to explanations: a small, reviewed vocabulary of reasons, each tied to a checkable fact about the input. The logging controls and tamper evidence for such stores are covered in audit logging for LLM systems.

Explanations that are true

With large language models there is a tempting shortcut: ask the model why it decided. Do not ship that as the explanation. A model's stated rationale is generated text, not a trace of its computation, and it can be fluent, plausible and wrong. An explanation that misstates the reason is worse than none, because the person acts on it and the record contradicts it.

Use the model where it is strong and constrain it where it is not. Have the decision step produce structured reason codes against a fixed rubric, validated by code (a certification is missing or not; experience is above the minimum or not). Then generate the explanation text from those codes, either from reviewed templates or by letting an LLM rephrase only the supplied codes and facts, with a check that the output mentions no reason outside the record. Include what the person can do: which information to correct, how to request human review, and the deadline. Keep a copy of the exact text shown, tied to the decision id.

Measuring discrimination: a worked hiring example

Suppose an LLM-assisted screener decides which applicants advance to interview. The standard first measurement is the selection rate per group and the impact ratio: each group's selection rate divided by the rate of the most selected group. The four-fifths guideline from US employment practice treats a ratio below 0.8 as evidence of adverse impact worth investigating. It is a screening heuristic, not a safe harbor, and a ratio above 0.8 does not prove fairness.

def impact_ratios(outcomes, min_n=30):
    """outcomes: {group: (applicants, selected)} -> {group: (rate, ratio, flag)}"""
    rates = {g: s / n for g, (n, s) in outcomes.items() if n >= min_n}
    best = max(rates.values())
    report = {}
    for g, (n, s) in outcomes.items():
        if n < min_n:
            report[g] = (None, None, "too few to measure; aggregate or wait")
            continue
        ratio = rates[g] / best
        report[g] = (round(rates[g], 3), round(ratio, 3),
                     "INVESTIGATE" if ratio < 0.8 else "ok")
    return report

impact_ratios({"group_a": (200, 60), "group_b": (150, 33), "group_c": (12, 2)})
# group_a: rate 0.300, ratio 1.000, ok
# group_b: rate 0.220, ratio 0.733, INVESTIGATE
# group_c: too few to measure

Group B's ratio of 0.733 triggers an investigation, not an automatic conclusion. The next step is to find which reason codes drive the gap. If most group B rejections carry EXPERIENCE_BELOW_MIN and the minimum was set without a job-related justification, the requirement itself is the problem. If the gap appears only when the LLM summarizes free-text résumés, test for proxies: rerun the screener on the same applications with names, addresses and school names redacted and compare. Small groups such as group C need care: rates from twelve people swing wildly, so aggregate over time rather than drawing conclusions or silently dropping them. Protected attributes are often not collected at decision time; bias audits typically use voluntarily provided demographic data held separately under strict access control, which is itself a data privacy decision to document.

Human alternatives that are real

A human review channel only satisfies the principle if the human can and does change outcomes. Automation bias, the tendency to defer to the machine, turns reviewers into a rubber stamp. Design against it: show reviewers the underlying inputs and reason codes before the system's recommendation, give them authority and time to overturn, and track the overturn rate. An overturn rate near zero over hundreds of appeals is a signal to investigate, not a sign of a perfect model. The broader design of approval steps is in human-in-the-loop controls.

Set service levels for appeals, route them away from the team whose metrics depend on the original decision, and feed overturned cases back into evaluation sets so the same error is caught before the next release. Where an opt-out is offered, make the non-automated path genuinely equivalent in speed and outcome, or the opt-out is a penalty.

Data privacy and safe operation

LLM systems create new copies of personal data in prompts, retrieval indexes, traces and evaluation sets. The data inventory should cover all of them, with purpose and retention for each, and the decision record should hold references rather than raw personal data wherever possible. Redact before logging, and apply the controls in PII handling for LLM applications to prompts and outputs alike.

For safe and effective operation, evaluate before launch on data that resembles the real population, including the groups you will monitor, and keep evaluating after launch because inputs drift. Keep a documented way to stop automated decisions and fall back to manual processing within hours. For how these controls map onto NIST AI RMF, ISO/IEC 42001 and the EU AI Act as management-system requirements, see AI safety frameworks.

Failure modes and trade-offs

FailureWhat it looks likePrevention
Unreconstructable decisionCannot say which model or prompt produced an outcomePin system_version in every record
Confabulated explanationStated reason contradicts the recordExplain from reason codes only
Rubber-stamp reviewOverturn rate near zeroShow inputs first; track overturns
Proxy discriminationGap persists without protected fieldsRedaction reruns; audit features and text
Notice driftProduct copy changed, notice versions unknownVersion notices; record which one was shown
Compliance by deadlineControls built after the date, untestedBuild the layer once, map it to each rule

The trade-offs are real. Reason-code explanations are less nuanced than free-form ones. Meaningful review costs reviewer time. Demographic monitoring requires holding sensitive data you would otherwise avoid. Structured decisions constrain how freely an LLM can be used. Each of these is cheaper than rebuilding a decision system under a regulator's timeline, and each makes the system easier to debug, which is its own return.

What to do next

  1. Inventory every place your products use automated systems to make or materially influence consequential decisions about people.
  2. For each, write down which jurisdictions' users it touches and have counsel map the applicable rules and dates.
  3. Implement an append-only decision record with pinned system versions and reason codes.
  4. Version your pre-use and adverse-outcome notices and record which version each person saw.
  5. Generate explanations from reason codes, never from the model's self-description.
  6. Stand up an appeal queue with authority to override, a service level and an overturn-rate metric.
  7. Compute selection rates and impact ratios per group on a schedule, with small-sample handling and a documented remediation path.
  8. Map the data each system copies into prompts, logs and indexes, and set retention for each.
Key takeaway: The Blueprint for an AI Bill of Rights was never law and is no longer federal policy, but its five principles describe what binding rules keep demanding: the EU AI Act's transparency, oversight and explanation duties, GDPR Article 22, New York City's bias audits and Colorado's notice, adverse-outcome and human-review requirements, each on its own and moving timeline. Build them once as a rights layer around the decision system: append-only decision records with pinned versions and reason codes, versioned notices, explanations generated from recorded reasons, a human review path that genuinely overturns, impact-ratio monitoring with care for small groups, and a data inventory with retention. Then map that layer to each rule as the rules change.