Banks, insurers and lenders were regulated as model users long before anyone said the word AI. Credit scorecards, market-risk models and anti-money-laundering engines all sit under supervisory expectations, consumer protection law and operational resilience rules. Recently those rules moved: the US banking agencies replaced their model risk guidance in April 2026, a circular on algorithmic credit denials was withdrawn, Colorado rewrote its AI law before it took effect, and the EU high-risk timetable slipped. Meanwhile firms are putting large language models into underwriting and servicing work the rules were not written for.

This article maps those rules onto engineering controls, works through a lender with three AI systems, and shows code for the two artefacts regulators ask for most: a use-case inventory and an adverse-action reason generator. Regulatory facts were checked on 2026-10-03, against primary text where possible. This is engineering guidance, not legal advice.

The map in one picture

AI use casesCredit scoring modelgradient-boosted treesLLM memo draftingunderwriter assistantLLM service chatcustomer facingFraud detectiontransaction scoringRegimes that attachSR 26-2 / PRA SS1/23model risk managementECOA and Regulation Bspecific adverse reasonsEU AI Act Annex IIIcredit, life and healthDORAICT third-party riskInternal AI policywhere guidance is silentControls and evidenceUse-case inventoryowner, purpose, tierValidation fileeffective challengeDecision recordsinputs, version, reasonsVendor registercontracts, exit plan
Each AI use case attracts different regimes; the regimes converge on four control artefacts: an inventory, a validation file, decision records and a vendor register.

Why finance is different

Three features make finance different from a generic AI governance programme. First, supervisors already examine models. A bank examiner can ask for the inventory, the validation report and the evidence that someone independent challenged the model, and a weak answer becomes a supervisory finding. Second, decisions about people carry explanation duties. When a lender declines credit, United States law requires the specific principal reasons, and EU law treats creditworthiness scoring of individuals as high-risk AI. Third, outsourcing does not move accountability. A model API from a cloud provider is an ICT third-party service, and the firm stays answerable for what it does. So you rarely need a new AI framework; you need the existing model risk, fair lending and third-party machinery to see AI systems, including ones that do not look like models.

Model risk management: SR 26-2 replaces SR 11-7

On 17 April 2026 the Federal Reserve, OCC and FDIC issued revised interagency guidance on model risk management, published by the Fed as SR 26-2. It supersedes SR 11-7, the 2011 guidance that most bank model risk functions were built around, and SR 21-8 on BSA/AML models; the OCC rescinded its companion bulletin at the same time. The guidance says it is generally aimed at banking organisations with more than $30 billion in total assets, while noting it may be relevant to smaller ones with significant model risk.

Four parts matter for AI teams. The definition: a model is a complex quantitative method, system or approach that applies statistical, economic or financial theories to turn input data into quantitative estimates, excluding simple spreadsheet arithmetic and deterministic rules. The materiality test: model purpose together with model exposure determines materiality, and immaterial models may need little more than identification and performance monitoring. Effective challenge still sits at the centre: critical analysis by objective experts with the expertise, independence and standing to force changes. And there is a section on vendor and third-party products, which recognises that proprietary components limit what you can validate.

The sentence most AI teams will miss is footnote 3: Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance. The same footnote says the principles do apply to traditional statistical models and to non-generative, non-agentic AI, and that the firm's own risk management and governance should determine controls for anything the guidance does not cover. So a gradient-boosted credit model is squarely in scope, while an LLM drafting credit memos is outside SR 26-2 but still expected to be governed. Treat that boundary as temporary.

In the UK, the PRA's SS1/23 model risk principles took effect on 17 May 2024 for banks with internal-model approval for regulatory capital, and explicitly cover AI and machine learning. It has no generative AI carve-out, so the US exclusion does not travel.

Credit and insurance decisions about people

The Equal Credit Opportunity Act requires a creditor that takes adverse action to give the applicant a statement of the specific reasons, and Regulation B implements that duty. In 2022 the CFPB issued Circular 2022-03, saying a creditor cannot escape the duty because its algorithm is too complex to explain. The Bureau withdrew that circular in May 2025 (90 Fed. Reg. 20,084), along with dozens of other guidance documents. Withdrawing the circular did not amend the statute or the regulation: the reasons must still be specific and accurate, and a model whose reasons cannot be derived is a model you cannot lawfully use to decline people. Regulation B's commentary also suggests that listing more than four reasons is unlikely to help the applicant.

In the EU, the AI Act lists in Annex III point 5(b) AI systems used to evaluate the creditworthiness of natural persons or establish their credit score, except systems used to detect financial fraud, and in point 5(c) risk assessment and pricing for life and health insurance, which has no fraud exception. These high-risk systems carry risk management, data governance, logging, human oversight and accuracy duties. The Digital Omnibus on AI moves Annex III obligations from 2 August 2026 to 2 December 2027; commentary differs on its adoption status, so confirm the Official Journal text before relying on the later date.

State law adds a third layer. Colorado's SB 24-205, the first broad US state AI law, never took effect: according to published legal commentary it was replaced in May 2026 by SB 26-189, an automated decision-making law with notice, adverse-outcome explanation and human review duties effective 1 January 2027. For engineers, explanation and human-review plumbing is now demanded by several regimes at once, so build it once.

Operational resilience: DORA and AI vendors

The EU's Digital Operational Resilience Act, Regulation (EU) 2022/2554, has applied since 17 January 2025 to banks, insurers, investment firms, payment institutions and others. It is not an AI law, but it governs how you consume AI services. Article 28(3) requires a register of information covering every contractual arrangement with ICT third-party providers, reported to the competent authority. A hosted LLM API, an embedding service, a vector database as a service and the cloud that runs your own model all belong in that register, with the functions they support and whether those functions are critical or important.

DORA also expects exit strategies, incident reporting and resilience testing. For AI that means: can you switch model provider within the exit plan's time limit, and does testing cover the provider failing or degrading? A feature tied to one provider's proprietary tool-calling format fails the first question.

The inventory as code

Every regime starts from the same question: what AI do you use, for what, and who owns it? Build the inventory as code so obligations are derived consistently and change when the rules do.

from dataclasses import dataclass, field

@dataclass
class UseCase:
    name: str
    owner: str
    markets: set            # {"US", "EU", "UK"}
    purpose: str            # credit_underwriting, life_health_pricing, fraud, drafting, chat
    technique: str          # logistic, gbm, ml_other, llm, agent
    decides_about_people: bool
    vendor: str | None = None
    obligations: list = field(default_factory=list)

def classify(u: UseCase, bank_assets_usd_bn: float, pra_in_scope: bool) -> UseCase:
    ob = u.obligations
    generative = u.technique in {"llm", "agent"}
    if "US" in u.markets and bank_assets_usd_bn > 30:
        ob.append("internal AI policy (SR 26-2 fn.3 excludes gen/agentic AI)" if generative
                  else "SR 26-2 model risk management: inventory, validation, challenge")
    if "UK" in u.markets and pra_in_scope:
        ob.append("PRA SS1/23 model risk management (no generative carve-out)")
    if "US" in u.markets and u.purpose == "credit_underwriting":
        ob.append("ECOA/Reg B: specific principal reasons on adverse action")
    if "EU" in u.markets and u.purpose in {"credit_underwriting", "life_health_pricing"}:
        ob.append("EU AI Act Annex III high-risk (5(b) fraud carve-out does not apply)")
    if "EU" in u.markets and u.vendor:
        ob.append(f"DORA Art. 28(3) register entry and exit plan for {u.vendor}")
    if generative and u.decides_about_people:
        ob.append("POLICY BLOCK: generative output may not be the decision of record")
    return u

The function fails toward more obligations, not fewer. The final rule is internal policy, not law: a generative model may assist a human who decides, but may not itself decide about a person, because its reasons cannot yet be shown to the standard explanation duties need.

Worked example: a lender with three AI systems

Take a $45 billion lender selling consumer loans in the US and EU with three AI systems: a gradient-boosted model that approves or declines applications, a hosted LLM drafting credit memos for underwriters, and an LLM servicing chat assistant. The inventory gives this picture.

SystemRegimesKey controls
Credit model (gbm)SR 26-2; ECOA/Reg B; EU AI Act Annex III 5(b)Independent validation, outcome monitoring by segment, reason codes, decision records, human oversight design
Memo drafting (llm, vendor)Internal AI policy; DORA register; Annex III if the memo effectively decidesUnderwriter signs the decision, citation of source documents, prompt and output logging, provider exit plan
Service chat (llm, vendor)Internal AI policy; DORA register; consumer protection and complaints rulesNo credit decisions in chat, hand-off to a human, retrieval restricted to approved content, logging

The credit model needs reasons it can defend. A common method is to compute each feature's contribution to the applicant's score relative to a reference point (for tree models, often a SHAP-style attribution) and map the most negative contributions to approved reason text. The mapping must fail closed: a feature without approved wording, or a decline with no negative contribution, goes to a human rather than producing a vague notice.

REASON_TEXT = {
    "utilisation": "Proportion of balances to credit limits is too high",
    "delinquency_12m": "Delinquency on one or more accounts in the last 12 months",
    "history_months": "Length of credit history is insufficient",
    "inquiries_6m": "Number of recent requests for credit",
    "dti": "Debt obligations are too high relative to income",
}

def principal_reasons(contribs: dict, k: int = 4) -> list:
    """contribs maps feature -> signed contribution versus the reference applicant;
    negative values pushed the score toward decline."""
    adverse = sorted((v, f) for f, v in contribs.items() if v < 0)
    if not adverse:
        raise ValueError("decline without adverse contribution: route to manual review")
    reasons = []
    for _, f in adverse[:k]:
        if f not in REASON_TEXT:
            raise KeyError(f"no approved reason text for {f}: route to manual review")
        reasons.append(REASON_TEXT[f])
    return reasons

Correlated features are the trap: attribution can split one signal between utilisation and total balance and list two half-reasons. Group correlated features into one reason family before ranking, and test that reasons stay stable under small perturbations. For the memo LLM, validate factual accuracy against the loan file, underwriter change rates, and consistency across comparable applicants.

Decision records as evidence

Regulators ask what happened to a specific person on a specific day. Write one immutable decision record per automated or assisted decision.

{
  "decision_id": "app-2026-10-03-000481",
  "system": "consumer-credit-gbm",
  "model_version": "4.2.1",
  "feature_snapshot_uri": "s3://decisions/2026/10/03/000481/features.parquet",
  "score": 0.31,
  "threshold": 0.45,
  "outcome": "decline",
  "reasons": ["Proportion of balances to credit limits is too high",
              "Delinquency on one or more accounts in the last 12 months"],
  "human_reviewer": null,
  "assist": {"memo_model": "vendor-llm@2026-08", "prompt_hash": "sha256:9c1e...",
             "underwriter_changed_conclusion": false},
  "notice_sent_at": "2026-10-03T14:02:11Z"
}

The fields answer an examiner's questions: which version, which inputs, what threshold, which reasons, and whether a person was involved. For LLM assistance, keep the prompt, output and provider model identifier, because hosted models change behind stable endpoint names.

Failure modes

  • Assuming the SR 26-2 exclusion means no rules. The footnote moves generative and agentic AI out of that guidance and into the firm's own governance; an examiner will still ask how you control them, and UK and EU rules have no matching carve-out.
  • Shadow models. A team uses an LLM to rank collections cases and nobody inventories it. Discover AI through procurement records, cloud billing and egress logs, not just surveys.
  • Explanations that do not match the model. Reasons from a separate surrogate model can contradict the real one. Derive reasons from the deployed model.
  • Automation drift. Underwriters stop reading and approve whatever the memo LLM recommends. Measure override rates; near zero is a warning.
  • Silent provider changes. The vendor updates the model behind the same API name. Pin versions where possible, run a fixed evaluation set daily, and treat a provider change as a model change.

Trade-offs

The central trade-off is model power against explainability. A more complex credit model may rank risk better, but each gain must survive reason testing, fair lending analysis and validation. Many lenders keep the decision model explainable and use complex models as challengers or for fraud. A single LLM provider is cheaper to build on, but DORA's exit expectations push toward a provider abstraction and a tested fallback. And because SR 26-2 lets you do less for immaterial models, a well-kept materiality rating saves more effort than any tool.

Related reading on this site: an obligation register across AI regulations, EU AI Act compliance, the NIST AI Risk Management Framework, audit logging for LLM systems, fairness testing and third-party AI risk.

What to do next

  1. Build the AI use-case inventory as code, including internal LLM tools, and tag each obligation with its source and effective date.
  2. For every model that declines or prices for people, generate reasons from the deployed model, fail closed on unmapped features, and test reason stability in validation.
  3. Write an internal policy for generative and agentic AI that covers what SR 26-2 now leaves out: approval, evaluation, logging, human sign-off and change control.
  4. Add every model API and AI platform to the DORA register of information and write an exit plan with a tested fallback.
  5. Store one immutable decision record per decision, including model version and, for LLM assistance, prompt and output.
  6. Track underwriter override rates and a fixed evaluation set for each LLM, and alert on drift.
  7. Recheck the EU Annex III date and Colorado's replacement law against primary text before your next planning cycle.
Key takeaway: AI in finance is governed mostly by rules that already existed for models, credit decisions and outsourcing. SR 26-2 replaced SR 11-7 in April 2026 and leaves generative and agentic AI to the firm's own governance, ECOA still demands specific reasons for declines, the EU AI Act makes credit scoring and life and health insurance pricing high-risk, and DORA puts every model API in the third-party register. Inventory every AI system as code, derive reasons from the deployed model, log each decision, keep a human as the decision maker for generative assistance, and recheck dates that moved in 2026.