A credit score is a ranking, and a lending decision is a threshold applied to that ranking. Fairness problems in credit rarely come from a model that reads race or sex directly; lenders know not to do that. They come from features that stand in for protected characteristics, from training labels that only exist for people who were approved in the past, and from a cutoff chosen without anyone looking at who falls on each side of it. This article is about building the score itself so that those problems are measured, explained and reduced, and about where large language models now enter the pipeline and add new ways to get it wrong.

Two neighbouring pages cover adjacent ground. AI fairness in depth explains the general metrics and why they conflict, and AI credit decisions covers reason codes, the decision record and attacks on the decision pipeline. Here the focus is the score itself: testing it without the protected attribute, auditing proxies, and finding a less discriminatory model that still predicts default. Nothing here is legal advice; involve counsel.

The legal and practical frame

In the United States the Equal Credit Opportunity Act and its implementing rule, Regulation B, prohibit discrimination on characteristics including race, colour, religion, national origin, sex, marital status and age. Two theories matter for model builders. Disparate treatment is using a protected characteristic, or an obvious stand-in for it, in the decision. Disparate impact is a facially neutral practice that falls more heavily on a protected group; the usual analysis asks whether the practice serves a legitimate business need and whether a less discriminatory alternative would serve it about as well. Enforcement posture on disparate impact has shifted over time and varies by agency and state; the technical work is the same either way.

In the European Union, the AI Act lists AI systems used to evaluate the creditworthiness of natural persons or establish their credit score as high-risk, with fraud detection carved out, which brings obligations on data governance, documentation, human oversight and accuracy. Everywhere, you must be able to show what the model does to each group and that you looked for a better option.

One constraint shapes everything below: outside mortgage lending, Regulation B generally prohibits a creditor from asking an applicant's race, national origin or sex. So the team that must test for racial disparity usually does not have race in its data. That is not an excuse; it is a design input.

Reference architecture

Fair credit scoring: where disparity is measured and where it is fixedApplicationsbureau, cash flow, docsFeature pipelineLLM extraction, joinsScore modelprobability of defaultCutoff + pricingapprove, limit, APRProxy auditcan features predict group?Inferred group (BISG)testing only, never a featureDisparity at the cutoffAIR, SMD, conditional ratesLess discriminatory alternativessearch, compare at equal volumeMonitoring + overridesmonthly AIR with intervalsYellow boxes measure; green boxes act. Protected attributes flow only into the testing side.
The scoring path runs left to right along the top. Testing artefacts hang off it, and only the testing side ever touches protected or inferred attributes.

Keep two strictly separated zones. The production zone holds application data, features, the model and the decision. The testing zone holds inferred or self-reported protected attributes, joined by application ID, and produces disparity reports. Nothing flows from testing back into production except human decisions. Enforce it with access controls and a pipeline check that fails if any testing column appears in a training frame.

Testing without the protected attribute

When the protected attribute is missing, the standard approach in US fair-lending testing is Bayesian Improved Surname Geocoding, or BISG. It combines two public sources: the probability of each race and ethnicity category given a surname, from the Census surname list, and the share of each group's population living in the applicant's census block group or tract. Bayes' rule joins them under the assumption that surname and location are independent once group is known.

def bisg(p_race_given_surname, p_geo_given_race):
    # p_race_given_surname: {race: P(race | surname)} from the Census surname table
    # p_geo_given_race:     {race: P(this tract | race)} = tract pop of race / national pop of race
    joint = {r: p_race_given_surname[r] * p_geo_given_race.get(r, 0.0)
             for r in p_race_given_surname}
    total = sum(joint.values())
    if total == 0:                       # unknown tract: fall back to surname alone
        return dict(p_race_given_surname)
    return {r: v / total for r, v in joint.items()}

def weighted_rate(outcomes, probs, group):
    # Prefer the probability as a weight over thresholding it into a hard label.
    num = sum(o * pr[group] for o, pr in zip(outcomes, probs))
    den = sum(pr[group] for pr in probs)
    return num / den

Two practices matter more than the formula. First, use the probabilities as weights, as weighted_rate does. Assigning each applicant to their most likely group and counting throws away uncertainty. Both approaches can bias disparity estimates, so weighting is standard mainly because it keeps every applicant and carries the uncertainty through. Second, report how good the proxy is. BISG works best where surnames and residential patterns are informative and worst where they are not, so its accuracy varies by group and region. If you have a portfolio with self-reported data, such as mortgage applications reported under HMDA, validate the proxy against it before trusting it elsewhere.

Measuring disparity at the cutoff

Disparity is measured on decisions, not on scores, because the harm happens at the cutoff. Three measures cover most reviews. The adverse impact ratio, or AIR, is the approval rate of the protected group divided by the approval rate of the reference group. Many teams flag values below 0.8, a heuristic borrowed from the four-fifths rule in US employment selection guidance; it is a screening threshold, not a legal safe harbour in lending. The standardised mean difference, or SMD, compares mean scores, or mean offered APR, in units of the pooled standard deviation, catching pricing disparities a yes-or-no rate misses. Conditional rates compare applicants with similar risk.

import numpy as np

def air(approved, w_prot, w_ref):
    # approved: 0/1 array; w_*: BISG weights for the protected and reference groups
    rate_p = (approved * w_prot).sum() / w_prot.sum()
    rate_r = (approved * w_ref).sum() / w_ref.sum()
    return rate_p / rate_r

def smd(values, w_prot, w_ref):
    def wmean(w): return (values * w).sum() / w.sum()
    def wvar(w):  return (w * (values - wmean(w)) ** 2).sum() / w.sum()
    pooled = np.sqrt((wvar(w_prot) + wvar(w_ref)) / 2)
    return (wmean(w_prot) - wmean(w_ref)) / pooled

Always report an interval, by resampling applicants and recomputing the statistic. A small portfolio can show an AIR of 0.78 one month and 0.86 the next with nothing changing but noise, and a weighted estimate built on BISG probabilities has wider intervals than its raw count suggests.

Worked example: from 0.75 to 0.83

Take an illustrative personal-loan portfolio with 14,000 applicants on the holdout set. After BISG weighting, the reference group carries a weight of 10,000 and the protected group 4,000. The production model approves 60 percent of the reference group and 45 percent of the protected group, so AIR is 0.45 / 0.60 = 0.75, with a bootstrap interval of roughly 0.71 to 0.79. That is below the 0.8 screening line, so the review continues.

Step one asks whether the gap reflects risk: suppose approval rates within the same observed-default band are close at low risk but diverge in the middle bands, where the cutoff sits. Step two is the proxy audit below, which finds that a feature counting distinct addresses in five years carries much of the group signal. Step three is the alternative search: a candidate that drops that feature and adds a monotone constraint on utilisation loses 0.005 of AUC, from 0.781 to 0.776, and at the same overall approval volume moves AIR to 0.83. Expected loss on the holdout rises by a small fraction of a percent. Business and compliance now decide on numbers, not rhetoric.

Auditing features for proxies

A proxy is a feature whose predictive power comes partly from its correlation with a protected characteristic. ZIP code is the famous one, but modern scoring pipelines have subtler candidates: residential stability, device type, email domain, shopping categories in cash-flow data, the language of uploaded documents. Three tests make the audit systematic.

  1. Group predictability. Train a small model to predict the inferred group from each feature alone, then from all features together. A feature set that predicts group with high AUC tells you the model could reconstruct group whether or not it means to.
  2. Contribution to disparity. Drop or neutralise one feature at a time, retrain or refit, and record the change in AIR and in AUC at fixed approval volume. Features that buy much disparity for little accuracy are the first to question.
  3. Business justification. For every surviving feature, write down why it should predict repayment. A feature nobody can explain is not defensible, however much lift it gives.

Run the audit on the feature pipeline's output, not on the raw application, because derived features and joins are where proxies sneak in.

Searching for a less discriminatory alternative

Searching for a less discriminatory alternative means generating several candidate models that differ in features, constraints and hyperparameters, and comparing them on accuracy and disparity together. The comparison must be fair to the business: hold approval volume or expected loss constant, because any model can raise AIR by approving everyone.

from sklearn.metrics import roc_auc_score

def compare_at_volume(candidates, X_val, y_default, w_prot, w_ref, approve_share=0.55):
    rows = []
    for name, model in candidates.items():
        pd_hat = model.predict_proba(X_val)[:, 1]          # probability of default
        cutoff = np.quantile(pd_hat, approve_share)          # approve the safest 55 percent
        approved = (pd_hat <= cutoff).astype(float)
        rows.append(dict(
            name=name,
            auc=roc_auc_score(y_default, pd_hat),
            bad_rate=y_default[approved == 1].mean(),        # losses among approved
            air=air(approved, w_prot, w_ref),
        ))
    return sorted(rows, key=lambda r: (-r["air"], r["bad_rate"]))

# Keep candidates that are not dominated: no other model has both higher AIR and a lower bad rate.

Useful levers include removing or coarsening proxy features, monotonic constraints so that a better payment history can never lower a score, regularisation that reduces reliance on brittle interactions, and in-processing penalties on disparity. Avoid per-group cutoffs: they use the protected attribute in the decision, which is disparate treatment. Record every candidate and why the chosen one won.

Reject inference and thin files

Credit models learn from loans that were made. Applicants declined in the past have no repayment outcome, so the training data is filtered by the previous policy. If that policy declined one group more often, the model has fewer and less representative examples of that group, and its errors there are larger and harder to see. This is the reject-inference problem, and it is a fairness problem as much as a statistical one.

Classic remedies reweight approved applicants or assign declined ones inferred labels, and each rests on assumptions you cannot fully test. The most reliable source is a small, budgeted exploration programme that approves a random slice of applicants just below the cutoff, with a loss cap agreed in advance, so that the model sees real outcomes where it is least certain. Thin-file applicants raise the same issue: cash-flow data from bank accounts can score people with no bureau history, which can widen access, but only if the new features pass the same proxy audit as the old ones.

Where language models enter the pipeline

Language models now sit in several places in the scoring pipeline: extracting income and employer from pay slips and bank statements, categorising transactions for cash-flow features, summarising files for underwriters, and drafting adverse action notices. Each adds a fairness surface that tabular testing will not catch on its own.

  • Extraction accuracy by group. Document parsers can be less accurate on names, scripts, languages or employer formats more common in some groups. Report extraction error rates per inferred group.
  • Redaction before the model sees text. Strip names, photos, addresses and free-text fields that reveal protected characteristics before prompting. An extraction prompt does not need the applicant's name to read a salary figure.
  • Counterfactual tests. Swap names, pronouns or language markers in otherwise identical documents and check that extracted features and downstream scores do not move.
  • No free-form judgement. Never ask a language model whether an applicant is creditworthy.

Monitoring after launch

A model that was fair on the holdout drifts. Applicant mix changes with marketing campaigns, a new data vendor shifts a feature's distribution, and underwriters override decisions. Monitor monthly: AIR and SMD on approvals, limits and APR, with bootstrap intervals; the group predictability of the live feature set; and override rates by inferred group, because a human who approves exceptions more often for one group can undo a fair model. Alert on a sustained move outside the interval, not on a single month. Version the testing datasets alongside the model so any past report can be regenerated.

A disparity alert should open a ticket with an owner and reach the model risk committee with numbers attached. See governance councils for how that committee can be run.

Failure modes

FailureWhat it looks likePrevention
Unvalidated proxyDisparity estimate skewed by BISG error or hard labelsWeight by probabilities; validate against self-reported data
Comparing at different volumesAlternative looks fairer only because it approves more peopleFix approval share or expected loss first
Testing data leaks into trainingInferred race column appears in a feature frameSeparate stores and a pipeline check that fails
Pricing ignoredApproval parity but higher APR for one groupRun SMD on APR, limits and fees
Silent LLM extraction biasIncome misread more often for some namesPer-group extraction error reporting

What to do next

  1. Build the testing zone: a separate store for inferred attributes keyed by application ID, with BISG probabilities and a validation report of the proxy's accuracy.
  2. Compute AIR and SMD with bootstrap intervals for approvals, limits and APR on your current model.
  3. Run the three-step proxy audit on the feature pipeline's output and write a justification for every surviving feature.
  4. Generate at least five candidate models and compare them at equal approval volume; record the frontier and the reason for the choice.
  5. Add per-group error reporting for any LLM extraction step, plus a name-swap counterfactual test.
  6. Set up monthly monitoring including override rates, and agree in advance what a sustained alert triggers.
  7. Read audit preparation to package the evidence for an examiner or internal audit.
Key takeaway: Fair credit scoring is a testing discipline wrapped around an ordinary risk model. Infer protected attributes only for testing and use them as weights, measure disparity at the cutoff with intervals, audit features for proxies, and search for less discriminatory alternatives at equal approval volume. Then keep measuring, including overrides and any language-model extraction, because a fair launch does not stay fair on its own.