A credit score is a ranking, and a lending decision is a threshold applied to that ranking. Fairness problems in credit rarely come from a model that reads race or sex directly; lenders know not to do that. They come from features that stand in for protected characteristics, from training labels that only exist for people who were approved in the past, and from a cutoff chosen without anyone looking at who falls on each side of it. This article is about building the score itself so that those problems are measured, explained and reduced, and about where large language models now enter the pipeline and add new ways to get it wrong.
Two neighbouring pages cover adjacent ground. AI fairness in depth explains the general metrics and why they conflict, and AI credit decisions covers reason codes, the decision record and attacks on the decision pipeline. Here the focus is the score itself: testing it without the protected attribute, auditing proxies, and finding a less discriminatory model that still predicts default. Nothing here is legal advice; involve counsel.
The legal and practical frame
In the United States the Equal Credit Opportunity Act and its implementing rule, Regulation B, prohibit discrimination on characteristics including race, colour, religion, national origin, sex, marital status and age. Two theories matter for model builders. Disparate treatment is using a protected characteristic, or an obvious stand-in for it, in the decision. Disparate impact is a facially neutral practice that falls more heavily on a protected group; the usual analysis asks whether the practice serves a legitimate business need and whether a less discriminatory alternative would serve it about as well. Enforcement posture on disparate impact has shifted over time and varies by agency and state; the technical work is the same either way.
In the European Union, the AI Act lists AI systems used to evaluate the creditworthiness of natural persons or establish their credit score as high-risk, with fraud detection carved out, which brings obligations on data governance, documentation, human oversight and accuracy. Everywhere, you must be able to show what the model does to each group and that you looked for a better option.
One constraint shapes everything below: outside mortgage lending, Regulation B generally prohibits a creditor from asking an applicant's race, national origin or sex. So the team that must test for racial disparity usually does not have race in its data. That is not an excuse; it is a design input.
Reference architecture
Keep two strictly separated zones. The production zone holds application data, features, the model and the decision. The testing zone holds inferred or self-reported protected attributes, joined by application ID, and produces disparity reports. Nothing flows from testing back into production except human decisions. Enforce it with access controls and a pipeline check that fails if any testing column appears in a training frame.
Testing without the protected attribute
When the protected attribute is missing, the standard approach in US fair-lending testing is Bayesian Improved Surname Geocoding, or BISG. It combines two public sources: the probability of each race and ethnicity category given a surname, from the Census surname list, and the share of each group's population living in the applicant's census block group or tract. Bayes' rule joins them under the assumption that surname and location are independent once group is known.
def bisg(p_race_given_surname, p_geo_given_race):
# p_race_given_surname: {race: P(race | surname)} from the Census surname table
# p_geo_given_race: {race: P(this tract | race)} = tract pop of race / national pop of race
joint = {r: p_race_given_surname[r] * p_geo_given_race.get(r, 0.0)
for r in p_race_given_surname}
total = sum(joint.values())
if total == 0: # unknown tract: fall back to surname alone
return dict(p_race_given_surname)
return {r: v / total for r, v in joint.items()}
def weighted_rate(outcomes, probs, group):
# Prefer the probability as a weight over thresholding it into a hard label.
num = sum(o * pr[group] for o, pr in zip(outcomes, probs))
den = sum(pr[group] for pr in probs)
return num / denTwo practices matter more than the formula. First, use the probabilities as weights, as weighted_rate does. Assigning each applicant to their most likely group and counting throws away uncertainty. Both approaches can bias disparity estimates, so weighting is standard mainly because it keeps every applicant and carries the uncertainty through. Second, report how good the proxy is. BISG works best where surnames and residential patterns are informative and worst where they are not, so its accuracy varies by group and region. If you have a portfolio with self-reported data, such as mortgage applications reported under HMDA, validate the proxy against it before trusting it elsewhere.
Measuring disparity at the cutoff
Disparity is measured on decisions, not on scores, because the harm happens at the cutoff. Three measures cover most reviews. The adverse impact ratio, or AIR, is the approval rate of the protected group divided by the approval rate of the reference group. Many teams flag values below 0.8, a heuristic borrowed from the four-fifths rule in US employment selection guidance; it is a screening threshold, not a legal safe harbour in lending. The standardised mean difference, or SMD, compares mean scores, or mean offered APR, in units of the pooled standard deviation, catching pricing disparities a yes-or-no rate misses. Conditional rates compare applicants with similar risk.
import numpy as np
def air(approved, w_prot, w_ref):
# approved: 0/1 array; w_*: BISG weights for the protected and reference groups
rate_p = (approved * w_prot).sum() / w_prot.sum()
rate_r = (approved * w_ref).sum() / w_ref.sum()
return rate_p / rate_r
def smd(values, w_prot, w_ref):
def wmean(w): return (values * w).sum() / w.sum()
def wvar(w): return (w * (values - wmean(w)) ** 2).sum() / w.sum()
pooled = np.sqrt((wvar(w_prot) + wvar(w_ref)) / 2)
return (wmean(w_prot) - wmean(w_ref)) / pooledAlways report an interval, by resampling applicants and recomputing the statistic. A small portfolio can show an AIR of 0.78 one month and 0.86 the next with nothing changing but noise, and a weighted estimate built on BISG probabilities has wider intervals than its raw count suggests.
Worked example: from 0.75 to 0.83
Take an illustrative personal-loan portfolio with 14,000 applicants on the holdout set. After BISG weighting, the reference group carries a weight of 10,000 and the protected group 4,000. The production model approves 60 percent of the reference group and 45 percent of the protected group, so AIR is 0.45 / 0.60 = 0.75, with a bootstrap interval of roughly 0.71 to 0.79. That is below the 0.8 screening line, so the review continues.
Step one asks whether the gap reflects risk: suppose approval rates within the same observed-default band are close at low risk but diverge in the middle bands, where the cutoff sits. Step two is the proxy audit below, which finds that a feature counting distinct addresses in five years carries much of the group signal. Step three is the alternative search: a candidate that drops that feature and adds a monotone constraint on utilisation loses 0.005 of AUC, from 0.781 to 0.776, and at the same overall approval volume moves AIR to 0.83. Expected loss on the holdout rises by a small fraction of a percent. Business and compliance now decide on numbers, not rhetoric.
Auditing features for proxies
A proxy is a feature whose predictive power comes partly from its correlation with a protected characteristic. ZIP code is the famous one, but modern scoring pipelines have subtler candidates: residential stability, device type, email domain, shopping categories in cash-flow data, the language of uploaded documents. Three tests make the audit systematic.
- Group predictability. Train a small model to predict the inferred group from each feature alone, then from all features together. A feature set that predicts group with high AUC tells you the model could reconstruct group whether or not it means to.
- Contribution to disparity. Drop or neutralise one feature at a time, retrain or refit, and record the change in AIR and in AUC at fixed approval volume. Features that buy much disparity for little accuracy are the first to question.
- Business justification. For every surviving feature, write down why it should predict repayment. A feature nobody can explain is not defensible, however much lift it gives.
Run the audit on the feature pipeline's output, not on the raw application, because derived features and joins are where proxies sneak in.
Searching for a less discriminatory alternative
Searching for a less discriminatory alternative means generating several candidate models that differ in features, constraints and hyperparameters, and comparing them on accuracy and disparity together. The comparison must be fair to the business: hold approval volume or expected loss constant, because any model can raise AIR by approving everyone.
from sklearn.metrics import roc_auc_score
def compare_at_volume(candidates, X_val, y_default, w_prot, w_ref, approve_share=0.55):
rows = []
for name, model in candidates.items():
pd_hat = model.predict_proba(X_val)[:, 1] # probability of default
cutoff = np.quantile(pd_hat, approve_share) # approve the safest 55 percent
approved = (pd_hat <= cutoff).astype(float)
rows.append(dict(
name=name,
auc=roc_auc_score(y_default, pd_hat),
bad_rate=y_default[approved == 1].mean(), # losses among approved
air=air(approved, w_prot, w_ref),
))
return sorted(rows, key=lambda r: (-r["air"], r["bad_rate"]))
# Keep candidates that are not dominated: no other model has both higher AIR and a lower bad rate.Useful levers include removing or coarsening proxy features, monotonic constraints so that a better payment history can never lower a score, regularisation that reduces reliance on brittle interactions, and in-processing penalties on disparity. Avoid per-group cutoffs: they use the protected attribute in the decision, which is disparate treatment. Record every candidate and why the chosen one won.
Reject inference and thin files
Credit models learn from loans that were made. Applicants declined in the past have no repayment outcome, so the training data is filtered by the previous policy. If that policy declined one group more often, the model has fewer and less representative examples of that group, and its errors there are larger and harder to see. This is the reject-inference problem, and it is a fairness problem as much as a statistical one.
Classic remedies reweight approved applicants or assign declined ones inferred labels, and each rests on assumptions you cannot fully test. The most reliable source is a small, budgeted exploration programme that approves a random slice of applicants just below the cutoff, with a loss cap agreed in advance, so that the model sees real outcomes where it is least certain. Thin-file applicants raise the same issue: cash-flow data from bank accounts can score people with no bureau history, which can widen access, but only if the new features pass the same proxy audit as the old ones.
Where language models enter the pipeline
Language models now sit in several places in the scoring pipeline: extracting income and employer from pay slips and bank statements, categorising transactions for cash-flow features, summarising files for underwriters, and drafting adverse action notices. Each adds a fairness surface that tabular testing will not catch on its own.
- Extraction accuracy by group. Document parsers can be less accurate on names, scripts, languages or employer formats more common in some groups. Report extraction error rates per inferred group.
- Redaction before the model sees text. Strip names, photos, addresses and free-text fields that reveal protected characteristics before prompting. An extraction prompt does not need the applicant's name to read a salary figure.
- Counterfactual tests. Swap names, pronouns or language markers in otherwise identical documents and check that extracted features and downstream scores do not move.
- No free-form judgement. Never ask a language model whether an applicant is creditworthy.
Monitoring after launch
A model that was fair on the holdout drifts. Applicant mix changes with marketing campaigns, a new data vendor shifts a feature's distribution, and underwriters override decisions. Monitor monthly: AIR and SMD on approvals, limits and APR, with bootstrap intervals; the group predictability of the live feature set; and override rates by inferred group, because a human who approves exceptions more often for one group can undo a fair model. Alert on a sustained move outside the interval, not on a single month. Version the testing datasets alongside the model so any past report can be regenerated.
A disparity alert should open a ticket with an owner and reach the model risk committee with numbers attached. See governance councils for how that committee can be run.
Failure modes
| Failure | What it looks like | Prevention |
|---|---|---|
| Unvalidated proxy | Disparity estimate skewed by BISG error or hard labels | Weight by probabilities; validate against self-reported data |
| Comparing at different volumes | Alternative looks fairer only because it approves more people | Fix approval share or expected loss first |
| Testing data leaks into training | Inferred race column appears in a feature frame | Separate stores and a pipeline check that fails |
| Pricing ignored | Approval parity but higher APR for one group | Run SMD on APR, limits and fees |
| Silent LLM extraction bias | Income misread more often for some names | Per-group extraction error reporting |
What to do next
- Build the testing zone: a separate store for inferred attributes keyed by application ID, with BISG probabilities and a validation report of the proxy's accuracy.
- Compute AIR and SMD with bootstrap intervals for approvals, limits and APR on your current model.
- Run the three-step proxy audit on the feature pipeline's output and write a justification for every surviving feature.
- Generate at least five candidate models and compare them at equal approval volume; record the frontier and the reason for the choice.
- Add per-group error reporting for any LLM extraction step, plus a name-swap counterfactual test.
- Set up monthly monitoring including override rates, and agree in advance what a sustained alert triggers.
- Read audit preparation to package the evidence for an examiner or internal audit.