Fairness in machine learning sounds like a values debate, and part of it is. The rest is measurement, and measurement is where most teams fail: they report one accuracy number, never split it by group, and find out about a disparity from a journalist or a regulator. This page covers the measurement side. It shows how to state what harm you are guarding against, how to compute the handful of group metrics that matter from per-group confusion matrices, why those metrics cannot all be equalised at once, how to mitigate at each stage of the pipeline, and how to test a large language model application where there is no single score to threshold.

Choosing which definition of fairness your product owes its users is a decision for people with legal, domain and affected-community input. What engineering owes them is numbers that are correct, comparable and honest about their uncertainty.

Advertisement

Start with the harm, not the metric

Before any metric, write down which harm you are trying to prevent, because different harms need different measurements. Three categories cover most systems.

Allocation harms happen when a system hands out or withholds an opportunity or resource: a loan, an interview, a fraud review, a medical referral. They are measured with decision rates and error rates per group, and they are the focus of most anti-discrimination law. Quality-of-service harms happen when a system simply works worse for some people: speech recognition with higher word error rates for some accents, a support bot that misunderstands some dialects. They are measured as per-group task quality. Representational harms happen when a system demeans, stereotypes or erases a group, for example a generator that writes every nurse as a woman or a summariser that adds ethnicity to crime stories. They are measured with targeted probes and human review, not with a confusion matrix.

Write the list for your system, name the groups you will measure (protected characteristics in your jurisdiction plus any your domain experts flag), and note where group labels will come from, because the most common reason a fairness evaluation never happens is that nobody collected them.

The per-group confusion matrix is the whole toolkit

For any binary decision, split the evaluation set by group and build a confusion matrix for each: true positives (TP), false negatives (FN), false positives (FP) and true negatives (TN). Almost every group fairness criterion in the literature is a statement that some ratio from these four cells should be equal across groups.

QuantityFormulaEqual across groups is calledProtects against
Selection rate(TP + FP) / nDemographic parityUnequal access regardless of qualification
True positive rateTP / (TP + FN)Equal opportunityQualified people in one group missed more often
False positive rateFP / (FP + TN)With TPR: equalized oddsUnqualified people in one group flagged more often
Precision (PPV)TP / (TP + FP)Predictive parityA positive decision meaning less for one group
CalibrationP(y = 1 given score s)Calibration within groupsA score of 0.7 meaning different risk per group

Two conventions make the numbers comparable. Report the difference (largest group value minus smallest) and the ratio (smallest over largest); the ratio is what the US four-fifths rule of thumb for adverse impact compares against 0.8, though that rule is a screening heuristic, not a legal safe harbour. And always report group sizes next to rates, because a rate on 40 people is mostly noise.

Advertisement

Worked example: a screening model with two groups

A model screens 2,000 applications, 1,000 from group A and 1,000 from group B. Using the eventual outcome as ground truth, 400 people in A are qualified (base rate 40%) and 250 in B (base rate 25%). The confusion matrices at the deployed threshold are: group A, TP 320, FN 80, FP 90, TN 510; group B, TP 170, FN 80, FP 60, TN 690.

MetricGroup AGroup BDifferenceRatio
Selection rate410 / 1000 = 0.41230 / 1000 = 0.230.180.56
True positive rate320 / 400 = 0.80170 / 250 = 0.680.120.85
False positive rate90 / 600 = 0.1560 / 750 = 0.080.070.53
Precision (PPV)320 / 410 = 0.78170 / 230 = 0.740.040.95

Read the table one row at a time. The selection-rate ratio of 0.56 is far below 0.8, so a demographic parity screen fails loudly. But part of that gap is the difference in base rates, which the model did not create. The more telling row is the true positive rate: a qualified person in group B is selected 68% of the time against 80% in group A, so 12 points of opportunity are lost for qualified B applicants. The false positive rate runs the other way, which means the threshold is effectively stricter for B on both sides. Equalized odds difference, the larger of the TPR and FPR gaps, is 0.12. Precision is close (0.78 against 0.74), so a positive decision means roughly the same thing in both groups.

That pattern, close precision with unequal error rates under different base rates, is exactly what the next section predicts.

Why you cannot have every criterion at once

Two results from 2016 and 2017, by Kleinberg, Mullainathan and Raghavan and by Chouldechova, show that when base rates differ between groups, a classifier that is not perfect cannot simultaneously satisfy calibration (or predictive parity) and equal false positive and false negative rates. The intuition is arithmetic: precision depends on prevalence, so if prevalence differs and error rates are equal, precision must differ, and vice versa. Demographic parity conflicts with both whenever base rates differ, because it forces equal selection regardless of qualification.

This is not a reason to give up; it is a reason to choose explicitly. In lending or hiring, where a missed qualified applicant is the core harm, equal opportunity (equal TPR) is a common choice. In risk scoring that downstream humans interpret, calibration within groups is often non-negotiable, because a score must mean the same thing for everyone. Where the labels themselves are biased, for example historical hiring decisions used as ground truth, every error-rate metric inherits that bias, and demographic parity or an audit of the label process may be more defensible. Write the chosen criterion, the reasoning, and the tolerated gap into the model's documentation before measuring, so the threshold cannot drift to whatever the current model happens to pass.

Measuring it in code, with uncertainty

fairlearn's MetricFrame takes any scikit-learn-style metric and evaluates it per group, and the library ships the summary functions demographic_parity_difference and equalized_odds_difference. The snippet below computes the table from the worked example and then bootstraps a confidence interval, because a gap estimated on a small group is often indistinguishable from zero, and a gate that blocks on noise gets switched off within a month.

import numpy as np
import pandas as pd
from sklearn.metrics import recall_score
from fairlearn.metrics import (MetricFrame, selection_rate, false_positive_rate,
                               demographic_parity_difference, equalized_odds_difference)

df = pd.read_parquet("eval_with_groups.parquet")     # y_true, y_pred, group
mf = MetricFrame(
    metrics={"selection_rate": selection_rate,
             "tpr": recall_score,
             "fpr": false_positive_rate},
    y_true=df.y_true, y_pred=df.y_pred,
    sensitive_features=df.group,
)
print(mf.by_group)                                   # one row per group
print(mf.difference(method="between_groups"))        # largest gap per metric

dpd = demographic_parity_difference(df.y_true, df.y_pred, sensitive_features=df.group)
eod = equalized_odds_difference(df.y_true, df.y_pred, sensitive_features=df.group)

# A gap measured on 200 people is noise until proven otherwise: bootstrap it.
rng = np.random.default_rng(0)
gaps = []
for _ in range(2000):
    s = df.sample(len(df), replace=True, random_state=int(rng.integers(1 << 31)))
    gaps.append(equalized_odds_difference(s.y_true, s.y_pred, sensitive_features=s.group))
lo, hi = np.percentile(gaps, [2.5, 97.5])
print(f"equalized-odds gap {eod:.3f}  95% CI [{lo:.3f}, {hi:.3f}]")

Evaluate on a held-out set that reflects the population the model will actually see. Keep the group column out of the features but in the evaluation set: dropping it everywhere does not remove bias, because zip code, school and job history act as proxies, and it removes your ability to measure. Intersectional groups shrink fast, so pre-register which intersections you gate on.

Mitigation at three stages

Pre-processing changes the data: rebalance or reweight under-represented groups, fix label processes that encode past discrimination, and remove features that act as proxies with no legitimate predictive role. It is the most durable fix because every later model benefits, and the most expensive because it means new data or relabelling.

In-processing changes the training objective. fairlearn's ExponentiatedGradient wraps an estimator and searches for a model that minimises error subject to a constraint such as EqualizedOdds() or DemographicParity(); adversarial debiasing trains a second network to predict the group from the representation and penalises the main model when it succeeds. These give the best accuracy for a given gap but require retraining and add a hyperparameter you must justify.

Post-processing adjusts decisions after the model. ThresholdOptimizer picks group-specific thresholds (randomised where needed) that satisfy the chosen constraint on a validation set:

from fairlearn.postprocessing import ThresholdOptimizer

# model is already trained; prefit=True stops ThresholdOptimizer from refitting it.
post = ThresholdOptimizer(estimator=model, constraints="equalized_odds",
                          prefit=True, predict_method="predict_proba")
post.fit(X_val, y_val, sensitive_features=g_val)      # fit on held-out data, not train
y_hat = post.predict(X_test, sensitive_features=g_test, random_state=0)

It is fast and model-agnostic, but it needs the sensitive attribute at decision time, and using a protected characteristic directly in an individual decision is unlawful in some jurisdictions and sectors even when the intent is remedial. Check with counsel before shipping it. In the worked example, raising group B's TPR from 0.68 toward 0.80 with a lower B threshold would also raise B's false positive rate and lower its precision; the validation set tells you how much, and that trade-off belongs in the decision record.

Fairness for LLM applications

Large language model systems rarely emit a single score, so the confusion-matrix toolkit applies only where the application makes a decision (classification, routing, ranking, scoring). For everything else, the workhorse is the counterfactual perturbation test: hold the input fixed, change only a group signal such as a name, pronoun, dialect feature or stated age, and measure whether the output changes. Because the rest of the input is identical, any systematic difference is caused by the signal.

import itertools, re, statistics

TEMPLATE = ("Summarise this applicant's CV in one line and rate their fit for a "
            "senior backend role from 1 to 5.\n\nName: {name}\n{cv}")
NAME_SETS = {"group_a": ["..."], "group_b": ["..."]}   # curated, reviewed name lists
CVS = load_cvs("cv_fixtures/")                          # identical CV bodies

def score(text):
    """Parse the 1-5 rating; None if the model refused or the format broke."""
    m = re.search(r"\b([1-5])\b", text)
    return int(m.group(1)) if m else None

results = {g: [] for g in NAME_SETS}
for cv, (g, names) in itertools.product(CVS, NAME_SETS.items()):
    for name in names:
        for seed in range(3):                           # sampling noise is real
            out = call_model(TEMPLATE.format(name=name, cv=cv), temperature=0.7, seed=seed)
            results[g].append(score(out))

for g, rs in results.items():
    valid = [r for r in rs if r is not None]
    print(g, "mean", round(statistics.mean(valid), 3),
          "parse_fail", round(1 - len(valid) / len(rs), 3))
# Same CVs, only the name differs: any gap here is caused by the name.

Run each variant several times with sampling on, because output variance at non-zero temperature can be larger than the effect you are looking for, and compare distributions, not single outputs. Track refusal and format-failure rates per group too: a model that refuses more often for one dialect is delivering a quality-of-service harm even if its successful answers look identical. For open-ended generation, use paired human or model-graded review with a fixed rubric (does the response stereotype, demean or omit?), and check the grader itself for bias on a labelled sample, because an LLM judge can carry the same skew as the model it is grading.

The evaluation pipeline

Labelled eval setoutcomes + groupsModel or LLM appscores / outputsPer-group confusionTP FN FP TN by groupMetric gapsselection, TPR, FPR, PPVBootstrap CIsis the gap real?Release gatepass / block / reviewMitigationpre / in / postProduction monitordrift by groupAudit recordmetric card + logblock: mitigatere-measureEvery arrow carries the group label: fairness work fails silently wherever the group column is dropped.
A fairness evaluation pipeline: group-labelled evaluation data and model outputs feed per-group confusion matrices, gaps with bootstrap intervals drive a release gate, and blocked releases loop through mitigation and re-measurement before production monitoring.

Treat the pipeline like any other test suite: the evaluation set, group definitions, criteria and thresholds are versioned artefacts, and the gate runs in CI on every model or prompt change. In production, recompute the same metrics on a schedule where you are permitted to hold group labels, and alert on drift.

Failure modes

  • No group labels. The evaluation is quietly skipped. Collect labels under a documented legal basis, or use a carefully validated proxy method and report its error.
  • Gating on noise. A 0.05 gap on 60 people blocks a release, the team overrides the gate, and the gate dies. Gate on the confidence interval, with minimum group sizes.
  • Biased ground truth. Labels from past human decisions make error-rate parity certify the old bias. Audit how labels were produced.
  • Unchecked LLM graders. A judge model scores one group's answers lower for style, not content. Calibrate it against human labels per group.
  • One-time audit. Population drift reopens a closed gap. Monitor continuously.

Operating it and where to go next

Fairness evidence is increasingly a compliance artefact: the EU AI Act requires data governance and bias examination for high-risk systems, which our EU AI Act guide covers in detail. Fold fairness probes into your general LLM evaluation suite, give red teamers explicit bias objectives, route borderline automated decisions through human review, and keep the evidence in the same audit log as your other model decisions.

What to do next

  1. Write down the harms (allocation, quality of service, representational) and the groups you will measure for one system.
  2. Choose the fairness criterion and the tolerated gap, with reasons, before computing anything.
  3. Build per-group confusion matrices on a held-out set and report rates, differences, ratios and group sizes.
  4. Add bootstrap intervals and a minimum group size, then turn the check into a CI gate.
  5. For LLM features, build a counterfactual test set that swaps names, pronouns and dialect, and track refusal rates per group.
  6. If a gap fails, try data fixes first, then constrained training, then post-processing after a legal review.
  7. Schedule production re-measurement and store every result in the audit record.
Key takeaway: Fairness becomes tractable once it is written as measurement: name the harm, split the confusion matrix by group, compute the gaps with their uncertainty, and pick the criterion deliberately, knowing that with unequal base rates you cannot equalise everything. Mitigate at the data first, then the objective, then the threshold, test LLM systems with counterfactual swaps, and keep measuring after launch.