A fairness audit is a structured, evidence-backed judgment about whether an AI system produces disparate outcomes across groups of people, carried out against stated criteria by someone independent enough to be believed. It is not the same as fairness monitoring, a dashboard the product team watches, or fairness testing, the checks engineers run before release. An audit produces a report that others rely on: regulators, candidates, customers and the board.
The best-specified legal example is New York City's Local Law 144. It requires employers and employment agencies using an automated employment decision tool (AEDT) for NYC candidates to have an independent bias audit within one year before use, publish a summary of the results, and notify candidates. This article uses its rules as the worked regime, because they force the hard decisions: which rate, which categories, what to do with small groups and unknowns. It then goes beyond them to the uncertainty analysis, diagnosis and evidence pack that make an audit worth trusting. The metrics themselves are explained in AI Fairness, in depth.
What makes it an audit
Three properties separate an audit from a test. Criteria are fixed before the numbers are seen: which metric, which groups, what counts as a finding. Independence means the auditor is not the builder. Under the LL144 rules an independent auditor cannot have been involved in using, developing or distributing the tool, cannot work for the employer or vendor, and cannot hold a financial interest in either. Evidence means another competent person can reproduce every number from the archived inputs. General audit craft, such as scoping, sampling and writing findings, is covered in AI Compliance Audit, in depth. This article covers what is specific to fairness.
Scoping the engagement
Write a scope memo and get it signed before touching data. It is the document that stops the audit drifting toward whichever cut looks best.
| Field | Example | Why it matters |
|---|---|---|
| Decision point | Advance from application to recruiter screen | Each stage of a funnel has its own disparity |
| Tool output | Binary advance flag (selection) or 0-100 score (scoring) | Determines selection rate or scoring rate |
| Population and window | All NYC applicants, 1 Jan to 30 Jun | Defines the denominator |
| Categories | Sex, race/ethnicity, and their intersections | LL144 requires all three tables |
| Data sources | ATS export plus voluntary self-identification | Demographics usually live elsewhere |
| Flag rule | Impact ratio below 0.8, or a significant gap | Fixed before results |
| Deliverables | Published summary, full report, evidence pack | Different audiences, different detail |
Data: historical, test and unknown
Use historical data from real use where it exists. The LL144 rules allow test data only when historical data is insufficient, and the summary must then explain why and how the test data was built. A vendor can commission one audit across many employers' data, but an employer may rely on it only if it contributed its own historical data or has never used the tool.
Demographics are the hard part. Hiring systems rarely record sex or race at decision time, so the audit joins tool outputs to voluntary self-identification collected separately. Many candidates decline, and those rows form the unknown category. LL144 requires the summary to report how many individuals fall in it. Do not impute categories for the published tables. Imputation methods such as surname and geography inference can support internal diagnosis, with the caveats described in Fair AI Credit Scoring, but they are not the self-reported categories the rules expect.
Before computing anything, reconcile: rows exported, rows joined, rows with each category known, and the hash of every input file. Join failures that are not random, such as one source system that drops self-ID, can create a disparity or hide one.
Impact-ratio mechanics under Local Law 144
LL144 defines two rates. A selection rate is the share of a category that the tool selects to move forward or assigns a classification to. A scoring rate applies when the tool outputs a score: it is the share of a category that scores above the median score of the whole sample. The impact ratio divides each category's rate by the rate of the most selected or highest-scoring category, so the top group is 1.0. Ratios are computed separately for sex, for race/ethnicity using the EEO-1 categories, and for each sex and race/ethnicity intersection.
A category making up under 2% of the audit data may be left out of the impact-ratio calculation. The auditor must justify the exclusion, and the summary still reports that category's count and rate. The law sets no pass threshold. The common reference is the four-fifths rule from the federal Uniform Guidelines on Employee Selection Procedures, under which a ratio below 0.8 is generally regarded as evidence of adverse impact. That is the screening line used here. The code below is a summary-table generator, not a single ratio function. It tracks unknowns per table, applies the exclusion, and handles both rate types:
def summarize(rows, kind="selection", min_share=0.02):
"""rows: dicts with sex, race and selected (bool) or score (float); None = unknown."""
if kind == "scoring": # share above the median of the whole sample
s = sorted(r["score"] for r in rows)
median = (s[(len(s) - 1) // 2] + s[len(s) // 2]) / 2
hit = lambda r: r["score"] > median
else:
hit = lambda r: bool(r["selected"])
dims = {"sex": lambda r: r["sex"], "race": lambda r: r["race"],
"sex_x_race": lambda r: (r["sex"], r["race"]) if r["sex"] and r["race"] else None}
report = {"kind": kind, "total": len(rows), "tables": {}}
for name, key in dims.items():
groups, unknown = {}, 0
for r in rows:
g = key(r)
if g is None:
unknown += 1
continue
n, k = groups.get(g, (0, 0))
groups[g] = (n + 1, k + hit(r))
known = sum(n for n, _ in groups.values())
keep = {g for g, (n, _) in groups.items() if n / known >= min_share}
best = max(groups[g][1] / groups[g][0] for g in keep)
report["tables"][name] = {"unknown": unknown, "rows": [
{"group": g, "n": n, "rate": round(k / n, 3),
"impact_ratio": round(k / n / best, 3) if g in keep else None,
"excluded_lt_2pct": g not in keep}
for g, (n, k) in sorted(groups.items(), key=lambda x: str(x[0]))]}
return reportTwo interpretation choices are visible in the code and must be written in the report: the 2% share is measured against records with a known category, and the median is taken over the whole sample, unknowns included. A different auditor might choose differently. What matters is that the choice is stated and applied consistently.
Worked example: an intersectional gap
Worked example. A resume screener's advance flag is audited over 1,250 applicants: 20 did not report sex, 29 did not report race/ethnicity, and 49 lack at least one of the two. Running the generator gives these intersectional results, with the sex and race/ethnicity tables computed the same way:
| Category | n | Selection rate | Impact ratio |
|---|---|---|---|
| Male, Asian | 100 | 0.360 | 1.000 |
| Male, White | 340 | 0.350 | 0.972 |
| Female, Asian | 90 | 0.333 | 0.926 |
| Female, White | 300 | 0.320 | 0.889 |
| Male, Hispanic | 60 | 0.317 | 0.880 |
| Male, Black | 110 | 0.300 | 0.833 |
| Female, Hispanic | 70 | 0.257 | 0.714 |
| Female, Black | 120 | 0.250 | 0.694 |
| Female, NHPI | 6 | 0.167 | excluded (under 2%) |
| Male, NHPI | 5 | 0.400 | excluded (under 2%) |
By sex alone the ratio for women is 0.931, which looks fine. By race/ethnicity, Black applicants come out at 0.789 and Hispanic applicants at 0.819. The intersections show where the gap is concentrated: Black women at 0.694 and Hispanic women at 0.714, both well under 0.8. This is why intersectional tables are required. Single-axis averages spread a concentrated disparity thin. The Native Hawaiian and Pacific Islander categories, 11 people in total, are excluded with a stated justification, but their counts and rates still appear in the summary.
Is the gap real? Small cells and power
An impact ratio is a point estimate from finite counts. Is the 0.694 for Black women a real disparity or noise? A one-sided Fisher exact test compares the category with the top category without large-sample approximations:
from math import comb
def fisher_lower(k1, n1, k2, n2):
"""P(group 1 gets k1 or fewer selections if both groups share one rate)."""
K, N = k1 + k2, n1 + n2
return sum(comb(K, i) * comb(N - K, n1 - i) for i in range(k1 + 1)) / comb(N, n1)
fisher_lower(30, 120, 36, 100) # Black women vs Asian men -> 0.0522
fisher_lower(18, 70, 36, 100) # Hispanic women vs Asian men -> 0.1051Neither falls under 0.05. Concluding "no significant disparity" would be a mistake. Simulating a true rate of 0.25 against 0.36 at these sample sizes shows the test detects the gap only about 47% of the time. The audit lacks the power to confirm or rule out a large disparity. The right finding reports both facts: the ratio is under the four-fifths line, and the sample is too small for significance. The recommendation is to investigate the cause now and to pool a longer window for the next audit. Report intervals or exact p-values next to ratios. Never use "not significant" as a pass when power is low, and never use a ratio from a cell of 15 people as a conclusion.
Diagnosis and the evidence pack
The published tables say where a disparity is, not why. A useful audit adds diagnosis. Funnel decomposition computes ratios at each stage, such as parsing, knockout questions and ranking, to find the stage that creates the gap. Cutoff sensitivity matters because the median-based scoring rate may not match how the tool is used. If recruiters only see the top 10%, compute ratios at that cutoff too. A tool can look balanced at the median and be badly skewed at the operational threshold. Paired testing suits LLM-based screeners: submit matched resumes that differ only in names or other group-linked details, and compare outcomes. The method and its pitfalls are covered in the counterfactual section of the fairness article linked above.
Then build the evidence pack. It holds the signed scope memo, input file hashes, row reconciliation, the code at a pinned commit, the environment lockfile, raw outputs, the exclusion justifications, the auditor's independence statement and the dates. Someone should be able to rerun it a year later and get identical numbers. The published summary includes the date of the audit, the data source and explanation, the unknown count, the counts, rates and impact ratios for every category, and the date the tool was first distributed. It must be posted before the tool is used and stay up for at least six months after the employer's latest use of the tool.
Failure modes
- Single-axis only. Sex and race tables both look acceptable while one intersection sits far below 0.8.
- Silent unknowns. A third of candidates decline self-ID and the report never says so, which hides selection bias in who reports.
- Significance as a pass. Small cells produce high p-values, and a large disparity is waved through as noise.
- Wrong decision point. The audit measures the score at the median while recruiters act on the top decile.
- Vendor audit, borrowed. An employer relies on a multi-employer audit without having contributed its own data.
- Irreproducible numbers. The export query was not saved, so nobody can explain why next year's baseline differs.
Trade-offs
Legal minimum or useful audit. The LL144 tables can be produced in a day. Diagnosis and power analysis take weeks. The minimum satisfies the publication duty, and only the full audit tells you what to fix.
Historical or test data. Historical data reflects real use, including real disparities. Test data allows controlled comparisons but may not match the applicant pool. Use historical data for the published numbers and test data for diagnosis.
Independence or depth. An outside auditor is credible but needs access to internals to diagnose. Agree on data and code access in the engagement letter, not mid-audit.
Window length. Longer windows give power but mix tool versions. Audit per version where volumes allow, and pool only across versions with unchanged logic. The rights-layer design in AI Bill of Rights and Regulation shows how to log decisions so this is possible.
What to do next
- Write and sign a scope memo: decision point, output type, population, window, categories, flag rule.
- Inventory where self-reported demographics live, and measure the unknown rate before the audit starts.
- Run the summary generator on sex, race/ethnicity and intersections, and record the interpretation choices.
- Add exact tests and a power estimate to every flagged or borderline category.
- Decompose the funnel and compute ratios at the cutoff where humans actually act.
- Assemble the evidence pack with hashes, pinned code and an independence statement, and rerun it from scratch.
- Publish the summary, assign each finding an owner, and schedule the next audit before the one-year deadline.