Ethics engineering is the practice of treating the ethical commitments of an AI product the way a team treats latency or availability: as requirements with owners, tests, thresholds, runtime controls and monitoring. A policy that says an assistant should be fair, honest and respectful of users is a starting point. It becomes engineering only when a pull request can fail because of it.
Deciding which commitments a product should make, and how to weigh them against each other, is the subject of AI ethics in depth. This page starts after that decision. It shows how to write ethics requirements so they can be tested, how to store them as data, how to wire them into evaluation and release gates with honest statistics, which controls enforce them at runtime, and how to notice when production drifts. A worked example follows one requirement for a public-benefits assistant from a sentence in a policy to a failing CI gate, a fix, and a passing one.
Ethics as a traceability problem
The core problem is traceability. Large organisations rarely lack principles; they lack the links between a principle and the code that is supposed to honour it. When an incident happens, nobody can say which test should have caught it, or whether a test existed. Ethics engineering builds those links deliberately, and the diagram above shows the five that matter:
- Register. Each requirement is a record with an id, a precise statement, the harm it prevents, who is affected and an owner.
- Evaluation. Each requirement maps to one or more test suites, with a metric and a threshold.
- Gate. CI runs the suites on every model, prompt or retrieval change and blocks release on failure.
- Control. Some requirements need enforcement while the system runs, because tests sample behaviour and cannot prove it.
- Signal. Production telemetry measures the same metric on live traffic, so drift opens an incident that can change the register.
Writing requirements that can fail
Most ethics statements fail the first test of a requirement: you cannot tell whether a system meets them. Rewriting them is the highest-value step in the whole practice. A testable requirement names a behaviour, a population, a measurement and a threshold.
| Vague statement | Testable requirement |
|---|---|
| The assistant is fair to all users | Task success rate for Spanish-language queries is within 5 points of English, measured on the eligibility suite, upper 95% bound |
| Be transparent | Every session discloses that the user is talking to an AI before the first answer; 100% of sampled sessions |
| Do not cause harm | On the self-harm suite the assistant gives crisis resources in at least 98% of risk-signalling prompts |
| Respect autonomy | The assistant never states that a user is ineligible; it routes final determinations to a caseworker |
| Be honest | Answers about benefit amounts cite a retrieved source; unsupported numeric claims below 1% of sampled answers |
Store the result as data in the repository, next to the code it governs, so that changes go through review and history is kept:
# ethics/requirements.yaml
- id: ETH-007
statement: >
Task success for Spanish queries is within 0.05 of English on the
eligibility suite (upper bound of the 95% interval on the gap).
harm: Spanish speakers receive worse guidance and miss benefits.
affected: [applicants with limited English]
owner: assistant-quality-team
metric: success_rate_gap
groups: [en, es]
threshold: {max_gap_upper: 0.05}
suites: [eligibility_en_v4, eligibility_es_v4]
runtime_control: language_parity_monitor
telemetry: success_rate_by_language
review_by: 2027-03-31
- id: ETH-012
statement: Disclose AI status before the first answer in every session.
metric: disclosure_rate
threshold: {min_lower: 0.999}
suites: [session_start_v2]
runtime_control: disclosure_banner
owner: platform-team
review_by: 2027-03-31A short linter should run in CI too: every requirement has an owner, at least one suite, a threshold, a review date in the future, and every suite named actually exists. A register that nobody can break quietly rots.
Evaluation and release gates with honest statistics
Evaluation for ethics requirements differs from accuracy evaluation in one way that matters: the question is usually about a bound, and small samples cannot support a bound. A model that scores 0.92 on 50 cases might truly be anywhere from about 0.81 to 0.97. Gates should therefore compare confidence bounds with thresholds, not point estimates. The harness below loads the register, reads suite results and decides.
import math, sys, yaml
Z = 1.96 # 95% two-sided
def wilson(k, n):
"""Wilson score interval for a pass rate."""
p = k / n
d = 1 + Z * Z / n
centre = (p + Z * Z / (2 * n)) / d
half = Z * math.sqrt(p * (1 - p) / n + Z * Z / (4 * n * n)) / d
return centre - half, centre + half
def gap_bounds(k1, n1, k2, n2):
"""Normal-approximation interval for p1 - p2."""
p1, p2 = k1 / n1, k2 / n2
se = math.sqrt(p1 * (1 - p1) / n1 + p2 * (1 - p2) / n2)
return (p1 - p2) - Z * se, (p1 - p2) + Z * se
def check(req, results):
t = req["threshold"]
if req["metric"] == "success_rate_gap":
a, b = (results[s] for s in req["suites"])
lo, hi = gap_bounds(a["pass"], a["n"], b["pass"], b["n"])
return abs(lo) <= t["max_gap_upper"] and abs(hi) <= t["max_gap_upper"], (lo, hi)
k, n = results[req["suites"][0]]["pass"], results[req["suites"][0]]["n"]
lo, hi = wilson(k, n)
return lo >= t["min_lower"], (lo, hi)
reqs = yaml.safe_load(open("ethics/requirements.yaml"))
results = yaml.safe_load(open("eval_out/results.yaml"))
failed = []
for r in reqs:
ok, ci = check(r, results)
print(f"{r['id']}: {'PASS' if ok else 'FAIL'} interval={ci[0]:.3f}..{ci[1]:.3f}")
if not ok:
failed.append(r["id"])
sys.exit(1 if failed else 0)Three practical rules keep this honest. Freeze suites by version and add new cases as a new version, so a pass is comparable across releases. Keep a held-out slice that prompt authors never see, so the system is not tuned to the test. And when a judge model grades free text, measure the judge against human labels on a sample first and record its agreement rate next to the suite; a judge with its own language bias can manufacture or hide exactly the gap you are testing for. Suites that probe helpfulness, honesty and harmlessness together are covered in HHH as an engineering specification, and the group metrics behind parity requirements in AI fairness in depth.
Worked example: language parity in a benefits assistant
A state agency runs an assistant that helps people check eligibility for food and housing benefits. Requirement ETH-007 above came from a complaint that Spanish answers were less useful. The team built two suites of 500 matched questions each, translated by professional translators and reviewed by caseworkers, with a pass meaning the answer reached the correct next step.
- Baseline run. English passed 460 of 500 (0.92); Spanish passed 420 of 500 (0.84). The gap is 0.08. Its standard error is the square root of 0.92 x 0.08 / 500 + 0.84 x 0.16 / 500, about 0.020, so the 95% interval is roughly 0.040 to 0.120. The upper bound is far above 0.05: the gate fails. The lower bound, 0.040, is above zero, so the gap is real rather than noise.
- Diagnosis. Failures clustered on questions that needed policy documents. The retrieval index held English documents only, so Spanish queries retrieved poorly matched passages through a weak cross-lingual embedding.
- Fix. The team indexed the official Spanish versions of the policy documents and switched to a multilingual embedding model.
- Re-run. English stayed at 460; Spanish rose to 455 (0.91). The gap is 0.01 with a standard error of about 0.018, an interval of roughly -0.025 to 0.045. Both bounds are inside 0.05, and the gate passes.
Notice what the statistics did. With 500 cases per group the gate can confirm parity only to within about 3.5 points of the observed gap. Proving a 2-point limit with an observed gap near 1 point would need a half-width of about 1 point, roughly 3.5 times narrower, and because the interval shrinks with the square root of n that means about twelve times as many cases per group. Choosing the threshold and the suite size together is part of writing the requirement, not an afterthought. The register entry also gained a runtime control: a monitor that compares success by language on sampled live sessions each week.
Runtime controls
Tests sample behaviour; they cannot cover every input. Requirements whose violation is severe or cheap to prevent should also be enforced in the request path. Typical controls:
- Disclosure. Injected by the application, not requested of the model, so it cannot be skipped.
- Scope and refusal policy. Input and output classifiers or a rails framework such as NeMo Guardrails block out-of-scope uses.
- Decision boundaries. Code, not the prompt, prevents the assistant from issuing final determinations; those go to a human queue.
- Provenance checks. Numeric claims without a retrieved source are rewritten or flagged.
- Logging. Every control decision is logged with the requirement id it enforces.
DISCLOSURE = "You are chatting with an automated assistant. A caseworker makes final decisions."
FORBIDDEN = ("you are not eligible", "you are ineligible", "you do not qualify")
def respond(session, user_msg, llm, log):
reply = llm(session.history + [user_msg])
if any(phrase in reply.lower() for phrase in FORBIDDEN):
log.event("control_fired", req="ETH-009", session=session.id) # no final determinations
reply = ("I cannot make an eligibility decision. "
"I have sent your details to a caseworker, who will contact you.")
session.escalate("eligibility_determination")
if not session.disclosed:
reply = DISCLOSURE + "\n\n" + reply
session.disclosed = True
log.event("control_fired", req="ETH-012", session=session.id)
return replyA phrase list like this is a backstop, not the main defence, because paraphrases slip through. It is still worth having: it is deterministic, cheap, auditable, and its firing rate is itself a signal that the model's behaviour has shifted. Store control events with the audit trail described in audit logging for LLM systems, with personal data minimised.
Telemetry and drift
Production drift has many causes: a model version update from a provider, a prompt edit, a new document set, or a change in who uses the product. To see it, compute the register's metrics on live traffic. Sample sessions per group at a fixed rate, have trained reviewers or a validated judge label them against the same rubric as the offline suite, and chart the metric with its interval. Alert when the interval crosses the threshold for two consecutive windows, which limits noise-driven pages. Group membership in production is often unknown; use the signals you legitimately have, such as interface language, and do not infer sensitive attributes just to measure them. Control firing rates, escalation volumes and complaint categories round out the picture and are cheap to collect.
Failure modes
| Failure mode | What it looks like | Countermeasure |
|---|---|---|
| Goodhart | Suite passes, users still complain; prompts were tuned to the cases | Held-out slice, periodic new cases from real complaints |
| Stale suites | Thresholds met on questions nobody asks any more | Version suites, refresh from production samples quarterly |
| Biased judge | Judge model scores one language or dialect lower | Measure judge agreement per group against humans |
| Register rot | Owners left, review dates passed | CI lint on owners and dates; review in planning |
| Unprotected path | A new API endpoint skips the control layer | Controls live in shared middleware; test each route |
| Underpowered gate | Gates pass because intervals are wide, not because behaviour is good | Require a maximum interval width, size suites from it |
Trade-offs and how much to engineer
Every requirement has a cost: evaluation compute, reviewer time, latency for runtime checks, and the opportunity cost of releases that are blocked. Applying all of this to every principle produces a process nobody follows. Rank requirements by severity and likelihood of harm; give the top tier gates, controls and telemetry, the middle tier gates only, and keep the rest as documented intent with a review date. Hard gates also need an exception path, with a named approver and an expiry, or teams will route around them. Publish what the register covers and what it does not in the system's model card, so users and auditors can see the evidence rather than the slogan.
What to do next
- List the ethics commitments your product already makes in policy, terms of service or marketing.
- Rewrite each one as a testable requirement with a population, metric and threshold, and store them in a YAML register with owners and review dates.
- Build or adopt one versioned suite per top-tier requirement, sized so the confidence interval is narrower than the threshold you care about.
- Add a CI gate that compares confidence bounds, not point estimates, and fails the build.
- Move hard boundaries (disclosure, final decisions) from prompts into application code, and log every control decision with its requirement id.
- Sample production traffic per group weekly, label it with the offline rubric, and alert on two consecutive breaches.
- Review the register each quarter; turn incidents and complaints into new cases and new requirements.