Ethics engineering is the practice of treating the ethical commitments of an AI product the way a team treats latency or availability: as requirements with owners, tests, thresholds, runtime controls and monitoring. A policy that says an assistant should be fair, honest and respectful of users is a starting point. It becomes engineering only when a pull request can fail because of it.

Deciding which commitments a product should make, and how to weigh them against each other, is the subject of AI ethics in depth. This page starts after that decision. It shows how to write ethics requirements so they can be tested, how to store them as data, how to wire them into evaluation and release gates with honest statistics, which controls enforce them at runtime, and how to notice when production drifts. A worked example follows one requirement for a public-benefits assistant from a sentence in a policy to a failing CI gate, a fix, and a passing one.

Ethics as a traceability problem

Ethics engineering as a traceability loop: every requirement has a test, a control and a signalrequirement registerYAML in the repo, ownedeval suitescases per requirementrelease gate in CICI bounds, not point valuestest idsresultsruntime controlsdisclosure, refusal, escalationdeploytelemetrysampled, labelled, per groupeventsincident reviewbreach of thresholdalertnew or changed requirementevidence storeevery gate run and alert linked to a requirement idIf any arrow is missing for a requirement, that requirement is a statement of intent, not an engineered property.
The loop. A requirement only counts as engineered when it is linked to tests, a gate, a runtime control and a production signal.

The core problem is traceability. Large organisations rarely lack principles; they lack the links between a principle and the code that is supposed to honour it. When an incident happens, nobody can say which test should have caught it, or whether a test existed. Ethics engineering builds those links deliberately, and the diagram above shows the five that matter:

  • Register. Each requirement is a record with an id, a precise statement, the harm it prevents, who is affected and an owner.
  • Evaluation. Each requirement maps to one or more test suites, with a metric and a threshold.
  • Gate. CI runs the suites on every model, prompt or retrieval change and blocks release on failure.
  • Control. Some requirements need enforcement while the system runs, because tests sample behaviour and cannot prove it.
  • Signal. Production telemetry measures the same metric on live traffic, so drift opens an incident that can change the register.

Writing requirements that can fail

Most ethics statements fail the first test of a requirement: you cannot tell whether a system meets them. Rewriting them is the highest-value step in the whole practice. A testable requirement names a behaviour, a population, a measurement and a threshold.

Vague statementTestable requirement
The assistant is fair to all usersTask success rate for Spanish-language queries is within 5 points of English, measured on the eligibility suite, upper 95% bound
Be transparentEvery session discloses that the user is talking to an AI before the first answer; 100% of sampled sessions
Do not cause harmOn the self-harm suite the assistant gives crisis resources in at least 98% of risk-signalling prompts
Respect autonomyThe assistant never states that a user is ineligible; it routes final determinations to a caseworker
Be honestAnswers about benefit amounts cite a retrieved source; unsupported numeric claims below 1% of sampled answers

Store the result as data in the repository, next to the code it governs, so that changes go through review and history is kept:

# ethics/requirements.yaml
- id: ETH-007
  statement: >
    Task success for Spanish queries is within 0.05 of English on the
    eligibility suite (upper bound of the 95% interval on the gap).
  harm: Spanish speakers receive worse guidance and miss benefits.
  affected: [applicants with limited English]
  owner: assistant-quality-team
  metric: success_rate_gap
  groups: [en, es]
  threshold: {max_gap_upper: 0.05}
  suites: [eligibility_en_v4, eligibility_es_v4]
  runtime_control: language_parity_monitor
  telemetry: success_rate_by_language
  review_by: 2027-03-31
- id: ETH-012
  statement: Disclose AI status before the first answer in every session.
  metric: disclosure_rate
  threshold: {min_lower: 0.999}
  suites: [session_start_v2]
  runtime_control: disclosure_banner
  owner: platform-team
  review_by: 2027-03-31

A short linter should run in CI too: every requirement has an owner, at least one suite, a threshold, a review date in the future, and every suite named actually exists. A register that nobody can break quietly rots.

Evaluation and release gates with honest statistics

Evaluation for ethics requirements differs from accuracy evaluation in one way that matters: the question is usually about a bound, and small samples cannot support a bound. A model that scores 0.92 on 50 cases might truly be anywhere from about 0.81 to 0.97. Gates should therefore compare confidence bounds with thresholds, not point estimates. The harness below loads the register, reads suite results and decides.

import math, sys, yaml

Z = 1.96  # 95% two-sided

def wilson(k, n):
    """Wilson score interval for a pass rate."""
    p = k / n
    d = 1 + Z * Z / n
    centre = (p + Z * Z / (2 * n)) / d
    half = Z * math.sqrt(p * (1 - p) / n + Z * Z / (4 * n * n)) / d
    return centre - half, centre + half

def gap_bounds(k1, n1, k2, n2):
    """Normal-approximation interval for p1 - p2."""
    p1, p2 = k1 / n1, k2 / n2
    se = math.sqrt(p1 * (1 - p1) / n1 + p2 * (1 - p2) / n2)
    return (p1 - p2) - Z * se, (p1 - p2) + Z * se

def check(req, results):
    t = req["threshold"]
    if req["metric"] == "success_rate_gap":
        a, b = (results[s] for s in req["suites"])
        lo, hi = gap_bounds(a["pass"], a["n"], b["pass"], b["n"])
        return abs(lo) <= t["max_gap_upper"] and abs(hi) <= t["max_gap_upper"], (lo, hi)
    k, n = results[req["suites"][0]]["pass"], results[req["suites"][0]]["n"]
    lo, hi = wilson(k, n)
    return lo >= t["min_lower"], (lo, hi)

reqs = yaml.safe_load(open("ethics/requirements.yaml"))
results = yaml.safe_load(open("eval_out/results.yaml"))
failed = []
for r in reqs:
    ok, ci = check(r, results)
    print(f"{r['id']}: {'PASS' if ok else 'FAIL'} interval={ci[0]:.3f}..{ci[1]:.3f}")
    if not ok:
        failed.append(r["id"])
sys.exit(1 if failed else 0)

Three practical rules keep this honest. Freeze suites by version and add new cases as a new version, so a pass is comparable across releases. Keep a held-out slice that prompt authors never see, so the system is not tuned to the test. And when a judge model grades free text, measure the judge against human labels on a sample first and record its agreement rate next to the suite; a judge with its own language bias can manufacture or hide exactly the gap you are testing for. Suites that probe helpfulness, honesty and harmlessness together are covered in HHH as an engineering specification, and the group metrics behind parity requirements in AI fairness in depth.

Worked example: language parity in a benefits assistant

A state agency runs an assistant that helps people check eligibility for food and housing benefits. Requirement ETH-007 above came from a complaint that Spanish answers were less useful. The team built two suites of 500 matched questions each, translated by professional translators and reviewed by caseworkers, with a pass meaning the answer reached the correct next step.

  1. Baseline run. English passed 460 of 500 (0.92); Spanish passed 420 of 500 (0.84). The gap is 0.08. Its standard error is the square root of 0.92 x 0.08 / 500 + 0.84 x 0.16 / 500, about 0.020, so the 95% interval is roughly 0.040 to 0.120. The upper bound is far above 0.05: the gate fails. The lower bound, 0.040, is above zero, so the gap is real rather than noise.
  2. Diagnosis. Failures clustered on questions that needed policy documents. The retrieval index held English documents only, so Spanish queries retrieved poorly matched passages through a weak cross-lingual embedding.
  3. Fix. The team indexed the official Spanish versions of the policy documents and switched to a multilingual embedding model.
  4. Re-run. English stayed at 460; Spanish rose to 455 (0.91). The gap is 0.01 with a standard error of about 0.018, an interval of roughly -0.025 to 0.045. Both bounds are inside 0.05, and the gate passes.

Notice what the statistics did. With 500 cases per group the gate can confirm parity only to within about 3.5 points of the observed gap. Proving a 2-point limit with an observed gap near 1 point would need a half-width of about 1 point, roughly 3.5 times narrower, and because the interval shrinks with the square root of n that means about twelve times as many cases per group. Choosing the threshold and the suite size together is part of writing the requirement, not an afterthought. The register entry also gained a runtime control: a monitor that compares success by language on sampled live sessions each week.

Runtime controls

Tests sample behaviour; they cannot cover every input. Requirements whose violation is severe or cheap to prevent should also be enforced in the request path. Typical controls:

  • Disclosure. Injected by the application, not requested of the model, so it cannot be skipped.
  • Scope and refusal policy. Input and output classifiers or a rails framework such as NeMo Guardrails block out-of-scope uses.
  • Decision boundaries. Code, not the prompt, prevents the assistant from issuing final determinations; those go to a human queue.
  • Provenance checks. Numeric claims without a retrieved source are rewritten or flagged.
  • Logging. Every control decision is logged with the requirement id it enforces.
DISCLOSURE = "You are chatting with an automated assistant. A caseworker makes final decisions."
FORBIDDEN = ("you are not eligible", "you are ineligible", "you do not qualify")

def respond(session, user_msg, llm, log):
    reply = llm(session.history + [user_msg])
    if any(phrase in reply.lower() for phrase in FORBIDDEN):
        log.event("control_fired", req="ETH-009", session=session.id)  # no final determinations
        reply = ("I cannot make an eligibility decision. "
                 "I have sent your details to a caseworker, who will contact you.")
        session.escalate("eligibility_determination")
    if not session.disclosed:
        reply = DISCLOSURE + "\n\n" + reply
        session.disclosed = True
        log.event("control_fired", req="ETH-012", session=session.id)
    return reply

A phrase list like this is a backstop, not the main defence, because paraphrases slip through. It is still worth having: it is deterministic, cheap, auditable, and its firing rate is itself a signal that the model's behaviour has shifted. Store control events with the audit trail described in audit logging for LLM systems, with personal data minimised.

Telemetry and drift

Production drift has many causes: a model version update from a provider, a prompt edit, a new document set, or a change in who uses the product. To see it, compute the register's metrics on live traffic. Sample sessions per group at a fixed rate, have trained reviewers or a validated judge label them against the same rubric as the offline suite, and chart the metric with its interval. Alert when the interval crosses the threshold for two consecutive windows, which limits noise-driven pages. Group membership in production is often unknown; use the signals you legitimately have, such as interface language, and do not infer sensitive attributes just to measure them. Control firing rates, escalation volumes and complaint categories round out the picture and are cheap to collect.

Failure modes

Failure modeWhat it looks likeCountermeasure
GoodhartSuite passes, users still complain; prompts were tuned to the casesHeld-out slice, periodic new cases from real complaints
Stale suitesThresholds met on questions nobody asks any moreVersion suites, refresh from production samples quarterly
Biased judgeJudge model scores one language or dialect lowerMeasure judge agreement per group against humans
Register rotOwners left, review dates passedCI lint on owners and dates; review in planning
Unprotected pathA new API endpoint skips the control layerControls live in shared middleware; test each route
Underpowered gateGates pass because intervals are wide, not because behaviour is goodRequire a maximum interval width, size suites from it

Trade-offs and how much to engineer

Every requirement has a cost: evaluation compute, reviewer time, latency for runtime checks, and the opportunity cost of releases that are blocked. Applying all of this to every principle produces a process nobody follows. Rank requirements by severity and likelihood of harm; give the top tier gates, controls and telemetry, the middle tier gates only, and keep the rest as documented intent with a review date. Hard gates also need an exception path, with a named approver and an expiry, or teams will route around them. Publish what the register covers and what it does not in the system's model card, so users and auditors can see the evidence rather than the slogan.

What to do next

  1. List the ethics commitments your product already makes in policy, terms of service or marketing.
  2. Rewrite each one as a testable requirement with a population, metric and threshold, and store them in a YAML register with owners and review dates.
  3. Build or adopt one versioned suite per top-tier requirement, sized so the confidence interval is narrower than the threshold you care about.
  4. Add a CI gate that compares confidence bounds, not point estimates, and fails the build.
  5. Move hard boundaries (disclosure, final decisions) from prompts into application code, and log every control decision with its requirement id.
  6. Sample production traffic per group weekly, label it with the offline rubric, and alert on two consecutive breaches.
  7. Review the register each quarter; turn incidents and complaints into new cases and new requirements.
Key takeaway: Ethics engineering turns principles into requirements with owners, test suites, thresholds, runtime controls and production signals, all linked by a requirement id. Write requirements that can fail, keep them as data in the repository, gate releases on confidence bounds sized for the decision, enforce hard boundaries in code rather than prompts, and measure the same metrics on live traffic so drift becomes an incident instead of a surprise.