An AI ethics officer is the person who owns one question for an organisation: should we build or deploy this AI system, and under what conditions? Security asks whether the system can be attacked, privacy asks whether personal data is handled lawfully, and legal asks whether it complies with the law. The ethics officer asks whether it is fair, honest and acceptable to the people it affects, and then turns the answer into conditions a team can ship against.
The role fails in two familiar ways. It becomes a figurehead who writes principles nobody reads, or a bottleneck who reviews everything and slows every team. This article treats the role as an engineering function. It covers what the officer owns and what they do not, the decision rights that make the role real, intake and tiering as code, a worked review, the artifacts and metrics to keep, and the failure modes to design out. It complements the ethics board and governance council articles, which cover the bodies this person works with.
What the role owns, and what it does not
The officer is one accountable individual. A board deliberates and a council allocates decision rights across functions, but someone has to own the queue, write the decisions and follow up. The clearest way to define the role is against its neighbours.
| Role | Owns | Ethics officer relationship |
|---|---|---|
| CISO | Attack surface, security controls, incident response | Consulted on misuse and abuse; shared incident process |
| Privacy lead / DPO | Lawful basis, data minimisation, data subject rights | Joint review when personal data is used to train or infer |
| Legal / compliance | Regulatory obligations, contracts, liability | Legal says what is required; the officer decides what is acceptable above that |
| Product owner | The use case, its benefits and its delivery | Accountable for meeting conditions; can escalate a decision |
| Ethics board | Hard cases, policy, appeals | Officer brings cases to it and executes its decisions |
| AI ethics officer | Tiering, reviews, conditions, decision records, metrics | Single owner of the ethics review pipeline |
There is usually no legal requirement for the title. The EU AI Act, for example, does not create an ethics officer post the way GDPR creates a data protection officer. It does put obligations on organisations that someone must own: its AI literacy duty in Article 4 has applied since February 2025, and its high-risk obligations need risk management and documentation. ISO/IEC 42001 clause 5.3 requires roles, responsibilities and authorities for the AI management system to be assigned, and the NIST AI RMF's GOVERN function asks for documented accountability structures. In the US federal government, agencies have been required since 2024 to designate a Chief AI Officer, a role that stayed in place when the original OMB memo was replaced in 2025. An ethics officer is a common way to meet the accountability part of these frameworks, but the frameworks do not require that title.
Decision rights and reporting line
A role without decision rights is advice. Write the mandate down in a charter signed by an executive sponsor, and make three rights explicit.
- The right to be asked. Every AI use case above the lowest tier goes through intake before build, not after launch. A release gate in the delivery pipeline enforces this, so it does not depend on goodwill.
- The right to set conditions. Approval can carry conditions such as a human review step, a bias evaluation threshold, a disclosure, or a restricted population. Each condition has an owner and a check.
- The right to stop, with a clock. The officer can block release of a high-tier system. The product owner can appeal to the ethics board or the executive sponsor, who must decide within a fixed time, for example ten working days. Without the clock, a block becomes a silent veto, and teams learn to route around it.
Reporting line matters as much as the rights. An officer who reports to the head of the product line they review has a conflict of interest built in. Reporting to the general counsel, the chief risk officer or the CEO, with a dotted line to the board, keeps the role independent of any one product's targets. A CISO in the same position faces a similar trade-off, covered in the CISO role article.
Intake and tiering as code
Most use cases are low risk, and reviewing them in depth is how the role becomes a bottleneck. Encode tiering as rules so the easy cases clear themselves, and review effort goes where harm is plausible. The rules below are an example to adapt, not a standard.
from dataclasses import dataclass
@dataclass
class UseCase:
name: str
affects_people_decisions: bool # hiring, credit, health, housing, education, benefits
monitors_workers: bool # performance or behaviour evaluation
automated_final_decision: bool # no human review before the outcome
vulnerable_users: bool # minors, patients, people in crisis
generates_public_content: bool
personal_data: bool
def tier(u: UseCase) -> int:
"""1 = self-serve checklist, 2 = officer review, 3 = officer + board."""
if u.affects_people_decisions or u.monitors_workers or u.vulnerable_users:
return 3 if u.automated_final_decision else 2
if u.generates_public_content or u.personal_data:
return 2
return 1
REQUIRED = {
1: ["model card link", "owner named"],
2: ["impact assessment", "evaluation plan", "disclosure text", "rollback plan"],
3: ["impact assessment", "bias evaluation by group", "human review design",
"appeal route for affected people", "monitoring metrics", "board sign-off"],
}Keep the rules in version control next to the release gate, so a change to tiering is itself reviewed and visible. Review the tier-1 decisions every quarter: if incidents come from use cases that cleared themselves, the rules are too loose.
The review workflow
The review itself is a structured conversation, not an essay. The officer runs it, brings in privacy, security and legal where the tiering flags say so, and ends with a decision record. A review that does not end with a record did not happen.
Worked example: scoring contact-centre agents
A contact-centre team proposes a model that summarises each support call and scores the agent on empathy, accuracy and resolution. Scores will feed a weekly dashboard for team leads. The intake form marks it as monitoring workers and using personal data, with a human (the team lead) in the loop, so the rules return tier 2. The officer bumps it to tier 3 by judgement, because the scores will plainly inform performance reviews. The EU AI Act lists AI used to monitor and evaluate the performance and behaviour of workers among its high-risk use cases, so legal is consulted as well.
The review finds three issues. First, the empathy score is a model judgement with no validated link to the outcome the business cares about. Second, accent and speech patterns could bias transcription accuracy, and so the accuracy score. Third, agents have no way to see or contest a score. The decision is approval with conditions, recorded like this:
decision: approve-with-conditions
use_case: call-summary-agent-scoring
tier: 3
conditions:
- id: C1 # drop the unvalidated metric
text: Remove the empathy score from the dashboard until validated against outcomes.
owner: product-owner
check: dashboard schema has no empathy field
- id: C2 # measure the bias risk
text: Report word error rate and score distributions by accent group before launch.
owner: ml-lead
check: eval report attached; max group gap under agreed threshold
- id: C3 # contestability
text: Agents see their own scores and can flag a call for human re-review.
owner: product-owner
check: flag button live; re-review queue staffed
- id: C4 # no automated consequences
text: Scores cannot trigger performance actions without a lead reviewing the calls.
owner: hr-partner
check: written HR policy reference
review_date: six months after launch
escalation: none (product owner accepted)Each condition is a test the release gate or an auditor can check, rather than a sentiment. That is what separates an ethics function that changes systems from one that comments on them.
Metrics for the function
Measure the function the way you would measure any service, and publish the numbers to the board.
- Coverage: share of production AI systems with a tier and a decision record. Find the denominator through the AI inventory, not self-reporting.
- Latency: median and 90th-percentile days from intake to decision, by tier. A rising tier-2 latency is the first sign of a bottleneck.
- Condition closure: share of conditions verified closed by their due date.
- Escalations and overturns: how often decisions are appealed, and how often the board overturns them. Zero appeals can mean deference rather than agreement.
- Incident linkage: incidents traced to systems that were reviewed, reviewed with conditions, or never reviewed. This is the closest thing to an outcome metric.
Failure modes
- Ethics washing. Principles are published but nothing is ever blocked or conditioned. Count conditions issued; if the number is zero, the role is decoration.
- Bottleneck. Every case gets a deep review, so teams wait weeks and start skipping intake. Fix the tiering, and train embedded reviewers in large product groups.
- Late engagement. Review happens a week before launch, when only cosmetic changes are possible. Move intake to design time and tie it to funding or project approval.
- Conflict of interest. The officer reports into the product organisation they review, or is measured on launch velocity.
- Conditions without owners. Approvals carry vague expectations that nobody checks. Every condition needs an owner, a check and a date.
- Single point of failure. One person holds all the context. Keep decision records searchable, and name a deputy.
- Scope creep into security or legal. The officer re-does work other functions own and creates conflicting guidance. Use the scope table and consult instead.
Skills and team shape
Good officers combine enough technical depth to read an evaluation report and challenge a metric, enough legal literacy to know when to bring in counsel, and the facilitation skills to run a review that ends in a decision. Hiring for only one of these produces a philosopher nobody consults, an engineer who misses social harm, or a lawyer who treats compliance as the ceiling. In larger organisations the officer leads a small team and a network of trained reviewers embedded in product groups, who handle tier 2 and bring tier 3 cases forward. For the principles the reviews apply, see AI ethics in practice.
A workable operating rhythm keeps the function predictable for the teams it serves. Hold a weekly intake triage where new cases are tiered and assigned, with a published service level for each tier. Run a monthly review of open conditions with their owners, and close or escalate anything overdue. Each quarter, report the metrics to the board, recalibrate the tiering rules against incidents, and refresh the AI literacy training the reviewers and product teams rely on. After any serious incident involving an AI system, run a blameless review that asks whether the case was tiered correctly, whether the conditions were the right ones, and whether they were actually met.
What to do next
- Write a one-page charter with the three decision rights, the reporting line and the escalation clock, and get it signed by an executive sponsor.
- Build or borrow an AI inventory so you know the denominator for coverage.
- Encode the tiering rules, put them in version control, and connect them to the release gate.
- Create a decision record template with conditions that each have an owner, a check and a date.
- Run three reviews on existing systems to calibrate tiering before taking new intake.
- Publish coverage, latency, condition closure and incident linkage to the board every quarter.
- Agree the consult boundaries with the CISO, privacy lead and legal in writing.