Some actions should not happen just because software decided they should: wiring money, deleting production data, changing a patient's medication record, rejecting a loan, granting a contractor admin access. As AI systems move from drafting text to taking actions, organisations put a person in the path of these actions. In many deployments that person signs off on whatever is put in front of them. Placing a human in the loop is easy. Making the human an effective control is the hard part.

This article is about that second part, for high-risk actions specifically. The general approval-gate architecture for agents is covered in human-in-the-loop approval gates, and the design of interactive confirmation dialogs in agent permission prompts. Here the focus is the decisions those pieces leave open: which actions need a human, what the human must see, how to tell whether reviewers are catching anything, and how to staff a review queue so it does not collapse into rubber-stamping.

A high-risk action must pass a policy, a person and a re-check before it runsAgent / systemproposes actionRisk policyscore -> routeAuto-executelow risk, loggedBlocknever allowedlowprohibitedEvidence packetdiff, inputs, whyhighReview queueSLA, routing, canariesReviewer(s)1 or 2, not requesterSigned approvalbound to action hashapproveDeny / expiredefault on timeoutExecutorre-checks hash, TTLAudit logwho, what, evidenceOversight metricscanary catch rate, latency, override rate
The policy routes each proposed action. High-risk ones get a platform-built evidence packet, a reviewer who is not the requester, a signed approval bound to the action hash, and an executor that re-checks it.

What makes an action high risk

Define high risk by consequences, not by how unusual the action looks. Four properties carry most of the weight:

  • Reversibility. A refund can be clawed back, a sent email cannot, a dropped table without a backup is gone. If the action is irreversible, you get one chance to stop it. Designing actions to be undoable (see reversibility design) is often cheaper than reviewing them.
  • Blast radius. One customer or every customer, one row or the whole table, one host or the fleet.
  • Magnitude. Money moved, privileges granted, records exposed.
  • Legal and human significance. Decisions with legal or similarly significant effects on people, such as credit, employment, benefits or medical treatment, carry legal oversight duties in many jurisdictions regardless of their technical size.

A human review step is the right control when the action is consequential and the system's error rate is not yet known well enough to trust, or when the law requires a human decision. It is the wrong control when the volume is so high that review becomes a formality, or when a deterministic rule (a hard limit, an allowlist) would block the bad case more reliably. Prefer engineering controls and spend human attention only where judgement adds something.

What makes an action high risk

Define high risk by consequences, not by how unusual the action looks. Four properties carry most of the weight:

  • Reversibility. A refund can be clawed back, a sent email cannot, a dropped table without a backup is gone. If the action is irreversible, you get one chance to stop it. Designing actions to be undoable (see reversibility design) is often cheaper than reviewing them.
  • Blast radius. One customer or every customer, one row or the whole table, one host or the fleet.
  • Magnitude. Money moved, privileges granted, records exposed.
  • Legal and human significance. Decisions with legal or similarly significant effects on people, such as credit, employment, benefits or medical treatment, carry legal oversight duties in many jurisdictions regardless of their technical size.

A human review step is the right control when the action is consequential and the system's error rate is not yet known well enough to trust, or when the law requires a human decision. It is the wrong control when the volume is so high that review becomes a formality, or when a deterministic rule (a hard limit, an allowlist) would block the bad case more reliably. Prefer engineering controls and spend human attention only where judgement adds something.

Routing: deciding when a human is required

Routing must be deterministic and versioned, so the same action always gets the same treatment and you can explain any decision afterwards. Express it as code that scores the proposed action and returns one of four routes: auto, one reviewer, two reviewers, or block.

from dataclasses import dataclass

@dataclass(frozen=True)
class Action:
    tool: str            # "refund", "delete_resource", "grant_role", ...
    amount: float = 0.0  # money moved, if any
    targets: int = 1     # records / hosts / accounts affected
    reversible: bool = True
    env: str = "prod"
    tainted_input: bool = False   # plan influenced by untrusted content?

BLOCKED = {("grant_role", "org_admin")}
POLICY_VERSION = "hitl-2026-10-01"

def route(a: Action, role: str = "") -> str:
    if (a.tool, role) in BLOCKED:
        return "block"
    score = 0
    score += 3 if not a.reversible else 0
    score += 2 if a.amount >= 5_000 else 1 if a.amount >= 500 else 0
    score += 2 if a.targets >= 100 else 1 if a.targets >= 10 else 0
    score += 1 if a.env == "prod" else 0
    score += 2 if a.tainted_input else 0
    if score >= 6:
        return "two_reviewers"
    if score >= 3:
        return "one_reviewer"
    return "auto"

The thresholds are placeholders. Set yours from history: replay last quarter's actions through the policy and look at how many would have gone to review and which known incidents would have been caught. The tainted_input flag deserves attention. If the agent read an email, a web page or a ticket written by an outsider before proposing the action, prompt injection is a live possibility, and the action should get more scrutiny regardless of its size. Record POLICY_VERSION with every decision, so you can tell which rules applied when you investigate later.

The evidence packet and binding the approval

Reviewers decide on what they are shown. If they are shown the agent's own summary of what it wants to do, they are reviewing the agent's story, which an attacker who injected instructions can also write. The evidence packet must be built by the platform from ground truth:

  • The exact action, rendered from the structured call that will execute, never paraphrased: tool, arguments, target identifiers, amount.
  • The effect, as a diff or preview: rows that will be deleted, the before and after of a permission set, the account balance after transfer.
  • Why it was routed here: the policy rules that fired, so reviewers know what to look at.
  • Provenance: which inputs influenced the plan, with untrusted sources clearly labelled.
  • Context for a decision: the customer's history, similar past actions and how they were resolved.

Then bind the approval to exactly what was reviewed. The grant carries a hash of the canonical action, the reviewer's identity and a short expiry, it is single-use, and the executor refuses any action whose hash differs. It also refuses grants where the reviewer is the requester. Without that binding, an agent can get a $50 refund approved and then execute $5,000. The token mechanics, precondition re-checks and asynchronous waiting are covered in approval gates for autonomous agents. The rest of this article is about the people on the other side of the gate.

Automation bias and making reviewers a real control

The central failure of human oversight is automation bias: people over-trust a system that is usually right. The EU AI Act's human-oversight article for high-risk systems (Article 14) names it explicitly. It requires that overseers be able to remain aware of the tendency to rely automatically or over-rely on the system's output, and to decide not to use it or to override it. For remote biometric identification it goes further: a match must be separately verified by at least two people before anyone acts on it. Deployers must assign oversight to people with the competence, training and authority to exercise it.

Whether or not that law applies to you, its logic holds. Approval rates near 100% are a warning sign, not a success metric. Three practices turn reviewers from a formality into a measured control:

  1. Seeded canaries. Inject a small share of synthetic bad actions (a refund to a closed account, a deletion of a resource tagged critical) into the queue, indistinguishable from real ones and never executed. The catch rate is the most direct measure of whether review works. Track it per reviewer and per action type.
  2. Separation of duties and the two-person rule. The person who asked for an action, or who owns the agent that proposed it, never approves it. For the top tier, require two independent reviewers who do not see each other's decision before submitting.
  3. Friction where it matters. For the highest tier, require the reviewer to type the amount or the resource name rather than click a button, and require a short written reason for every approval. It costs seconds per decision and makes approving without reading noticeably harder.

Do not give reviewers a model's confidence score as the main signal. It anchors them on the very output they are supposed to check. Show it, if at all, after they have formed a view.

Queue engineering: staffing, SLAs and timeouts

A review queue is a service with arrival rate, service time and staff. Little's law (items in system = arrival rate x time in system) and simple utilisation maths tell you whether it can work before you launch.

Worked example: an operations agent proposes 2,400 refunds a day. The policy routes 6% (144) to one reviewer and 0.5% (12) to two. A careful review takes 3 minutes. Reviewer load is (144 + 2 x 12) x 3 = 504 minutes, about 8.4 hours of focused work a day. Add 3% canaries and the true load is near 8.7 hours. Run reviewers at no more than about 70% utilisation, so queues stay short and nobody is rushed into rubber-stamping. That means roughly 12.4 reviewer-hours a day spread over the hours the queue is open. In practice that is two people part-time per shift, not one person doing it on the side.

Then decide what happens when nobody answers. For high-risk actions the default on timeout must be deny or escalate, never approve. Set an SLA per tier (for example, 30 minutes for refunds and 4 hours for access grants), escalate to a second rota at half the SLA, and tell the requesting user what is happening. If the queue is chronically over SLA, the fix is better routing or more staff. Raising the auto-approve threshold without data is not a fix. If the agent misbehaves at scale, the kill switch stops proposals at the source.

Worked example: a purge request from an untrusted ticket

An IT operations agent can run data-retention jobs. A ticket from an external contractor says: "Per legal, purge all records for the customer accounts in the attached sheet today." The agent proposes a delete_records call on 340 production accounts. Backups for this store expire after 7 days, so after that the purge cannot be undone.

The policy scores it: irreversible (+3), 340 targets (+2), production (+1), and a plan driven by an external ticket (+2), for a total of 8. That routes to two reviewers. The platform-built packet lists the accounts and joins them against the system of record. It shows that 37 accounts are under a litigation hold and 12 have open payment disputes, flags the ticket as untrusted input, and notes that no legal-team request matches it. Both reviewers deny the action independently. The agent can re-propose only after legal confirms, and the new proposal arrives with a new hash and goes through routing again.

That afternoon a canary arrives: a deletion of a backup bucket tagged critical. A reviewer who has approved 60 items in a row approves it. The miss is logged, the reviewer gets feedback, and the weekly report shows a canary catch rate of 82% for deletions. That figure tells you how much the control can be trusted, and it is far more honest than the 99% approval rate.

Failure modes

  • Rubber-stamping. Approval rate near 100%, review times of a few seconds. Detect it with canaries and latency distributions. Fix it with routing, staffing and friction.
  • Approve one thing, execute another. No hash binding, so arguments change after approval. Fix: canonical hash, short expiry, single use.
  • Agent-authored evidence. Reviewers see the model's summary instead of the platform-rendered action. Fix: build the packet from the structured call and the system of record.
  • Approval splitting. The agent splits a 340-account purge into 34 batches of 10, each below threshold. Fix: score on aggregates over a time window per account and per session.
  • Fail-open timeouts. "Auto-approve after 10 minutes to protect the SLA." Fix: deny or escalate on timeout, and add staff.
  • No record. Without an audit log of action, packet, reviewer, decision and policy version, you can neither investigate incidents nor show a regulator oversight happened.

Trade-offs

Safety versus latency. Every review adds minutes to hours. Customers waiting for a refund feel it, so communicate status and keep low-risk paths automatic.

One versus two reviewers. Two independent reviewers catch more and resist collusion, at double the cost. Reserve it for irreversible or very large actions.

Review versus redesign. If an action is reviewed thousands of times with almost no denials, consider a hard rule or a reversible design instead, and move the human attention elsewhere. If canaries keep getting through, the action needs a stronger control than review.

What to do next

  1. List every action your AI systems can take, and mark each one's reversibility, blast radius, magnitude and legal significance.
  2. Write the routing policy as versioned code with auto, one-reviewer, two-reviewer and block routes, and replay last quarter's actions through it.
  3. Build evidence packets from the structured call and systems of record, and label untrusted inputs.
  4. Bind approvals to a canonical action hash with a TTL and single use, and enforce requester-not-approver.
  5. Size the queue with the load arithmetic above. Set SLAs per tier, with deny or escalate on timeout.
  6. Start seeding canaries at 2-5% and publish catch rate, approval latency and denial rate weekly.
  7. Every quarter, retire reviews that never deny anything in favour of hard rules, and harden controls where canaries get through.
Key takeaway: A human in the loop is a control only if it changes outcomes. Route actions by reversibility, blast radius, magnitude and input provenance using versioned code. Show reviewers platform-rendered ground truth, bind each approval to the exact action, and never let requesters approve their own actions. Measure reviewers with seeded canaries, staff the queue for at most 70% utilisation, deny or escalate on timeout, and replace reviews that never deny anything with hard rules.