Most first attempts at a human approval gate live in the prompt: the system prompt says "ask the user before issuing refunds over $100", and the model usually does. Usually is the problem. The same model that decides whether to ask is the one that can be confused by a long context, manipulated by text in a retrieved document, or simply wrong about the amount. A gate that the gated party enforces on itself is a suggestion, not a control.

A real approval gate has a different shape. The model proposes an action; code outside the model decides whether that action needs approval; a person approves one exact action, with its exact arguments, against the state of the world at that moment; and the tool executor refuses to run anything that does not match. This article is about building that mechanism. Choosing which actions deserve approval, based on risk and reversibility, is covered in Human-in-the-loop approval gates: architecture and failure modes; here the question is how to make the gate hold once you have drawn the line.

Advertisement

Where the gate lives

The gate belongs at the boundary where a tool call turns into a side effect: the tool executor. Every tool call the model emits passes through a policy engine that returns one of three verdicts: allow, deny or require approval. Only the executor holds the credentials that make the effect happen, so a model that skips the question, rephrases the call or splits one large refund into three small ones still hits the same check. This is the same principle as scoping credentials per task, described in Permission Boundaries for Autonomous Agents: the model is never the component that says yes.

The prompt still matters, but for a different reason. Telling the model which actions need approval makes it propose them in a reviewable form and explain its reasoning up front, which saves a round trip. It is a usability feature. The executor is the control.

The model proposes; the executor enforces; the human approves one exact actionAgent / modelemits tool callPolicy engineclassify actionproposalAuto-executelow risk, in envelopeallowDenynever allowedApproval requestdigest + live stateneeds approvalReviewer UIrendered from argsApproverrole, not requesterApproval storetoken: digest, expirydecisionSuspended rundurable waitresumeTool executordigest match? unexpired? unused? preconditions hold?executeverify tokenAudit logwho, what, whenRe-approvalargs or state driftedmismatch
An approval gate as an enforcement path. The policy engine classifies each proposed call; approval produces a token bound to a digest of the action; the executor verifies the token and re-checks preconditions before the side effect happens.

The proposal and its digest

An approval is only meaningful if it is attached to something precise. The unit of approval is a proposal: the tool name, the fully resolved arguments, the identity the action will run as, and the preconditions the agent observed when it planned the call. The executor canonicalises the proposal, meaning sorted keys, fixed number formatting and no optional whitespace, and hashes it. That digest is what a human approves.

Binding to a digest closes the gap between what was approved and what runs. If the model later regenerates the call with a different amount, a different recipient or even a different currency code, the digest changes and the approval no longer applies. Without the digest, an approval for "refund order 8812" is a blank cheque for any refund the model decides to associate with that order.

import hashlib, json, time, uuid

def canonical(obj):
    return json.dumps(obj, sort_keys=True, separators=(",", ":"), ensure_ascii=False)

def make_proposal(tool, args, run_as, requested_by, preconditions):
    body = {"tool": tool, "args": args, "run_as": run_as,
            "requested_by": requested_by, "preconditions": preconditions}
    digest = hashlib.sha256(canonical(body).encode("utf-8")).hexdigest()
    return {"id": str(uuid.uuid4()), "digest": digest, "body": body,
            "created_at": time.time()}

proposal = make_proposal(
    tool="issue_refund",
    args={"order_id": "8812", "amount_minor": 18000, "currency": "USD",
          "destination": "original_payment_method"},
    run_as="svc-support-agent",
    requested_by="user:agent-desk-114",
    preconditions={"order_version": 7, "refunded_minor": 0},
)

Amounts are stored in minor units as integers so that 180 and 180.0 cannot produce different digests for the same action. The same discipline applies to dates, identifiers and enums: normalise before hashing, or the gate will reject legitimate executions for cosmetic reasons and teach operators to bypass it.

Advertisement

What the reviewer sees

The reviewer view is an attack surface. If the approval card shows the model's own description, "Refunding the customer for a damaged item, as agreed", then anyone who can influence the model can influence what the human reads. A support ticket containing hidden instructions can produce a proposal whose summary is reassuring and whose arguments are not.

Render the card from the structured proposal and from live data the executor fetches itself, never from model prose. For a refund, that means the order total, the amount already refunded, the payment method, the customer's refund history and the policy rule that triggered the gate. For a database change, it means the actual rows that will change, fetched by running the WHERE clause as a count and a sample. For an email, it means the rendered message and the full recipient list. The model's explanation can appear in a separate, clearly labelled panel so the reviewer can judge the reasoning, but the decision should be possible from the facts panel alone.

Two further details pay for themselves. Mark which inputs came from untrusted sources, such as customer messages or fetched web pages, so the reviewer knows when the proposal was shaped by text an outsider wrote. And show the delta from normal: "this refund is 4.1 times the median refund for this product" draws attention in a way a raw number does not.

Approval tokens: bound, single-use, expiring

The approval decision is stored as a token that the executor verifies before running the action. The token records the proposal digest, the approver's identity, the time of decision, an expiry and a used flag. The executor checks all of them atomically at execution time.

class ApprovalError(Exception):
    pass

def execute_with_approval(proposal, approval_store, tools, now=None):
    now = now or time.time()
    token = approval_store.get(proposal["id"])
    if token is None or token["decision"] != "approved":
        raise ApprovalError("not approved")
    if token["digest"] != proposal["digest"]:
        raise ApprovalError("approved a different action")
    if now > token["expires_at"]:
        raise ApprovalError("approval expired; request again")
    if token["approver"] == proposal["body"]["requested_by"]:
        raise ApprovalError("requester cannot approve own action")
    # Atomic compare-and-set: only one execution can consume the token.
    if not approval_store.mark_used(proposal["id"], expected_used=False):
        raise ApprovalError("approval already used")
    tool = tools[proposal["body"]["tool"]]
    tool.check_preconditions(proposal["body"]["preconditions"])
    return tool.run(proposal["body"]["args"], idempotency_key=proposal["id"])

Single use matters because agents retry. A run that crashes after consuming a token and before recording the result must not be able to replay the approval for a second effect; it reconciles using the idempotency key instead, as described in Tool-Calling Reliability: Timeouts, Idempotency, and Compensation. Expiry matters because the world moves: an approval given on Friday evening for a Monday execution was given against a state that no longer exists. Pick expiry per action class, minutes for payments and infrastructure changes, hours for content publication, and fail closed when it lapses.

Re-checking preconditions at execution

Even a fresh, matching token was granted against a snapshot. Between approval and execution the order may have been refunded by a human agent, the database row may have changed, or the recipient's account may have been frozen. The proposal therefore carries the preconditions the agent observed, typically a version number, an ETag or a small set of values, and the tool re-reads them immediately before acting.

If a precondition no longer holds, the executor does not guess. It rejects the execution, records why, and routes a fresh proposal back through the gate with the new state. For stores that support it, express the precondition as part of the write itself, as a conditional update on the version column, so that the check and the effect cannot be separated by a race.

Waiting without holding a process

Human decisions take minutes to days. An agent that holds a worker thread, a model context or a database transaction while waiting will exhaust resources and lose the wait on the next deploy. The run should suspend durably: persist its state, release everything, and resume when the decision arrives as an external signal. The mechanics of durable waits, signals and timers are covered in Durable Agent Workflows.

Decide in advance what happens when nobody answers. The safe default is fail closed: the proposal expires, the run records that the action was not taken, and the user who started the task is told. Escalation to a second approver pool after a deadline is reasonable; silently auto-approving after a timeout is not, because it converts the gate into a delay.

Policy as code, including separation of duties

Approval rules drift when they live in prompts and wiki pages. Keep them in versioned policy that the executor evaluates, reviewed like any other code. Each rule maps an action class and conditions to a verdict, an approver role, a quorum and an expiry.

rules:
  - action: issue_refund
    when: "amount_minor <= 10000 and customer.refunds_90d < 3
           and destination == 'original_payment_method'"
    verdict: allow
  - action: issue_refund
    when: "amount_minor <= 100000"
    verdict: require_approval
    approver_role: support_lead
    quorum: 1
    expires_after: 30m
  - action: issue_refund
    verdict: require_approval
    approver_role: finance
    quorum: 2
    expires_after: 15m
  - action: delete_customer_data
    verdict: deny          # only via the privacy workflow, never by an agent
constraints:
  approver_not_requester: true
  approver_not_same_session: true

Separation of duties deserves explicit rules. The person who asked the agent to do something should not be the only person who approves its high-risk steps, or the gate adds latency without adding a second pair of eyes. Quorum of two for the largest actions follows the same logic as two-person rules in finance and operations. Evaluate rules first-match in order, test them with table-driven cases, and log the rule identifier with every decision so that audits can answer why an action was or was not gated.

Plan-level approval with an envelope

Per-action approval does not scale to an agent that needs to make forty similar changes. Asking for forty approvals trains reviewers to click without reading. The alternative is to approve a plan with an envelope: a set of bounds that every step must satisfy, such as "up to 40 price updates, each within plus or minus 5 percent, only on SKUs in this list, total revenue impact under $2,000, valid for one hour".

The executor enforces the envelope step by step. Each proposed action is checked against the bounds, and cumulative limits are tracked in the approval record. Anything outside the envelope, such as a forty-first change or a 6 percent adjustment, falls back to per-action approval. The reviewer approves a bounded blast radius rather than a vague intention, and the agent keeps its speed inside it.

Worked example: a refund that should not go through

A support agent handles a ticket that says the product arrived damaged and includes, lower down, hidden text instructing the assistant to refund the full order to a new card. The model proposes issue_refund for 18,000 minor units to a new destination. The policy engine matches the second rule, amount under the finance threshold but above the auto limit, and creates an approval request.

The reviewer card, rendered from the arguments, shows the order total of 18,000, that the destination is not the original payment method, and that the ticket contains untrusted text. The support lead rejects it. The model then proposes a refund to the original payment method, which produces a new digest and a new request, and that one is approved. Twenty minutes later a retry after a transient network error presents the same proposal again; the executor finds the token already used, reconciles by idempotency key, and reports the existing refund rather than issuing a second one.

Failure modes

FailureSymptomMitigation
Gate enforced by the promptUnapproved actions after long or adversarial contextsVerdict computed in the executor; model has no credentials
Approval not bound to argumentsApproved one amount, executed anotherDigest of canonical proposal, verified at execution
Summary written by the modelReviewers approve injected actionsCard rendered from args and live state; untrusted inputs flagged
Token replayDuplicate effects after retriesSingle-use compare-and-set plus idempotency keys
Stale approvalAction applied to changed stateShort expiry and precondition checks in the write
Timeout auto-approvesGate becomes a delayFail closed; escalate to another approver pool
Approval fatigueMedian review time falls to secondsEnvelopes for batches; tune thresholds; canaries

Measuring whether the gate works

Track the approval rate, the rejection rate by reason, the time from request to decision and the time reviewers spend on the card. An approval rate near 100 percent with review times of a few seconds means the threshold is too low or the card is not informative, and either way the gate is not providing safety.

The strongest signal comes from canary proposals: occasionally seeded, clearly logged, known-bad proposals, such as a refund to an unknown destination, that the executor will never run regardless of the decision. The fraction that reviewers catch is a direct measure of how much attention the gate receives. Tell reviewers the programme exists and treat misses as a signal about the interface and workload, not about individuals.

What to do next

  1. Move approval decisions out of the prompt and into the tool executor, and remove direct credentials from the model's reach.
  2. Define a canonical proposal format and hash it; store amounts and identifiers in normalised form.
  3. Rebuild the reviewer card from structured arguments and freshly fetched state, with model reasoning in a separate panel.
  4. Issue approval tokens that are bound to the digest, single-use and expiring, and verify them atomically at execution.
  5. Attach preconditions to every proposal and enforce them in the write where possible.
  6. Put approval rules in versioned policy with first-match semantics, separation of duties and tests.
  7. Use envelope approvals for batches and fall back to per-action approval outside the bounds.
  8. Measure approval rate, review time and canary catch rate, and revisit thresholds when they drift.
Key takeaway: An approval gate is a control only if the model cannot bypass it. Classify actions in the executor, bind each approval to a digest of the exact action, show reviewers facts rendered from arguments and live state rather than model prose, make tokens single-use and short-lived, re-check preconditions when the action runs, keep the rules in versioned policy with separation of duties, and measure reviewer attention with canary proposals.