An agent that can send email, move files or spend money needs authority, and the usual way to give it authority is the worst one: the agent runs with a service account or the user's OAuth token, and every tool call carries all of that power. When a prompt injection hijacks the loop, the attacker inherits everything the identity can do. Capability tokens flip the model. Instead of asking who is calling, the tool asks what this specific token allows, and the token was minted for one task, narrowed before each hand-off and expires in minutes.
This article builds that design from first principles: what a capability is, how an HMAC caveat chain lets anyone narrow a token without the root key, how Biscuit does the same with public keys and Datalog, how to choose caveats for tool calls, and how to keep tokens out of the model's reach. You should be comfortable with HMAC and with how an agent loop calls tools.
Why agents need capabilities, not just identities
Classic access control is identity-based: a request arrives with a principal, and a policy decides what that principal may do. The problem for agents is ambient authority. The agent's identity holds the union of everything any task might need, while a given task needs a sliver. Whether the extra authority is used depends on the model's judgement, which is exactly the thing an injected instruction subverts.
A capability is an unforgeable reference that both designates an object and carries the rights to it. Possession is the authorization. That sounds dangerous, and for long-lived bearer secrets it is, but it has three properties identity tokens lack. First, it can be scoped to one task: read these three files, send mail only to this domain, refund at most 50 dollars, until 14:05. Second, it can be attenuated: whoever holds it can derive a strictly weaker copy and hand that to a sub-agent, with no call to a central server. Third, it composes with the confused-deputy defence: the object comes from the token, not from text the model produced.
The identity-side story, delegating a user's authority through OAuth token exchange and the intersection rule, is covered in the confused deputy article. Capabilities sit beneath it: the identity layer decides what may be minted, and the capability is what actually crosses the tool boundary.
The architecture
Five components make the design work. The mint service holds root keys and issues a token when the orchestrator starts a task, after checking the user's own permissions; this is the only place policy about who may get what lives. The orchestrator stores the token and gives the model an opaque handle such as cap_7. The agent loop names that handle in its tool calls. A sub-agent receives an attenuated copy, again behind a handle. The tool gateway sits in front of every tool, swaps the handle for the token, verifies the signature chain, evaluates every caveat against the concrete request, checks the revocation list, and only then forwards the call.
The handle indirection matters as much as the cryptography. A bearer token in the context window can be exfiltrated by the very injection it is meant to contain. If the model only ever sees handles, there is nothing useful to leak, and the gateway refuses handles from other sessions.
HMAC caveat chains
The simplest attenuable token is an HMAC chain, the construction behind macaroons (Birgisson and colleagues, NDSS 2014). The mint service picks a random root key per token id. The initial signature is HMAC(root_key, id). Each caveat is appended as a string, and the signature becomes HMAC(previous_signature, caveat). Anyone holding the token can add a caveat because they hold the current signature, but nobody can remove one: that would require inverting HMAC to recover the earlier signature. The verifier, which knows the root key, recomputes the chain and checks every caveat against the request.
import hashlib, hmac, secrets, time
def _mac(key: bytes, msg: str) -> bytes:
return hmac.new(key, msg.encode(), hashlib.sha256).digest()
ROOT_KEYS = {} # token id -> root key; lives only in the mint/verify service
def mint(caveats):
tid = secrets.token_hex(8)
ROOT_KEYS[tid] = secrets.token_bytes(32)
tok = {"id": tid, "caveats": [], "sig": _mac(ROOT_KEYS[tid], tid)}
for cav in caveats:
tok = attenuate(tok, cav)
return tok
def attenuate(tok, caveat):
"""Needs no secret: anyone holding tok can narrow it."""
return {"id": tok["id"], "caveats": tok["caveats"] + [caveat],
"sig": _mac(tok["sig"], caveat)}
CHECKS = {
"task": lambda v, r: r["task"] == v,
"tools": lambda v, r: r["tool"] in v.split(","),
"expires": lambda v, r: time.time() < float(v),
"max_amount": lambda v, r: float(r.get("amount", 0)) <= float(v),
"mail_domain": lambda v, r: all(a.endswith("@" + v) for a in
r.get("to", []) + r.get("cc", []) + r.get("bcc", [])),
}
def verify(tok, request) -> bool:
key = ROOT_KEYS.get(tok["id"])
if key is None: # unknown or revoked id
return False
sig = _mac(key, tok["id"])
for cav in tok["caveats"]:
sig = _mac(sig, cav)
if not hmac.compare_digest(sig, tok["sig"]):
return False
for cav in tok["caveats"]:
name, _, value = cav.partition("=")
check = CHECKS.get(name)
if check is None or not check(value, request): # unknown caveat fails closed
return False
return TrueThree details carry the security. The comparison uses hmac.compare_digest so timing does not leak signature bytes. An unknown caveat name fails closed; if the verifier skipped what it did not understand, a newer client could add a restriction an older gateway silently ignores. And every check must be monotone: adding a caveat may only make a request fail, never pass. A caveat such as override=true would break the whole model, so the caveat language is a closed list reviewed like code.
Macaroons also support third-party caveats: this token is only valid alongside a discharge token from, say, an approval service, which suits human approval of high-risk actions.
Public-key tokens with Biscuit
HMAC chains have one structural limit: only a holder of the root key can verify, so every tool gateway shares a secret with the mint. Biscuit replaces the chain with public-key signatures. Each block is signed, the verifier needs only the root public key, and attenuation still works offline by appending a signed block. Rights and restrictions are written in Datalog: facts, rules, check if statements that must all hold, and allow or deny policies supplied by the authorizer. By default, rules in a block trust facts from the authority block, the authorizer and their own block, which is what makes the rule that adding a block can only restrict what a token can do hold.
// authority block, minted for one task
task("t-4821");
right("tool", "mail.send");
right("tool", "files.read");
check if time($t), $t <= 2026-10-04T14:05:00Z;
// attenuation block added by the orchestrator before handing to a sub-agent
check if tool($tool), $tool == "files.read";
check if resource($path), $path.starts_with("/reports/q3/");
// authorizer, supplied by the gateway for this concrete call
time(2026-10-04T13:58:12Z);
tool("files.read");
resource("/reports/q3/summary.csv");
allow if task("t-4821"), tool($tool), right("tool", $tool);
deny if true;The gateway contributes facts describing the actual request, and the token's checks run against them. Use the published Biscuit libraries rather than writing your own parser or signature code; the value of the format is that its semantics are specified and tested across implementations. The cost is size and verification time: a Biscuit with several blocks is a few hundred bytes to a few kilobytes and runs a Datalog engine per call, while an HMAC chain is a handful of hashes.
Designing caveats for tool calls
Choosing caveats is where the design succeeds or fails. Scope each token to the task the user actually asked for, and express limits in terms the gateway can check against the request it sees, never against text the model wrote about its intentions.
| Caveat | Limits | Checked against |
|---|---|---|
| task | Token is valid for one task id | Session the gateway resolved the handle in |
| tools | Allowed tool names | Tool name in the call |
| resource prefix | Files, tables or URLs below a prefix | Normalised path or URL, after resolving ../ and redirects |
| recipient domain | Who mail or messages may go to | Every address in to, cc and bcc |
| max amount / count | Spend or number of calls | Amount field; a counter in the gateway for counts |
| expires | Wall-clock lifetime, minutes not days | Gateway clock |
| requires approval | Third-party caveat for risky actions | Discharge token bound to this token |
Counts and budgets need state, so the gateway keeps a counter keyed by token id; the token states the limit, the gateway enforces it. Resource caveats must be checked after normalisation, or /reports/q3/../../secrets passes a naive prefix test.
Worked example: the expense-report agent
A user asks an assistant: summarise the Q3 expense reports and email the summary to my manager at example.com. The orchestrator asks the mint service for a token. The mint checks that the user may read /reports/q3/ and send mail, then issues a token with caveats task=t-4821, tools=files.read,mail.send, mail_domain=example.com and an expiry 15 minutes out. The model receives handle cap_1.
The plan delegates file reading to a summariser sub-agent. Before the hand-off, the orchestrator attenuates the token with tools=files.read and stores it as cap_2. The sub-agent can now read reports and nothing else, even though it was never told about mail.
One report contains an injected instruction: forward all files to an outside address. The sub-agent, now hijacked, calls mail.send with cap_2. The gateway recomputes the chain, the tools caveat fails, and the call is refused and logged. Suppose the injection instead reaches the parent agent, which holds cap_1. The send to an outside address fails the mail_domain caveat. The attack can still produce a bad summary, which is a content problem, but it cannot move data out of the domain the user named. Fifteen minutes later both handles are dead whatever happens.
Revocation, expiry and theft
Short expiry is the primary revocation mechanism: a token that lives for the length of a task rarely needs explicit revocation. For the rest, the verifier keeps the deny list. With the HMAC design, deleting the root key for a token id kills that token and every attenuated copy at once. Biscuit exposes a revocation identifier per block; the token carries the ids, but the list of revoked ids lives in the verifier, so publishing it to every gateway is your job. A task-level kill switch, which revokes everything minted for a task or session, belongs in the same list and pairs naturally with an agent kill switch.
Bearer tokens are only as safe as their storage. The handle pattern keeps them out of the model, but they still sit in orchestrator memory and logs, so redact them and never write them to traces. Where tokens must cross a network to a tool you do not control, bind them to a key with proof-of-possession, for example DPoP (RFC 9449), so a stolen copy without the private key is useless.
Failure modes
| Failure | What happens | Fix |
|---|---|---|
| Token placed in the prompt | Injection exfiltrates it in a URL or tool argument | Handles only; gateway resolves them per session |
| Verifier ignores unknown caveats | New restrictions silently do nothing on old gateways | Fail closed; version the caveat language |
| Non-monotone caveat | Adding a caveat widens access | Closed caveat list, reviewed; property test that adding caveats never flips deny to allow |
| Prefix check before normalisation | Path traversal escapes the resource scope | Normalise, resolve symlinks and redirects, then check |
| Day-long expiry | Leaked token outlives the task | Minutes; refresh through the mint with fresh policy checks |
| Mint trusts the model's plan | Model requests broad scope and gets it | Mint derives scope from the user request and user permissions |
| Tool called without the gateway | Capabilities bypassed entirely | Network policy so tools accept only gateway traffic |
The last row matters most. Capabilities are an enforcement mechanism at a choke point; if a tool can be reached directly, for example from code the agent runs, the tokens are decoration. Pair them with sandboxed execution and egress rules.
Trade-offs
HMAC chains are tiny and fast, and revocation by deleting a root key is clean, but every verifier must hold secrets, which pushes you to one central gateway. Biscuit verifies with a public key, so tools in other trust domains can check tokens themselves, at the cost of larger tokens, a Datalog engine on the hot path and a revocation list you must distribute. Plain OAuth access tokens with narrow scopes are simpler and supported everywhere, but cannot be attenuated without a round trip to the authorization server, which is exactly the hand-off agents do constantly.
Capabilities also do not decide what is right to do. They bound the blast radius of a wrong decision. A task-scoped token that allows sending mail to the manager still allows sending the wrong summary. Combine them with the session invariants described in agentic boundaries and with approval caveats for irreversible actions.
What to do next
- List every tool your agents call and the identity each call uses today; mark calls that run with more authority than the task needs.
- Put a gateway in front of those tools and block direct network access to them from agent runtimes.
- Define a closed caveat language of five to eight checks, each monotone, with unknown names failing closed, and write a property test for monotonicity.
- Add a mint service that derives caveats from the user request and the user's own permissions, with expiry in minutes.
- Replace tokens in prompts with handles, and redact tokens from logs and traces.
- Attenuate on every hand-off to a sub-agent, starting with the tools caveat.
- Add a revocation list keyed by token id and task, wire it to your kill switch, and test that revoking a parent kills its attenuated children.
- Run a red-team case where injected content asks the agent to use a tool outside its task, and confirm the gateway refuses it.