Diagrams of agent payments show a tidy line of arrows: the agent builds a cart, the user approves, the merchant charges, everyone gets a receipt. Production is the space between those arrows. A receipt never arrives. Two autonomous purchases race against the same budget. The protocol defines what each message means and who must verify it; it does not tell you what to believe when a message is missing.
This article treats the Agent Payments Protocol (AP2) flow as a distributed state machine you have to build: identifiers, receipts, state, retries and budgets, with code, a worked example and a checklist. It follows AP2 v0.2 as documented at ap2-protocol.org on 2026-10-01; AP2 is young, so pin the version you implement. The roles and the hop-by-hop protocol walk are covered in AP2 architecture. Where this article goes beyond the specification it says so, because several of the most important behaviours are left to implementers.
The flow in one paragraph
Five roles take part: the Shopping Agent, the Trusted Surface (which must be non-agentic), the Credential Provider, the Merchant and the Merchant Payment Processor. In the human-present flow the merchant signs a checkout JWT describing the finalized cart; the agent builds the content of a closed Checkout Mandate and a closed Payment Mandate; the Trusted Surface shows them to the user and obtains the user's signature; the Credential Provider verifies the Payment Mandate and returns a payment credential; the agent gives the credential and the Checkout Mandate to the merchant, which verifies the mandate and starts the payment with its processor. The merchant returns a Checkout Receipt to the agent, and the processor returns a Payment Receipt to the agent, the Credential Provider and the network.
In the human-not-present flow the user signs open mandates earlier, constraints plus the agent's public key in a cnf claim. Later the agent signs the closed mandates itself with that key and presents both the user-signed open mandate and its own closed mandate to the verifiers, disclosing only the constraint elements each verifier needs.
Two hashes, two jobs
AP2 ties messages to purchases with two different hashes. Confusing them gives a flow that works in a demo and misroutes receipts in production.
| Identifier | Hash of | Appears in | Use it to |
|---|---|---|---|
checkout_hash / transaction_id | the merchant-signed checkout_jwt value | closed Checkout Mandate, closed Payment Mandate | prove both mandates cover the same cart; key the purchase in your own store |
reference in a receipt | the closed mandate the receipt answers, computed the way SD-JWT computes sd_hash | Checkout Receipt, Payment Receipt | match a receipt to the exact authorization that produced it |
The distinction matters when there is more than one authorization for one cart. If a payment fails and the agent obtains a new closed Payment Mandate for the same checkout, perhaps with a different instrument, both mandates carry the same transaction_id but produce different receipt references. Key the purchase on transaction_id and each attempt on the mandate hash.
import base64
import hashlib
def b64url_sha256(value: str) -> str:
digest = hashlib.sha256(value.encode("ascii")).digest()
return base64.urlsafe_b64encode(digest).rstrip(b"=").decode("ascii")
class Purchase:
def __init__(self, checkout_jwt: str):
self.checkout_jwt = checkout_jwt
self.transaction_id = b64url_sha256(checkout_jwt) # one per cart
self.attempts = {} # mandate hash -> attempt state
def add_attempt(self, presented_payment_mandate: str) -> str:
# Assumption: the receipt reference is computed over the mandate exactly as presented,
# like SD-JWT's sd_hash. Hash the bytes you sent, never a re-serialized copy, and check
# the algorithm and input against the specification's examples before relying on it.
ref = b64url_sha256(presented_payment_mandate)
self.attempts[ref] = "SUBMITTED"
return refThe code hard-codes SHA-256 for brevity; the specification uses the SD-JWT's declared algorithm, SHA-256 by default. Refuse algorithms you do not support.
What each party has to remember
AP2 tells each role what to verify; the table lists what each role must therefore persist. These are engineering choices, not protocol text.
| Role | Persists | Why |
|---|---|---|
| Shopping Agent | Purchase by transaction_id, attempts by mandate hash, receipts | To retry safely, to report outcomes to the user and to obey the rejection-receipt rule below |
| Trusted Surface | What was displayed and signed, with user authentication evidence | Consent is the evidence a dispute will ask for; it must match the mandate byte for byte |
| Credential Provider | Credentials issued per transaction_id; open-mandate spend and occurrence counters | To avoid issuing twice for one checkout and to evaluate budget and recurrence constraints |
| Merchant | Its signed checkout JWTs and their expiry; Checkout Receipts issued | To confirm a mandate refers to a cart it signed and still intends to honour |
| Merchant Payment Processor | Payments by mandate hash, with network outcomes | To make a resubmission return the original result instead of a second charge |
Receipts are the only terminal events
In AP2 a purchase ends when receipts arrive. Both receipt types carry a status that is either success or error, the issuer iss, an issue time iat and the reference hash. The Payment Receipt also carries a payment_id; on success it may carry psp_confirmation_id and network_confirmation_id, and on error an error code and an error_description. The Checkout Receipt carries an order_id on success and the same error fields on failure.
Treat a receipt as a signed claim to verify. Check the issuer is the party you expected and the reference matches an attempt you made, then move state. Log a receipt that matches nothing: it is a bug or an attack.
def apply_payment_receipt(purchase, receipt, expected_issuer):
claims = verify_jwt(receipt, issuer=expected_issuer) # signature, iss, iat window
ref = claims["reference"]
if ref not in purchase.attempts:
raise UnknownReceipt(ref) # never guess which attempt it meant
if purchase.attempts[ref] in ("SUCCEEDED", "FAILED"):
return purchase.attempts[ref] # duplicate delivery: no-op
if claims["status"] == SUCCESS: # compare against the spec's enum value
purchase.attempts[ref] = "SUCCEEDED"
purchase.payment_id = claims["payment_id"]
purchase.network_ref = claims.get("network_confirmation_id")
else:
purchase.attempts[ref] = "FAILED"
purchase.last_error = (claims.get("error"), claims.get("error_description"))
return purchase.attempts[ref]The two receipts come from different parties and can disagree: an order id with a failed payment, or a payment with no order. Decide in advance what each combination means; a payment without an order is a refund case. The receipt format and its verification are covered in depth in AP2 receipts.
Timeouts, retries and the UNKNOWN state
The specification does not define retries, timeouts or a status query, which leaves UNKNOWN to you: the agent sent the credential and Checkout Mandate and heard nothing, so the payment may have succeeded, failed or never started. Two rules follow from the identifiers.
- Resend, do not re-authorize. Retry the same request carrying the same mandate. A processor that stores payments by mandate hash can return the original outcome instead of charging again. Creating a fresh closed mandate to retry turns one purchase into two authorizations, and both may succeed.
- Resolve before replacing. Only start a new attempt for the same transaction_id after the previous one ended in an error receipt or an out-of-band status check confirmed it failed. Until then the purchase stays in UNKNOWN and the user sees pending, not failed.
async def submit(purchase, attempt_ref, request, deadline_s=60):
delay = 1.0
while True:
try:
return await merchant.complete_checkout(request, timeout=10) # same mandate every time
except (TimeoutError, ConnectionError):
if loop_time() > purchase.started + deadline_s:
purchase.attempts[attempt_ref] = "UNKNOWN"
schedule_reconciliation(purchase.transaction_id) # status check later
return None
await asyncio.sleep(delay)
delay = min(delay * 2, 8.0)This is the exactly-once problem every payment API faces, with more hops; see idempotency in agent payments. Mandates can carry exp, so a retry after expiry should be rejected and the checkout rebuilt.
Autonomous purchases: budgets, recurrence and concurrency
Open Payment Mandates carry constraints that only make sense with memory. The payment.budget constraint has a max and a currency, and the rule is that the requested amount plus the total of amounts from previously closed Payment Mandates must not exceed the maximum, with the approved amount added to the total afterwards. The payment.agent_recurrence constraint has a frequency such as WEEKLY or MONTHLY and an optional max_occurrences; a presentation passes only if it is sufficiently separated in time from the previous one and the count is not exceeded. Other constraints, such as payment.amount_range, payment.allowed_payees and payment.execution_date, are stateless checks on a single mandate.
The specification also says that Shopping Agents must not present any subsequent open Payment or Checkout Mandates without receiving a rejection receipt from the previous one. Read plainly, the agent must serialise its use of an open mandate: one presentation in flight at a time, and a new one only after the previous was rejected. That prevents a well-behaved agent from double-spending. It does not protect you from a compromised or buggy agent, which is exactly the party AP2 assumes may be wrong.
So the verifier that evaluates the stateful constraints must enforce them atomically. The specification does not say which party holds the counters; the Credential Provider is a natural choice because it already verifies the Payment Mandate and controls whether a credential is issued, but that is a design decision. A reserve-then-settle pattern keeps the counter correct when payments fail.
def reserve(db, open_mandate_id, closed_hash, amount_minor, budget_max_minor):
with db.transaction(isolation="serializable"):
row = db.select_for_update("budgets", open_mandate_id) # lock this mandate's counter
if db.exists("reservations", closed_hash):
return "already_reserved" # retry of the same attempt
if row.settled + row.reserved + amount_minor > budget_max_minor:
return "reject" # becomes an error receipt
row.reserved += amount_minor
db.insert("reservations", closed_hash, amount_minor, state="held")
return "ok"
def on_payment_receipt(db, open_mandate_id, closed_hash, succeeded):
with db.transaction(isolation="serializable"):
row = db.select_for_update("budgets", open_mandate_id)
res = db.get("reservations", closed_hash)
if res.state != "held":
return # duplicate receipt
row.reserved -= res.amount
if succeeded:
row.settled += res.amount # the spec's accumulated total
res.state = "settled" if succeeded else "released"Counting reservations is stricter than the specification's wording, which counts approved amounts; it trades headroom for never overspending while an outcome is unknown. Held reservations need a reconciliation job, or a lost receipt quietly shrinks the budget.
What AP2 leaves to the payment rail
AP2 does not move money, and the documents examined here do not define step-up authentication, decline codes, capture, refunds or disputes; those stay with the commerce protocol and the rail. Cardholder challenges remain the domain of 3-D Secure: in a human-not-present purchase nobody is there to answer one, so the attempt should end in an error receipt and a user notification, not a silent wait. Map issuer declines onto Payment Receipt error fields and keep the rail's decline code too. Refunds reference the original payment through rail identifiers, so store payment_id and the confirmation ids; see refund architecture.
Worked example: a weekly grocery mandate
The user approves open mandates on their phone's Trusted Surface: two allowed grocers, one allowed card and a budget of 60 euros, with the agent's public key bound in cnf. The Credential Provider keeps a counter for the open Payment Mandate in minor units: settled 0, reserved 0.
On Tuesday the agent receives a signed checkout from grocer A for 38.20 euros and signs closed mandates with its key. The Credential Provider verifies the open mandate, the disclosed payee and instrument, and reserves 3,820. The checkout completes, the processor's success receipt arrives with a payment id, and the reservation settles: settled 3,820.
On Thursday the agent tries grocer B for 19.90 euros: 3,820 + 1,990 = 5,810, within 6,000, so 1,990 is reserved. The request reaches the merchant and the connection drops. The agent resends the same mandate twice, hears nothing, and marks the attempt UNKNOWN; the hold stays. An hour later reconciliation learns by transaction_id that the issuer declined. The reservation is released, settled stays at 3,820, and only now may the agent start a new attempt. A later 25-euro basket fails the budget check even though each order alone is small, and the user is asked to approve it directly.
Failure modes
| Failure | What goes wrong | Defence |
|---|---|---|
| Retry by re-authorizing | Two closed mandates for one cart, two charges | Resend the same mandate; new attempts only after a confirmed failure |
| Receipts keyed on transaction_id | A late success receipt is applied to the wrong attempt | Key attempts on the mandate hash; store both |
| Unverified receipts | A forged success marks an unpaid order paid | Verify signature and issuer; reject unknown references |
| Budget counted by the agent only | A compromised agent spends past the limit | Atomic counters at a verifier, independent of agent behaviour |
| Lost receipt, held reservation | Budget shrinks silently | Reconciliation job for held reservations and UNKNOWN attempts |
| Re-serialized hashing | Hashes never match across parties | Hash the exact bytes received or sent |
Daily matching of receipts against processor and ledger records catches what the online path misses; see payment reconciliation.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Budget accounting | Count approved amounts only: matches the spec's wording, can overspend while outcomes are unknown | Reserve then settle: never overspends, needs reconciliation for stuck holds |
| Who holds counters | Credential Provider: already verifies the mandate, single choke point | Merchant or processor: closer to the money, but one user's budget spans many merchants |
| UNKNOWN handling | Long wait and status query: fewer duplicates, slower feedback | Fail fast and ask the user: faster, risks duplicate purchases if the first succeeded |
What to do next
- Pin the AP2 version you implement and record it with every stored mandate and receipt.
- Model the purchase explicitly: key it on transaction_id, key attempts on the mandate hash, and include an UNKNOWN state with a reconciliation path.
- Make retries resend the same mandate, and allow a new closed mandate only after a confirmed failure.
- Verify every receipt's signature, issuer and reference before changing state, and store payment_id and confirmation ids for refunds.
- Put budget and recurrence counters behind an atomic reserve-then-settle API at one verifier, and test it with concurrent presentations.
- Decide and document what each combination of Checkout Receipt and Payment Receipt outcomes means for the user.