Diagrams of agent payments show a tidy line of arrows: the agent builds a cart, the user approves, the merchant charges, everyone gets a receipt. Production is the space between those arrows. A receipt never arrives. Two autonomous purchases race against the same budget. The protocol defines what each message means and who must verify it; it does not tell you what to believe when a message is missing.

This article treats the Agent Payments Protocol (AP2) flow as a distributed state machine you have to build: identifiers, receipts, state, retries and budgets, with code, a worked example and a checklist. It follows AP2 v0.2 as documented at ap2-protocol.org on 2026-10-01; AP2 is young, so pin the version you implement. The roles and the hop-by-hop protocol walk are covered in AP2 architecture. Where this article goes beyond the specification it says so, because several of the most important behaviours are left to implementers.

Advertisement

The flow in one paragraph

Five roles take part: the Shopping Agent, the Trusted Surface (which must be non-agentic), the Credential Provider, the Merchant and the Merchant Payment Processor. In the human-present flow the merchant signs a checkout JWT describing the finalized cart; the agent builds the content of a closed Checkout Mandate and a closed Payment Mandate; the Trusted Surface shows them to the user and obtains the user's signature; the Credential Provider verifies the Payment Mandate and returns a payment credential; the agent gives the credential and the Checkout Mandate to the merchant, which verifies the mandate and starts the payment with its processor. The merchant returns a Checkout Receipt to the agent, and the processor returns a Payment Receipt to the agent, the Credential Provider and the network.

In the human-not-present flow the user signs open mandates earlier, constraints plus the agent's public key in a cnf claim. Later the agent signs the closed mandates itself with that key and presents both the user-signed open mandate and its own closed mandate to the verifiers, disclosing only the constraint elements each verifier needs.

The Shopping Agent's view of one AP2 purchase: every arrow is a message that can be lostCARTitems agreedCHECKOUT_SIGNEDmerchant checkout_jwtAUTHORIZEDclosed mandates signedCREDENTIALscoped by CPTS or agent keySUBMITTEDmerchant, MPPUNKNOWNtimeout, no receipttimeoutSUCCEEDEDboth receipts okFAILEDerror receiptsuccesserrorstatus: failedstatus: paidRECONCILEDreceipt ids stored, budget hold settled or releasedKeys: transaction_id = hash of the merchant's checkout_jwt (identifies the purchase);receipt reference = hash of the closed mandate it answers (identifies the authorization).State names are this article's, not AP2 vocabulary. AP2 defines the messages; you own the states.
One purchase from the agent's side. The happy path runs along the top and down the right; UNKNOWN is the state the protocol diagrams leave out, and it is where most engineering effort goes.

Two hashes, two jobs

AP2 ties messages to purchases with two different hashes. Confusing them gives a flow that works in a demo and misroutes receipts in production.

IdentifierHash ofAppears inUse it to
checkout_hash / transaction_idthe merchant-signed checkout_jwt valueclosed Checkout Mandate, closed Payment Mandateprove both mandates cover the same cart; key the purchase in your own store
reference in a receiptthe closed mandate the receipt answers, computed the way SD-JWT computes sd_hashCheckout Receipt, Payment Receiptmatch a receipt to the exact authorization that produced it

The distinction matters when there is more than one authorization for one cart. If a payment fails and the agent obtains a new closed Payment Mandate for the same checkout, perhaps with a different instrument, both mandates carry the same transaction_id but produce different receipt references. Key the purchase on transaction_id and each attempt on the mandate hash.

import base64
import hashlib

def b64url_sha256(value: str) -> str:
    digest = hashlib.sha256(value.encode("ascii")).digest()
    return base64.urlsafe_b64encode(digest).rstrip(b"=").decode("ascii")

class Purchase:
    def __init__(self, checkout_jwt: str):
        self.checkout_jwt = checkout_jwt
        self.transaction_id = b64url_sha256(checkout_jwt)  # one per cart
        self.attempts = {}                                 # mandate hash -> attempt state

    def add_attempt(self, presented_payment_mandate: str) -> str:
        # Assumption: the receipt reference is computed over the mandate exactly as presented,
        # like SD-JWT's sd_hash. Hash the bytes you sent, never a re-serialized copy, and check
        # the algorithm and input against the specification's examples before relying on it.
        ref = b64url_sha256(presented_payment_mandate)
        self.attempts[ref] = "SUBMITTED"
        return ref

The code hard-codes SHA-256 for brevity; the specification uses the SD-JWT's declared algorithm, SHA-256 by default. Refuse algorithms you do not support.

Advertisement

What each party has to remember

AP2 tells each role what to verify; the table lists what each role must therefore persist. These are engineering choices, not protocol text.

RolePersistsWhy
Shopping AgentPurchase by transaction_id, attempts by mandate hash, receiptsTo retry safely, to report outcomes to the user and to obey the rejection-receipt rule below
Trusted SurfaceWhat was displayed and signed, with user authentication evidenceConsent is the evidence a dispute will ask for; it must match the mandate byte for byte
Credential ProviderCredentials issued per transaction_id; open-mandate spend and occurrence countersTo avoid issuing twice for one checkout and to evaluate budget and recurrence constraints
MerchantIts signed checkout JWTs and their expiry; Checkout Receipts issuedTo confirm a mandate refers to a cart it signed and still intends to honour
Merchant Payment ProcessorPayments by mandate hash, with network outcomesTo make a resubmission return the original result instead of a second charge

Receipts are the only terminal events

In AP2 a purchase ends when receipts arrive. Both receipt types carry a status that is either success or error, the issuer iss, an issue time iat and the reference hash. The Payment Receipt also carries a payment_id; on success it may carry psp_confirmation_id and network_confirmation_id, and on error an error code and an error_description. The Checkout Receipt carries an order_id on success and the same error fields on failure.

Treat a receipt as a signed claim to verify. Check the issuer is the party you expected and the reference matches an attempt you made, then move state. Log a receipt that matches nothing: it is a bug or an attack.

def apply_payment_receipt(purchase, receipt, expected_issuer):
    claims = verify_jwt(receipt, issuer=expected_issuer)   # signature, iss, iat window
    ref = claims["reference"]
    if ref not in purchase.attempts:
        raise UnknownReceipt(ref)                           # never guess which attempt it meant
    if purchase.attempts[ref] in ("SUCCEEDED", "FAILED"):
        return purchase.attempts[ref]                       # duplicate delivery: no-op
    if claims["status"] == SUCCESS:                         # compare against the spec's enum value
        purchase.attempts[ref] = "SUCCEEDED"
        purchase.payment_id = claims["payment_id"]
        purchase.network_ref = claims.get("network_confirmation_id")
    else:
        purchase.attempts[ref] = "FAILED"
        purchase.last_error = (claims.get("error"), claims.get("error_description"))
    return purchase.attempts[ref]

The two receipts come from different parties and can disagree: an order id with a failed payment, or a payment with no order. Decide in advance what each combination means; a payment without an order is a refund case. The receipt format and its verification are covered in depth in AP2 receipts.

Timeouts, retries and the UNKNOWN state

The specification does not define retries, timeouts or a status query, which leaves UNKNOWN to you: the agent sent the credential and Checkout Mandate and heard nothing, so the payment may have succeeded, failed or never started. Two rules follow from the identifiers.

  1. Resend, do not re-authorize. Retry the same request carrying the same mandate. A processor that stores payments by mandate hash can return the original outcome instead of charging again. Creating a fresh closed mandate to retry turns one purchase into two authorizations, and both may succeed.
  2. Resolve before replacing. Only start a new attempt for the same transaction_id after the previous one ended in an error receipt or an out-of-band status check confirmed it failed. Until then the purchase stays in UNKNOWN and the user sees pending, not failed.
async def submit(purchase, attempt_ref, request, deadline_s=60):
    delay = 1.0
    while True:
        try:
            return await merchant.complete_checkout(request, timeout=10)  # same mandate every time
        except (TimeoutError, ConnectionError):
            if loop_time() > purchase.started + deadline_s:
                purchase.attempts[attempt_ref] = "UNKNOWN"
                schedule_reconciliation(purchase.transaction_id)          # status check later
                return None
            await asyncio.sleep(delay)
            delay = min(delay * 2, 8.0)

This is the exactly-once problem every payment API faces, with more hops; see idempotency in agent payments. Mandates can carry exp, so a retry after expiry should be rejected and the checkout rebuilt.

Autonomous purchases: budgets, recurrence and concurrency

Open Payment Mandates carry constraints that only make sense with memory. The payment.budget constraint has a max and a currency, and the rule is that the requested amount plus the total of amounts from previously closed Payment Mandates must not exceed the maximum, with the approved amount added to the total afterwards. The payment.agent_recurrence constraint has a frequency such as WEEKLY or MONTHLY and an optional max_occurrences; a presentation passes only if it is sufficiently separated in time from the previous one and the count is not exceeded. Other constraints, such as payment.amount_range, payment.allowed_payees and payment.execution_date, are stateless checks on a single mandate.

The specification also says that Shopping Agents must not present any subsequent open Payment or Checkout Mandates without receiving a rejection receipt from the previous one. Read plainly, the agent must serialise its use of an open mandate: one presentation in flight at a time, and a new one only after the previous was rejected. That prevents a well-behaved agent from double-spending. It does not protect you from a compromised or buggy agent, which is exactly the party AP2 assumes may be wrong.

So the verifier that evaluates the stateful constraints must enforce them atomically. The specification does not say which party holds the counters; the Credential Provider is a natural choice because it already verifies the Payment Mandate and controls whether a credential is issued, but that is a design decision. A reserve-then-settle pattern keeps the counter correct when payments fail.

def reserve(db, open_mandate_id, closed_hash, amount_minor, budget_max_minor):
    with db.transaction(isolation="serializable"):
        row = db.select_for_update("budgets", open_mandate_id)       # lock this mandate's counter
        if db.exists("reservations", closed_hash):
            return "already_reserved"                                # retry of the same attempt
        if row.settled + row.reserved + amount_minor > budget_max_minor:
            return "reject"                                          # becomes an error receipt
        row.reserved += amount_minor
        db.insert("reservations", closed_hash, amount_minor, state="held")
    return "ok"

def on_payment_receipt(db, open_mandate_id, closed_hash, succeeded):
    with db.transaction(isolation="serializable"):
        row = db.select_for_update("budgets", open_mandate_id)
        res = db.get("reservations", closed_hash)
        if res.state != "held":
            return                                                   # duplicate receipt
        row.reserved -= res.amount
        if succeeded:
            row.settled += res.amount                                # the spec's accumulated total
        res.state = "settled" if succeeded else "released"

Counting reservations is stricter than the specification's wording, which counts approved amounts; it trades headroom for never overspending while an outcome is unknown. Held reservations need a reconciliation job, or a lost receipt quietly shrinks the budget.

What AP2 leaves to the payment rail

AP2 does not move money, and the documents examined here do not define step-up authentication, decline codes, capture, refunds or disputes; those stay with the commerce protocol and the rail. Cardholder challenges remain the domain of 3-D Secure: in a human-not-present purchase nobody is there to answer one, so the attempt should end in an error receipt and a user notification, not a silent wait. Map issuer declines onto Payment Receipt error fields and keep the rail's decline code too. Refunds reference the original payment through rail identifiers, so store payment_id and the confirmation ids; see refund architecture.

Worked example: a weekly grocery mandate

The user approves open mandates on their phone's Trusted Surface: two allowed grocers, one allowed card and a budget of 60 euros, with the agent's public key bound in cnf. The Credential Provider keeps a counter for the open Payment Mandate in minor units: settled 0, reserved 0.

On Tuesday the agent receives a signed checkout from grocer A for 38.20 euros and signs closed mandates with its key. The Credential Provider verifies the open mandate, the disclosed payee and instrument, and reserves 3,820. The checkout completes, the processor's success receipt arrives with a payment id, and the reservation settles: settled 3,820.

On Thursday the agent tries grocer B for 19.90 euros: 3,820 + 1,990 = 5,810, within 6,000, so 1,990 is reserved. The request reaches the merchant and the connection drops. The agent resends the same mandate twice, hears nothing, and marks the attempt UNKNOWN; the hold stays. An hour later reconciliation learns by transaction_id that the issuer declined. The reservation is released, settled stays at 3,820, and only now may the agent start a new attempt. A later 25-euro basket fails the budget check even though each order alone is small, and the user is asked to approve it directly.

Failure modes

FailureWhat goes wrongDefence
Retry by re-authorizingTwo closed mandates for one cart, two chargesResend the same mandate; new attempts only after a confirmed failure
Receipts keyed on transaction_idA late success receipt is applied to the wrong attemptKey attempts on the mandate hash; store both
Unverified receiptsA forged success marks an unpaid order paidVerify signature and issuer; reject unknown references
Budget counted by the agent onlyA compromised agent spends past the limitAtomic counters at a verifier, independent of agent behaviour
Lost receipt, held reservationBudget shrinks silentlyReconciliation job for held reservations and UNKNOWN attempts
Re-serialized hashingHashes never match across partiesHash the exact bytes received or sent

Daily matching of receipts against processor and ledger records catches what the online path misses; see payment reconciliation.

Trade-offs

DecisionOption AOption B
Budget accountingCount approved amounts only: matches the spec's wording, can overspend while outcomes are unknownReserve then settle: never overspends, needs reconciliation for stuck holds
Who holds countersCredential Provider: already verifies the mandate, single choke pointMerchant or processor: closer to the money, but one user's budget spans many merchants
UNKNOWN handlingLong wait and status query: fewer duplicates, slower feedbackFail fast and ask the user: faster, risks duplicate purchases if the first succeeded

What to do next

  1. Pin the AP2 version you implement and record it with every stored mandate and receipt.
  2. Model the purchase explicitly: key it on transaction_id, key attempts on the mandate hash, and include an UNKNOWN state with a reconciliation path.
  3. Make retries resend the same mandate, and allow a new closed mandate only after a confirmed failure.
  4. Verify every receipt's signature, issuer and reference before changing state, and store payment_id and confirmation ids for refunds.
  5. Put budget and recurrence counters behind an atomic reserve-then-settle API at one verifier, and test it with concurrent presentations.
  6. Decide and document what each combination of Checkout Receipt and Payment Receipt outcomes means for the user.
Key takeaway: AP2 defines the messages of an agent purchase and who must verify them; the reliability of the flow is yours to build. Key the purchase on transaction_id and each authorization on its mandate hash, treat verified receipts as the only terminal events, give UNKNOWN a real state with reconciliation behind it, retry by resending rather than re-authorizing, and enforce budget and recurrence constraints atomically at a verifier rather than trusting the agent to count.