Most payment fraud signals were designed for a person at a checkout page: a device fingerprint, a browser session, someone typing a card number. When a software agent shops on someone's behalf, those signals describe the agent's servers rather than the person. The Agent Payments Protocol (AP2) replaces part of what is lost with something stronger: signed, verifiable statements about what the user approved, within which limits, and who signed the final transaction. It also leaves one free-form field for risk observations that cannot be proven cryptographically.
This article explains which signals an AP2 transaction carries, how to grade them by trustworthiness, how to extract them without leaking future information, and what to do with each class. Protocol facts follow the AP2 V2 documentation (release 0.2.0, dated 2026-04-28) in the google-agentic-commerce/AP2 repository, read on 2026-10-01. V2 replaced the intent and cart mandates of the first release with open and closed SD-JWT mandates, so older material, including some on this site, describes different objects. How scores are computed and served inside the authorization path is covered in fraud and risk scoring for agent payments; this page is about what goes in.
What makes something a fraud signal
A signal is an observation plus two pieces of metadata that are easy to drop: who asserted it, and when it became knowable. The first tells you whether the source could lie; the second, whether you could have used it at decision time. A feature table that keeps only the value answers neither.
It helps to grade every signal on a four-level scale. Proven signals come from verifying a signature or evaluating a constraint yourself; nobody upstream can forge the result. Attested signals are signed by a party, so you know who said it, but that party could be mistaken or compromised. Observed signals are measured from your own records, such as how often a key has been seen. Claimed signals are unsigned assertions, and in AP2 the most common source of claims is the shopping agent itself.
That is the protocol's own stance: its security considerations assume preventing prompt injection is infeasible and include all LLMs and agents in the threat model as potential attackers. Anything the shopping agent writes about itself or its user is a claim. What makes AP2 unusual is how much of a transaction is proven instead.
The objects that carry signals
AP2 V2 names a small set of roles. The Trusted Surface shows mandate content to the user and collects their approval and signature. The Shopping Agent assembles purchases. The Merchant signs a checkout_jwt describing the checkout. The Credential Provider verifies the Payment Mandate and releases a payment credential, and a Network and the Merchant Payment Processor may also verify it. Each verifier returns a signed Mandate Receipt.
Mandates are SD-JWTs, which lets a holder reveal only some fields. A closed Payment Mandate uses vct value mandate.payment.1 and describes one payment: transaction_id (a hash of the checkout_jwt), payee, an optional pisp, payment_amount (an ISO 4217 currency and an integer amount in minor units), payment_instrument, optional execution_date, optional risk_data, and iat/exp. An open Payment Mandate uses mandate.payment.open.1, is not yet bound to a transaction, and instead carries constraints and the agent's public key in a cnf claim.
In the direct, human-present mode the user signs the closed mandates. In the autonomous, human-not-present mode the user signs open mandates and the agent later signs closed ones with its own key, binding them to the open mandate with sd_hash.
Class 1: verdicts are gates, not features
The strongest signals are verification outcomes, and the right thing to do with most of them is not to score them. A Mandate Receipt carries iss, a result of success or error, a reference hash of the received mandate, and on error an error code with an optional description. The spec defines four codes:
| Code | Meaning in the spec | Correct response |
|---|---|---|
invalid_credential | The mandate fails verification; terminal | Decline. Never let a model outweigh it |
invalid_mandate | The mandate fails to approve the requested action; terminal | Decline. Closest fit for a constraint evaluating false |
unresolved_constraint | Unknown constraint, or conformance cannot be verified | Fall back to direct approval or a non-agentic flow |
mandates_not_supported | Verifier does not support mandates for this action | Fall back to a non-agentic flow |
Feeding a terminal verdict into a model as a weighted feature invites it to learn that a bad signature is fine when everything else looks normal. Keep verdicts as gates. Their rates are excellent fleet-level signals: a spike of invalid_credential from one agent provider's keys says something changed before individual transactions look odd.
Constraint outcomes are the other half of the proven layer. The V2 docs define ten types: checkout.allowed_merchants and checkout.line_items on open Checkout Mandates, and payment.agent_recurrence, payment.allowed_payees, payment.allowed_payment_instruments, payment.allowed_pisps, payment.amount_range, payment.budget, payment.reference and payment.execution_date on open Payment Mandates. Unknown constraints must be treated as failing. A constraint that fails is a verdict; a constraint that passes narrowly is a feature, which is the next class.
Class 2: structural signals derived from the chain
A verified presentation's structure still says a lot about risk, and because it is computed from verified fields it is proven. The most important structural signal is the mode, and it must be derived: a closed mandate signed by a user credential is direct, and one signed by an agent key backed by a user-signed open mandate is autonomous. Never read the mode from a field the agent supplied.
- Headroom: the amount divided by the
payment.amount_rangemaximum. Purchases clustering just under the ceiling are a classic probing pattern. - Budget use: spent plus requested, divided by the
payment.budgetmaximum. The spec requires verifiers to accumulate approved amounts, so you hold the running total. - Mandate breadth: does the open mandate restrict payees and merchants, or only amounts? Broad is not fraud, but widens the worst case.
- Lifetime:
expminusiaton the open mandate. The spec recommends the smallest expiry that lets the agent finish its task; a year-long autonomous mandate deserves more scrutiny.
One unit trap matters here. The closed mandate's payment_amount is documented in integer minor units, while the constraint examples in the docs write limits as decimals such as 100.50. Normalise both to minor units with the currency's exponent before computing any ratio, or a 120-dollar purchase looks like it used a hundredth of its limit.
Class 3: behavioural signals across presentations
Each presentation can be valid while the sequence is not. Agents must not present a subsequent open mandate without first receiving a rejection receipt for the previous one, and Credential Providers, Networks and processors may reject overlapping mandates or invalidate tokens already issued. So count presentations per open mandate over short windows: a second presentation while the first is still unresolved is a direct protocol violation, not a statistical oddity.
- Overlap count per open mandate reference in the last few minutes.
- Recurrence spacing: gaps between uses compared with the
frequencyandmax_occurrencesofpayment.agent_recurrence. - Agent key fan-out: distinct users whose open mandates name the same
cnfkey. One key serving thousands of users may be normal for a hosted agent and alarming for a self-hosted one; compare against the provider's baseline. - Fallback ratio: how often a given agent or merchant drives
unresolved_constraintfallbacks, which can be a way to push users into weaker flows. - Post-transaction outcomes: refunds and disputes. The spec notes that
checkout_hashlets the Checkout and Payment Mandates be joined in a dispute, which is also how you join outcomes back to signals for training.
Counter storage and windows are covered in payment velocity limits. Count on open mandate reference, agent key, payee and instrument, not IP address.
Class 4: risk_data, the one free-form field
The closed Payment Mandate schema defines risk_data as a map of relevant risk signals collected by the trusted surface at the time of mandate creation, and defines no keys. That makes it the channel for everything the protocol cannot prove: how the user authenticated, how old the device is, where the approval happened. Because it is untyped, two parties must agree its schema bilaterally. The example below is illustrative only; none of these keys come from the AP2 spec.
{
"risk_data": {
"schema": "example-risk/1",
"user_auth_method": "passkey",
"surface_session_age_s": 312,
"device_first_seen_days": 418,
"approval_country": "DE"
}
}How far to trust it depends on whose signature covers it. An open Payment Mandate may include any property of the closed one, so risk_data can appear in a user-signed open mandate produced on the Trusted Surface, or in a closed mandate. In the autonomous mode the closed mandate is signed by the agent's key, so risk_data there is at best a claim from a party the spec treats as a potential attacker. Prefer values from the user-signed object and version the schema explicitly.
Extracting signals in code
The extractor below runs only after verification passes. It tags each value with strength, source and time, derives mode from the signer, normalises units, and queries history strictly before the decision time so the same code yields leak-free training rows.
from dataclasses import dataclass
from enum import Enum
class Strength(Enum):
PROVEN = 4
ATTESTED = 3
OBSERVED = 2
CLAIMED = 1
@dataclass(frozen=True)
class Signal:
name: str
value: object
strength: Strength
source: str # who asserted it: "verifier", "trusted_surface", "agent", "history"
known_at: int # epoch seconds when this value became knowable
def extract(v, history, decided_at):
"""v is an already VERIFIED presentation; history is our own event log."""
out = []
# Structural: derived from who signed, never from a mode field the agent wrote.
autonomous = v.closed_signed_by_agent_key
out.append(Signal("mode_autonomous", autonomous, Strength.PROVEN, "verifier", decided_at))
amt, hi = v.amount_minor, v.constraint_max_minor # both normalised to minor units
if hi:
out.append(Signal("headroom_ratio", amt / hi, Strength.PROVEN, "verifier", decided_at))
if v.budget_max_minor:
spent = history.budget_spent(v.open_ref, before=decided_at)
out.append(Signal("budget_used", (spent + amt) / v.budget_max_minor,
Strength.OBSERVED, "history", decided_at))
# Behavioural: only events strictly before the decision time (no leakage).
out.append(Signal("open_presentations_10m",
history.presentations(v.open_ref, since=decided_at - 600, before=decided_at),
Strength.OBSERVED, "history", decided_at))
# risk_data: strength depends on whose signature covers it.
for key, val in (v.risk_data or {}).items():
s = Strength.ATTESTED if v.risk_data_signed_by_user else Strength.CLAIMED
out.append(Signal("rd." + key, val, s, v.risk_data_signer, decided_at))
return outThe before=decided_at arguments mean a replay over last month's traffic sees exactly what the live system saw. And strength travels with the value, so policy can state that no combination of claimed signals may lift a decline caused by observed ones.
Worked example: a household restock agent
A user approves, on their bank's Trusted Surface, open mandates for a grocery restock agent: checkout.allowed_merchants with two grocers, payment.amount_range from 10 to 120 USD, payment.budget of 600 USD, and payment.agent_recurrence with frequency MONTHLY and six occurrences. Expiry is set to six months. Purchases one and two, for 84.20 and 91.10 USD, verify cleanly; headroom is 0.70 and 0.76, budget use reaches 29 percent, spacing matches the frequency.
In month three a lookalike site injects instructions into the agent's context, and the agent assembles a checkout at a third merchant. The merchant constraint fails, the verifier returns an error receipt, and nothing is paid. The useful signal is not this decline but the cluster: the same agent provider's keys produced eleven merchant-constraint failures across unrelated users in an hour.
In month four the agent presents two closed mandates against one open mandate within seconds, each for 118 USD, to two verifiers before either issues a receipt. Evaluated alone, each passes recurrence and budget. The signal vector reads: mode autonomous (proven), headroom 0.98 (proven), overlap count 2 (observed, and a protocol violation), budget use would jump to 69 percent (observed), device age from user-signed risk_data 418 days (attested). Policy declines the second; rather than revoking the mandate, the next purchase is routed to direct, human-present approval.
Failure modes
- Trusting a mode or intent field the agent wrote. Derive mode from signatures; treat agent prose as claims.
- Unit mismatch between minor-unit amounts and decimal constraint limits, which hides near-limit probing.
- Leakage: computing features with data from after the decision, typically dispute outcomes or later presentations, which makes offline metrics look far better than production.
- Double counting budget when a retried presentation is recorded twice; key the accumulator on the mandate reference hash so retries are idempotent.
- Penalising selective disclosure. Holders reveal only the disclosures relevant to the transaction, and the Trusted Surface may insert decoy digests. Missing disclosures are by design, not evasion.
Operating the signal layer
Store every receipt you issue and receive; they are signed evidence and the join key to outcomes. Dashboard each signal by agent provider and merchant, because drift in agent traffic arrives in steps when a provider ships a new model. Run new signals in shadow mode for a full recurrence cycle before they can decline anything. Log derived signals and the mandate reference hash rather than full disclosed claims. When verification is in doubt, AP2 verification covers the authority chain in detail, and chargeback handling covers what happens after a dispute.
Trade-offs
Tight constraints make fraud expensive but also fail legitimate purchases, pushing users towards wider mandates. Falling back to human-present approval is safe but removes the convenience agents exist for, so reserve it for compound signals. Richer risk_data helps issuers but spreads personal data; send only what a counterparty has agreed to use.
What to do next
- List every signal you use today and tag each with strength, source and when it becomes knowable.
- Make verification outcomes hard gates and move their rates onto fleet dashboards by agent provider and merchant.
- Derive mode from the signer, and normalise every amount and constraint limit to minor units.
- Count presentations per open mandate reference and alert on any overlap before a rejection receipt.
- Agree a versioned risk_data schema with each counterparty, and prefer values from user-signed mandates.
- Rebuild training data with point-in-time queries and compare offline metrics before and after.