A fraud system answers one question many thousands of times a second: should this payment go through? It must answer inside the authorisation call, while a customer waits at a till or a checkout page, and it learns whether it was right only weeks later, when a chargeback arrives or does not. Those two clocks, milliseconds to decide and weeks to learn, shape the whole architecture.
This article designs the system from the request inward: the synchronous decision path, the features it reads and how streaming keeps them fresh, how rules and models combine into a decision, what happens when parts fail, and how delayed, biased labels feed training. The catalogue of individual signals for agent-initiated payments is covered in AP2 fraud signals; here the focus is the system that computes and acts on any signal.
The decision and its budget
The payment request reaches a decision service with the transaction attributes: card or account token, amount, currency, merchant and category, channel, device and network data where available. The service returns one of four outcomes: approve, approve after step-up authentication such as a one-time code or 3-D Secure challenge, queue for manual review where the business model allows a delay, or decline.
The service gets a slice of the authorisation latency, and the slice is small. An illustrative budget for a 50 ms allowance might be: 5 ms to parse and validate, 15 ms to fetch features in parallel, 5 ms for rules, 10 ms for model inference, 5 ms for the decision policy and response, and 10 ms of headroom for tail latency. The real number comes from your network and processor contracts; the point is that every component on the path needs a hard timeout and a defined behaviour when it expires.
Everything else, computing aggregates, training, investigating, runs off the critical path. The decision service writes each request, the features it saw and the decision it made to an event log. That log is the source for streaming features, for training data and for audit.
Features in three freshness classes
Fraud features fall into three classes by how fresh they must be, and each class has its own pipeline.
| Class | Examples | Computed by | Staleness tolerated |
|---|---|---|---|
| Profile | Account age, historical chargeback count, usual countries | Batch jobs into the online store | Hours to a day |
| Velocity | Transactions per card in 10 minutes, distinct merchants per hour, declines per device | Stream processor | Seconds |
| Request-time | Amount relative to the card's mean, distance from last location, new device flag | Decision service, at request time | None |
Profile and velocity features live in an online store keyed by entity: card, account, device, IP, merchant. A feature store gives you one definition per feature for both the online lookup and the offline training table, which is the main defence against training-serving skew.
Velocity features are where attacks are won or lost. Fraudsters test stolen cards with bursts of small payments, so a counter that lags by a minute gives them a minute. The stream processor keeps keyed state per entity, for example with Flink keyed state and timers to expire old buckets, and writes the current aggregates to the online store.
Velocity counters that tolerate duplicates and lag
Two problems recur. First, the event log delivers at least once, so a counter that increments on every event double-counts after a retry and starts declining honest customers. Second, the streaming aggregate never includes the transaction being scored right now, because that transaction has not been logged yet. The decision service must add it itself, or a burst of ten simultaneous attempts all see a count of zero.
A minute-bucketed counter with idempotent writes handles both. This sketch uses a Redis-style store; the same shape works in a stream processor's state:
BUCKET_S = 60
WINDOW_BUCKETS = 10 # 10-minute window
def record(store, card, txn_id, ts):
"""Called by the stream consumer; safe to replay.
Dedupe and increment run as ONE atomic script: a crash between them
followed by a replay would otherwise never count the transaction."""
key = f"vel:{card}:{ts // BUCKET_S}"
store.eval(RECORD_LUA, 2, f"seen:{txn_id}", key,
86400, BUCKET_S * (WINDOW_BUCKETS + 1))
RECORD_LUA = """
if redis.call('SET', KEYS[1], 1, 'NX', 'EX', ARGV[1]) then
redis.call('INCR', KEYS[2])
redis.call('EXPIRE', KEYS[2], ARGV[2])
end
"""
def count_10m(store, card, now):
"""Called by the decision service; includes the in-flight request."""
first = now // BUCKET_S - WINDOW_BUCKETS + 1
keys = [f"vel:{card}:{b}" for b in range(first, now // BUCKET_S + 1)]
logged = sum(int(v or 0) for v in store.mget(keys))
pending = store.incr(f"inflight:{card}:{now // BUCKET_S}") # concurrent attempts
store.expire(f"inflight:{card}:{now // BUCKET_S}", BUCKET_S * 2)
return logged + pendingThe in-flight counter is approximate: it is never decremented, so a scored transaction is counted twice, once in flight and once when logged, until its minute bucket ages out. That bias is in the safe direction for an attack feature; widen the window or decrement on logging if it causes false declines. Monitor consumer lag; when it exceeds the bucket width, velocity features are lying, and the decision policy should know it.
Rules, models and a decision policy
Mature systems use three layers rather than one. Rules encode hard constraints and fast responses: sanctions and block lists, a merchant category the business never serves, a rule an analyst writes within minutes of seeing a new attack. A model, typically gradient-boosted trees on tabular features, produces a fraud probability that generalises across patterns no one wrote a rule for. A policy layer turns rules, score and context into an outcome.
The policy should be explicit about costs. Approving fraud loses roughly the amount plus fees. Declining a good customer loses the margin on the sale and some probability that the customer never returns. Step-up costs friction and some abandonment. With a calibrated probability p you can compare expected costs directly:
def decide(p_fraud, amount, rules, ctx):
if rules.hard_decline: # block lists, sanctions
return "decline", rules.reason
if rules.hard_approve: # e.g. trusted recurring payment
return "approve", rules.reason
loss_if_approve = p_fraud * amount * ctx.loss_rate
loss_if_decline = (1 - p_fraud) * (amount * ctx.margin + ctx.customer_value_at_risk)
loss_if_stepup = (p_fraud * amount * ctx.loss_rate * ctx.stepup_fraud_pass
+ (1 - p_fraud) * amount * ctx.margin * ctx.stepup_abandon)
options = {"approve": loss_if_approve, "decline": loss_if_decline}
if ctx.stepup_available:
options["step_up"] = loss_if_stepup
outcome = min(options, key=options.get)
return outcome, f"p={p_fraud:.3f} costs={options}"This only works if the probability is calibrated, so check calibration whenever the model changes. Record the reason with every decision; disputes, regulators and analysts will all ask why a payment was declined.
Worked example: a card-testing burst
A batch of stolen card numbers is tested against a donation page: payments of one or two units of currency, each card once or twice, from a small set of IP addresses and a single device fingerprint, about forty attempts a minute.
Per-card velocity sees almost nothing, because each card is used only once or twice. The signal is on the other entities: declines per device in ten minutes jumps from zero to dozens, distinct cards per IP climbs, and the merchant's small-amount authorisation rate spikes. An analyst adds a rule on distinct cards per device within the hour; the model, if trained on similar bursts, already scores these attempts high because of the device and IP aggregates.
The lesson for design is that velocity must be keyed on several entities, not only the card, and that the rule path must deploy in minutes without a model release.
Linking entities without traversing a graph at request time
Organised fraud reuses infrastructure: the same device opens many accounts, the same shipping address receives goods bought with different cards, a ring of accounts pays each other to build history. Those links are a graph of accounts, cards, devices, addresses and IPs joined by shared use.
Walking that graph during authorisation is tempting and dangerous: a traversal from a popular node, such as a carrier-grade NAT address shared by thousands of honest users, has unbounded cost. Instead, compute graph features off the critical path and publish them to the online store like any other feature: the size of the connected component an account belongs to, how many confirmed-fraud entities sit within two hops, how many accounts share this device. Recompute them in batch, and update the most important ones incrementally from the stream when a new link appears. The decision service then reads a few numbers by key, at the same cost as a velocity lookup.
Cap the fan-out when building these features and exclude known shared infrastructure, or a single public Wi-Fi address will link half your customers and the feature will mean nothing.
Labels arrive late and biased
The truth for a transaction comes from chargebacks, customer reports and analyst decisions. Chargebacks can arrive weeks after the payment, so last week's data is not yet labelled: training must wait for labels to mature or treat recent unlabelled transactions carefully. Train on features as they were logged at decision time, not recomputed later, so the model learns from the values it will actually see.
Labels are also biased by the system itself. A declined transaction never produces a chargeback, so its true label is unknown, and a model trained only on approved traffic learns little about the region it already declines. Some teams approve a small, capped random sample of low-amount transactions that would have been declined to keep an unbiased view; others rely on analyst review of declines. Either way, write the choice down.
Fraud patterns move because adversaries adapt. Watch feature and score distributions with drift detection, and release new models first in shadow mode, scoring live traffic without acting, then compare decline and capture rates before promotion.
Failing safely
| Failure | Effect | Response |
|---|---|---|
| Feature fetch times out | Model sees missing values | Use defaults the model was trained with; flag the decision as degraded |
| Model server down | No score | Fall back to rules-only policy with tighter limits |
| Whole service unreachable | No decision | Fail open below an amount limit, fail closed above it, per merchant category |
| Stream consumer lag | Velocity counts stale | Alert on lag; raise step-up rate while lag exceeds bucket width |
| Duplicate events | Inflated counters, false declines | Idempotent writes keyed by transaction ID |
| Hot merchant key | One store shard saturates | Split merchant counters across sub-keys and sum |
| Bad rule deployed | Decline rate spike | Rules versioned, shadow-tested and instantly revertible |
Fail-open versus fail-closed is a business decision, not an engineering one. Make it explicit, per segment, before the outage forces it.
Operating it
- Track decline, step-up and review rates per segment, with alerts on sudden change in either direction.
- Track fraud losses in basis points of volume and the ratio of false declines to caught fraud as labels mature.
- Track p99 latency of the whole decision and of each dependency against its timeout.
- Track feature freshness: stream consumer lag and the age of the newest batch profile.
- Track the share of degraded decisions; a slow rise usually precedes an incident.
Trade-offs
- Tighter thresholds catch more fraud and decline more good customers; the cost-based policy makes the trade visible instead of implicit.
- Richer features raise accuracy and add latency and failure points on the synchronous path.
- Rules react in minutes but accumulate into an unmaintainable pile; models generalise but need weeks of labels to learn a new pattern.
- Exploration samples reduce label bias at a known, capped fraud cost.
What to do next
- Write down the decision latency budget and give every dependency a timeout and a fallback.
- Log every request with the exact features used and the decision reason to an append-only event log.
- Key velocity features on card, account, device, IP and merchant, with idempotent updates.
- Add the in-flight transaction to velocity counts at request time.
- Replace a single score threshold with a cost-based policy and verify model calibration.
- Decide fail-open and fail-closed limits per segment and test them in a game day.
- Define label maturity and how declined transactions are treated in training data.
- Put drift monitoring and shadow scoring in front of every model release.