A fraud system answers one question many thousands of times a second: should this payment go through? It must answer inside the authorisation call, while a customer waits at a till or a checkout page, and it learns whether it was right only weeks later, when a chargeback arrives or does not. Those two clocks, milliseconds to decide and weeks to learn, shape the whole architecture.

This article designs the system from the request inward: the synchronous decision path, the features it reads and how streaming keeps them fresh, how rules and models combine into a decision, what happens when parts fail, and how delayed, biased labels feed training. The catalogue of individual signals for agent-initiated payments is covered in AP2 fraud signals; here the focus is the system that computes and acts on any signal.

Advertisement

The decision and its budget

The payment request reaches a decision service with the transaction attributes: card or account token, amount, currency, merchant and category, channel, device and network data where available. The service returns one of four outcomes: approve, approve after step-up authentication such as a one-time code or 3-D Secure challenge, queue for manual review where the business model allows a delay, or decline.

The service gets a slice of the authorisation latency, and the slice is small. An illustrative budget for a 50 ms allowance might be: 5 ms to parse and validate, 15 ms to fetch features in parallel, 5 ms for rules, 10 ms for model inference, 5 ms for the decision policy and response, and 10 ms of headroom for tail latency. The real number comes from your network and processor contracts; the point is that every component on the path needs a hard timeout and a defined behaviour when it expires.

Everything else, computing aggregates, training, investigating, runs off the critical path. The decision service writes each request, the features it saw and the decision it made to an event log. That log is the source for streaming features, for training data and for audit.

Two paths: decide in milliseconds, learn over weeksSYNCHRONOUS (inside the authorisation budget)PaymentauthorisationDecision serviceassemble, score, decideOnline featuresRules engineModel serverOutcomeapprove, step-up, review, declineASYNCHRONOUS (seconds to weeks)Event logevery request + decisionStream processorvelocity, aggregatesLakelogged featuresLabelschargebacks, casesTrainingpoint-in-time joinsShadow, then promotelogupdatedeploy
The synchronous path reads precomputed state and must answer within its slice of the authorisation budget. The asynchronous path turns logged events and late-arriving labels into fresh features and new models.

Features in three freshness classes

Fraud features fall into three classes by how fresh they must be, and each class has its own pipeline.

ClassExamplesComputed byStaleness tolerated
ProfileAccount age, historical chargeback count, usual countriesBatch jobs into the online storeHours to a day
VelocityTransactions per card in 10 minutes, distinct merchants per hour, declines per deviceStream processorSeconds
Request-timeAmount relative to the card's mean, distance from last location, new device flagDecision service, at request timeNone

Profile and velocity features live in an online store keyed by entity: card, account, device, IP, merchant. A feature store gives you one definition per feature for both the online lookup and the offline training table, which is the main defence against training-serving skew.

Velocity features are where attacks are won or lost. Fraudsters test stolen cards with bursts of small payments, so a counter that lags by a minute gives them a minute. The stream processor keeps keyed state per entity, for example with Flink keyed state and timers to expire old buckets, and writes the current aggregates to the online store.

Advertisement

Velocity counters that tolerate duplicates and lag

Two problems recur. First, the event log delivers at least once, so a counter that increments on every event double-counts after a retry and starts declining honest customers. Second, the streaming aggregate never includes the transaction being scored right now, because that transaction has not been logged yet. The decision service must add it itself, or a burst of ten simultaneous attempts all see a count of zero.

A minute-bucketed counter with idempotent writes handles both. This sketch uses a Redis-style store; the same shape works in a stream processor's state:

BUCKET_S = 60
WINDOW_BUCKETS = 10          # 10-minute window

def record(store, card, txn_id, ts):
    """Called by the stream consumer; safe to replay.
    Dedupe and increment run as ONE atomic script: a crash between them
    followed by a replay would otherwise never count the transaction."""
    key = f"vel:{card}:{ts // BUCKET_S}"
    store.eval(RECORD_LUA, 2, f"seen:{txn_id}", key,
               86400, BUCKET_S * (WINDOW_BUCKETS + 1))

RECORD_LUA = """
if redis.call('SET', KEYS[1], 1, 'NX', 'EX', ARGV[1]) then
  redis.call('INCR', KEYS[2])
  redis.call('EXPIRE', KEYS[2], ARGV[2])
end
"""

def count_10m(store, card, now):
    """Called by the decision service; includes the in-flight request."""
    first = now // BUCKET_S - WINDOW_BUCKETS + 1
    keys = [f"vel:{card}:{b}" for b in range(first, now // BUCKET_S + 1)]
    logged = sum(int(v or 0) for v in store.mget(keys))
    pending = store.incr(f"inflight:{card}:{now // BUCKET_S}")   # concurrent attempts
    store.expire(f"inflight:{card}:{now // BUCKET_S}", BUCKET_S * 2)
    return logged + pending

The in-flight counter is approximate: it is never decremented, so a scored transaction is counted twice, once in flight and once when logged, until its minute bucket ages out. That bias is in the safe direction for an attack feature; widen the window or decrement on logging if it causes false declines. Monitor consumer lag; when it exceeds the bucket width, velocity features are lying, and the decision policy should know it.

Rules, models and a decision policy

Mature systems use three layers rather than one. Rules encode hard constraints and fast responses: sanctions and block lists, a merchant category the business never serves, a rule an analyst writes within minutes of seeing a new attack. A model, typically gradient-boosted trees on tabular features, produces a fraud probability that generalises across patterns no one wrote a rule for. A policy layer turns rules, score and context into an outcome.

The policy should be explicit about costs. Approving fraud loses roughly the amount plus fees. Declining a good customer loses the margin on the sale and some probability that the customer never returns. Step-up costs friction and some abandonment. With a calibrated probability p you can compare expected costs directly:

def decide(p_fraud, amount, rules, ctx):
    if rules.hard_decline:                      # block lists, sanctions
        return "decline", rules.reason
    if rules.hard_approve:                      # e.g. trusted recurring payment
        return "approve", rules.reason
    loss_if_approve = p_fraud * amount * ctx.loss_rate
    loss_if_decline = (1 - p_fraud) * (amount * ctx.margin + ctx.customer_value_at_risk)
    loss_if_stepup  = (p_fraud * amount * ctx.loss_rate * ctx.stepup_fraud_pass
                       + (1 - p_fraud) * amount * ctx.margin * ctx.stepup_abandon)
    options = {"approve": loss_if_approve, "decline": loss_if_decline}
    if ctx.stepup_available:
        options["step_up"] = loss_if_stepup
    outcome = min(options, key=options.get)
    return outcome, f"p={p_fraud:.3f} costs={options}"

This only works if the probability is calibrated, so check calibration whenever the model changes. Record the reason with every decision; disputes, regulators and analysts will all ask why a payment was declined.

Worked example: a card-testing burst

A batch of stolen card numbers is tested against a donation page: payments of one or two units of currency, each card once or twice, from a small set of IP addresses and a single device fingerprint, about forty attempts a minute.

Per-card velocity sees almost nothing, because each card is used only once or twice. The signal is on the other entities: declines per device in ten minutes jumps from zero to dozens, distinct cards per IP climbs, and the merchant's small-amount authorisation rate spikes. An analyst adds a rule on distinct cards per device within the hour; the model, if trained on similar bursts, already scores these attempts high because of the device and IP aggregates.

The lesson for design is that velocity must be keyed on several entities, not only the card, and that the rule path must deploy in minutes without a model release.

Linking entities without traversing a graph at request time

Organised fraud reuses infrastructure: the same device opens many accounts, the same shipping address receives goods bought with different cards, a ring of accounts pays each other to build history. Those links are a graph of accounts, cards, devices, addresses and IPs joined by shared use.

Walking that graph during authorisation is tempting and dangerous: a traversal from a popular node, such as a carrier-grade NAT address shared by thousands of honest users, has unbounded cost. Instead, compute graph features off the critical path and publish them to the online store like any other feature: the size of the connected component an account belongs to, how many confirmed-fraud entities sit within two hops, how many accounts share this device. Recompute them in batch, and update the most important ones incrementally from the stream when a new link appears. The decision service then reads a few numbers by key, at the same cost as a velocity lookup.

Cap the fan-out when building these features and exclude known shared infrastructure, or a single public Wi-Fi address will link half your customers and the feature will mean nothing.

Labels arrive late and biased

The truth for a transaction comes from chargebacks, customer reports and analyst decisions. Chargebacks can arrive weeks after the payment, so last week's data is not yet labelled: training must wait for labels to mature or treat recent unlabelled transactions carefully. Train on features as they were logged at decision time, not recomputed later, so the model learns from the values it will actually see.

Labels are also biased by the system itself. A declined transaction never produces a chargeback, so its true label is unknown, and a model trained only on approved traffic learns little about the region it already declines. Some teams approve a small, capped random sample of low-amount transactions that would have been declined to keep an unbiased view; others rely on analyst review of declines. Either way, write the choice down.

Fraud patterns move because adversaries adapt. Watch feature and score distributions with drift detection, and release new models first in shadow mode, scoring live traffic without acting, then compare decline and capture rates before promotion.

Failing safely

FailureEffectResponse
Feature fetch times outModel sees missing valuesUse defaults the model was trained with; flag the decision as degraded
Model server downNo scoreFall back to rules-only policy with tighter limits
Whole service unreachableNo decisionFail open below an amount limit, fail closed above it, per merchant category
Stream consumer lagVelocity counts staleAlert on lag; raise step-up rate while lag exceeds bucket width
Duplicate eventsInflated counters, false declinesIdempotent writes keyed by transaction ID
Hot merchant keyOne store shard saturatesSplit merchant counters across sub-keys and sum
Bad rule deployedDecline rate spikeRules versioned, shadow-tested and instantly revertible

Fail-open versus fail-closed is a business decision, not an engineering one. Make it explicit, per segment, before the outage forces it.

Operating it

  • Track decline, step-up and review rates per segment, with alerts on sudden change in either direction.
  • Track fraud losses in basis points of volume and the ratio of false declines to caught fraud as labels mature.
  • Track p99 latency of the whole decision and of each dependency against its timeout.
  • Track feature freshness: stream consumer lag and the age of the newest batch profile.
  • Track the share of degraded decisions; a slow rise usually precedes an incident.

Trade-offs

  • Tighter thresholds catch more fraud and decline more good customers; the cost-based policy makes the trade visible instead of implicit.
  • Richer features raise accuracy and add latency and failure points on the synchronous path.
  • Rules react in minutes but accumulate into an unmaintainable pile; models generalise but need weeks of labels to learn a new pattern.
  • Exploration samples reduce label bias at a known, capped fraud cost.

What to do next

  1. Write down the decision latency budget and give every dependency a timeout and a fallback.
  2. Log every request with the exact features used and the decision reason to an append-only event log.
  3. Key velocity features on card, account, device, IP and merchant, with idempotent updates.
  4. Add the in-flight transaction to velocity counts at request time.
  5. Replace a single score threshold with a cost-based policy and verify model calibration.
  6. Decide fail-open and fail-closed limits per segment and test them in a game day.
  7. Define label maturity and how declined transactions are treated in training data.
  8. Put drift monitoring and shadow scoring in front of every model release.
Key takeaway: A real-time fraud system decides in milliseconds and learns over weeks, so it splits into a synchronous path that only reads precomputed state and an asynchronous path that builds features and models. Velocity features keyed on several entities are where attacks are caught, and they need idempotent updates plus the in-flight transaction to be accurate. Combine fast rules with a calibrated model through an explicit cost-based policy, define degraded behaviour for every dependency, and train on logged features with an honest treatment of delayed and biased labels.