A payment system looks simple from the checkout page: take an amount and a card, get back yes or no. Underneath, it is a distributed system where the most important component, the payment service provider (PSP) that talks to card networks and banks, is outside your control, sometimes slow, occasionally down, and able to leave you not knowing whether a customer was charged.
This site already covers the single-provider design, with its state machine, data model and reconciliation, in designing a payment system, and the card rails in the credit card payment flow. This article goes one layer up, to the problem most growing merchants hit next: orchestrating several providers. Why would you, how do you route a payment, how do you fail over without charging twice, and how do you resolve the attempt whose outcome you do not know? We will build the data model, the router and the recovery worker, and walk one payment through a provider outage.
Why orchestrate several providers
Teams add a second PSP for four reasons: availability (a provider outage should not stop checkout), authorisation rate (some acquirers approve more of a given card type or region), cost (fees differ by card, currency and volume) and coverage (local payment methods only some providers support). Each reason is real, and each adds the same problem: the system now has to choose, and when a choice fails, choose again safely.
The architecture separates responsibilities so that safety lives in one place. The orchestrator owns payment and attempt state and is the only component that decides to retry. The router is a pure function from payment and provider health to an ordered list of candidates. Adapters translate a provider-neutral request to each PSP's API and map responses back to a small set of outcomes. The token vault holds card data so that any provider can be used. The recovery worker resolves unknown outcomes, and an outbox publishes events to the ledger and the rest of the business in the same transaction as the state change.
Payments and attempts: the data model
The key modelling decision is to separate the payment (the customer's intent to pay an amount for an order) from attempts (each try against a specific provider). A payment succeeds at most once; it may have many attempts. Every attempt gets its own idempotency key, which is what you send to the provider, so a network retry of the same attempt is deduplicated by the PSP while a new attempt on a different provider is a deliberate, separate charge.
CREATE TABLE payments (
payment_id uuid PRIMARY KEY,
merchant_ref text NOT NULL,
idempotency_key text NOT NULL UNIQUE, -- from the client
amount_minor bigint NOT NULL,
currency char(3) NOT NULL,
status text NOT NULL, -- PENDING, SUCCEEDED, FAILED, REQUIRES_ACTION
vault_token text NOT NULL,
version int NOT NULL DEFAULT 0
);
CREATE TABLE payment_attempts (
attempt_id uuid PRIMARY KEY,
payment_id uuid NOT NULL REFERENCES payments,
seq int NOT NULL,
provider text NOT NULL,
provider_key text NOT NULL UNIQUE, -- idempotency key sent to the PSP
provider_ref text, -- PSP's charge id, once known
status text NOT NULL, -- CREATED, SENT, APPROVED, DECLINED_SOFT,
-- DECLINED_HARD, UNKNOWN, VOIDED
decline_code text,
UNIQUE (payment_id, seq)
);
-- at most one approved attempt per payment, enforced by the database
CREATE UNIQUE INDEX one_approval ON payment_attempts(payment_id)
WHERE status = 'APPROVED';The partial unique index is the last line of defence: whatever bug exists in the retry logic, the database refuses to record two approvals for one payment. Client-facing idempotency (the same checkout request twice returns the same payment) works as described in idempotency keys, and the event publication follows the outbox pattern.
Routing: choosing a provider
Routing turns a payment into an ordered candidate list. Keep it deterministic and explainable: hard rules first (which providers support this method, currency and country), then a score. Health comes from your own measurements of recent attempts, not from provider status pages, which lag.
def route(payment, providers, health):
"""Return providers to try, best first. Pure: no I/O, easy to test and replay."""
eligible = [p for p in providers
if payment.method in p.methods
and payment.currency in p.currencies
and health[p.name].circuit != "OPEN"]
def score(p):
h = health[p.name]
auth = h.auth_rate(payment.card_brand, payment.issuer_country) # last 15 min
fee = p.fee_minor(payment) # expected cost
latency_penalty = 0.0 if h.p95_ms < 1500 else 0.05
# value of an approval minus cost; an unapproved payment earns nothing
return auth * payment.amount_minor - fee - latency_penalty * payment.amount_minor
return sorted(eligible, key=score, reverse=True)Two cautions. First, auth-rate estimates from small samples are noisy; use a minimum sample size and blend with a long-run prior, or the router will chase noise. Second, always send a small share of traffic to every eligible provider, so that you keep measuring the ones you are not currently favouring, and a recovered provider is noticed.
The circuit state comes from a breaker per provider: open it when the rate of timeouts and 5xx responses over a short window crosses a threshold, send a trickle of probe traffic while half-open, and close it when those succeed. Declines do not count as failures; a provider that declines a stolen card is working.
The unknown outcome
Every attempt ends in one of three ways. Approved: done. Declined: the provider answered no. Unknown: you sent the request and did not get a definitive answer, because of a timeout, a dropped connection, or a 5xx after the request may have reached the card network. Unknown is the state that causes double charges, and the rule for it is absolute: never start a new attempt while any attempt is unknown.
An unknown attempt is resolved by asking the provider what happened, using the same idempotency key or your reference, until it answers definitively. If it cannot answer within your checkout budget, the payment stays pending and the customer sees "we are confirming your payment", not a failure and not a retry.
def resolve_unknown(attempt, adapter, max_wait_s=900):
"""Run by the recovery worker for attempts stuck in UNKNOWN."""
for delay in backoff(start=2, factor=2, cap=120, total=max_wait_s):
result = adapter.lookup(provider_key=attempt.provider_key)
if result.found and result.status in ("approved", "captured"):
return mark(attempt, "APPROVED", provider_ref=result.ref)
if result.found and result.status in ("declined", "failed"):
return mark(attempt, "DECLINED_HARD" if result.hard else "DECLINED_SOFT")
if not result.found and age(attempt) > adapter.not_found_is_final_after:
# the PSP says it never saw the request, and enough time has passed
# that a late arrival is no longer plausible for this provider
return mark(attempt, "DECLINED_SOFT", decline_code="never_received")
sleep(delay)
alert("payment attempt unresolved", attempt.attempt_id) # human follow-upNotice the not_found_is_final_after setting. "Not found" from a provider seconds after a timeout does not prove the request was lost; it may still be in flight. Agree per provider on how long it takes before not-found is final, and some teams additionally void or refund any approval that appears afterwards. Webhooks help, since providers usually send one on completion, but treat them as a hint that triggers a lookup, because they can arrive late, out of order, or twice.
Cascading and failover, worked through
Worked example. A customer pays EUR 84.00. The router returns [A, B]. The orchestrator creates attempt 1 on A and sends it. A's latency has been rising; after 8 seconds the call times out. Attempt 1 is UNKNOWN, so the orchestrator does not touch B. The recovery worker looks up attempt 1 about 2, 6 and 14 seconds after the timeout (sleeping 2, 4 and 8 seconds between tries): A returns not found twice, then approved. The payment succeeds on A, and the customer, who saw a spinner and then a confirmation, was charged once. Had the orchestrator cascaded to B on timeout, the customer would have been charged twice and you would be issuing a refund and an apology.
Now a different case: A answers quickly with a decline. Whether to cascade depends on the decline type:
| Outcome | Example reasons | Cascade to next provider? |
|---|---|---|
| Hard decline | stolen or lost card, closed account, invalid number | no; retrying elsewhere can breach network rules and annoys the issuer |
| Soft decline, issuer | insufficient funds, do not honour | rarely; same issuer decides regardless of acquirer |
| Soft decline, acquirer or provider | processing error, acquirer-side issue | yes, once, if the router has another acquirer |
| Authentication required | issuer wants 3-D Secure | no; return REQUIRES_ACTION to the customer |
| Unknown | timeout, connection reset, 5xx | never until resolved |
Card networks publish rules limiting how often a declined card may be retried and charge fees for excessive retries; the specific limits and advice codes change, so encode them as configuration your payments team owns rather than as constants. Each adapter must map its provider's response codes to these categories, and that mapping is the code most worth reviewing and testing.
Tokens, ledger and reconciliation
Multi-provider routing only works if any provider can charge the card. If each PSP tokenises the card in its own vault, your saved cards are locked to that provider. The options are: run your own token vault (a small, isolated, PCI DSS-scoped service that stores card numbers and forwards them to providers), use a third-party vault that forwards to many PSPs, or use network tokens issued by the card schemes, which are provider-portable and update automatically when a card is reissued. Whichever you pick, keep card data out of the orchestrator and every other service: they handle only a vault token, which keeps them out of the most demanding PCI scope.
Routing freedom ends at approval. Everything that happens to a payment afterwards (capture, partial capture, void, refund, and the dispute that may arrive weeks later) must go to the provider that approved it, referencing that attempt's provider_ref. You cannot refund on B a charge that A approved; B has never heard of it. So follow-up operations look up the approved attempt and use its adapter, never the router. This is also why attempts are permanent records rather than scratch state: a refund requested in March for a payment approved in January needs the provider, the provider reference and the captured amount, and if that provider's contract has ended you still need the adapter, or a documented manual process, for the refund window. When you retire a provider, stop routing new payments to it first, and keep its adapter running for refunds and disputes until the last of its payments is out of the dispute window.
Money then flows into the ledger from events, never from adapter responses directly. An approved attempt emits an event through the outbox; the ledger posts the double-entry lines; reconciliation compares each provider's settlement report against the ledger daily and flags anything on one side only. With several providers, reconciliation is per provider, and the unresolved-attempt alerts above should be zero before the day closes.
Failure modes and trade-offs
| Failure | What goes wrong | Defence |
|---|---|---|
| Cascade on timeout | customer charged twice | UNKNOWN blocks new attempts; lookup first |
| Router chases noisy auth rates | traffic flaps between providers | minimum samples, priors, exploration share |
| Adapter mis-maps a decline | hard declines retried, fees, issuer blocks | contract tests per code; review mapping changes |
| Webhook before API response | state set twice or out of order | state transitions guarded by version; webhooks trigger lookups |
| Provider-locked tokens | failover impossible for saved cards | own or third-party vault, network tokens |
| Late approval after not-found | charge with no order | per-provider finality window; auto-void stragglers |
The trade-off throughout is between conversion and safety. Aggressive cascading recovers some soft declines and increases double-charge and fee risk; a strict UNKNOWN rule occasionally leaves a customer waiting. Measure both: auth rate per provider and route, time-to-resolution of unknown attempts, refunds caused by duplicates, and cost per approved payment. Stripe's platform architecture shows how one provider solves the same problems internally.
What to do next
- Split your model into payments and attempts, with a per-attempt provider idempotency key and a unique index allowing one approval per payment.
- Add the rule that no attempt starts while another is UNKNOWN, and a recovery worker that resolves unknowns by lookup.
- Agree with each provider how long before "not found" is final, and auto-void late approvals.
- Write the decline-code mapping per adapter, with contract tests for every code you have seen in production.
- Build the router as a pure function with circuit breakers and an exploration share; log every routing decision.
- Move saved cards to a provider-neutral vault or network tokens before relying on failover.
- Reconcile each provider's settlement report against the ledger daily.