A payment system looks simple from the checkout page: take an amount and a card, get back yes or no. Underneath, it is a distributed system where the most important component, the payment service provider (PSP) that talks to card networks and banks, is outside your control, sometimes slow, occasionally down, and able to leave you not knowing whether a customer was charged.

This site already covers the single-provider design, with its state machine, data model and reconciliation, in designing a payment system, and the card rails in the credit card payment flow. This article goes one layer up, to the problem most growing merchants hit next: orchestrating several providers. Why would you, how do you route a payment, how do you fail over without charging twice, and how do you resolve the attempt whose outcome you do not know? We will build the data model, the router and the recovery worker, and walk one payment through a provider outage.

Why orchestrate several providers

Teams add a second PSP for four reasons: availability (a provider outage should not stop checkout), authorisation rate (some acquirers approve more of a given card type or region), cost (fees differ by card, currency and volume) and coverage (local payment methods only some providers support). Each reason is real, and each adds the same problem: the system now has to choose, and when a choice fails, choose again safely.

The architecture separates responsibilities so that safety lives in one place. The orchestrator owns payment and attempt state and is the only component that decides to retry. The router is a pure function from payment and provider health to an ordered list of candidates. Adapters translate a provider-neutral request to each PSP's API and map responses back to a small set of outcomes. The token vault holds card data so that any provider can be used. The recovery worker resolves unknown outcomes, and an outbox publishes events to the ledger and the rest of the business in the same transaction as the state change.

Payment orchestration: one payment, many attempts, many providersCheckout / APIIdempotency-KeyPayment orchestratorpayment + attempt stateRouterrules, health, costPSP adapter APSP adapter BPSP adapter CToken vaultPAN, network tokensPayments DBpayments, attemptsOutboxevents, same txnRecovery workerresolves UNKNOWNLedgerdouble entryReconciliationPSP reports vs ledgerdetokenisewriteeventsstatus queryThe orchestrator owns state; adapters translate; the recovery worker turns every timeout into a known outcome before money is retried.
Multi-provider payment orchestration: state and retry decisions live in the orchestrator, routing is a pure function, and a recovery worker resolves timeouts.

Payments and attempts: the data model

The key modelling decision is to separate the payment (the customer's intent to pay an amount for an order) from attempts (each try against a specific provider). A payment succeeds at most once; it may have many attempts. Every attempt gets its own idempotency key, which is what you send to the provider, so a network retry of the same attempt is deduplicated by the PSP while a new attempt on a different provider is a deliberate, separate charge.

CREATE TABLE payments (
  payment_id      uuid PRIMARY KEY,
  merchant_ref    text NOT NULL,
  idempotency_key text NOT NULL UNIQUE,      -- from the client
  amount_minor    bigint NOT NULL,
  currency        char(3) NOT NULL,
  status          text NOT NULL,             -- PENDING, SUCCEEDED, FAILED, REQUIRES_ACTION
  vault_token     text NOT NULL,
  version         int NOT NULL DEFAULT 0
);

CREATE TABLE payment_attempts (
  attempt_id      uuid PRIMARY KEY,
  payment_id      uuid NOT NULL REFERENCES payments,
  seq             int NOT NULL,
  provider        text NOT NULL,
  provider_key    text NOT NULL UNIQUE,      -- idempotency key sent to the PSP
  provider_ref    text,                      -- PSP's charge id, once known
  status          text NOT NULL,             -- CREATED, SENT, APPROVED, DECLINED_SOFT,
                                             -- DECLINED_HARD, UNKNOWN, VOIDED
  decline_code    text,
  UNIQUE (payment_id, seq)
);
-- at most one approved attempt per payment, enforced by the database
CREATE UNIQUE INDEX one_approval ON payment_attempts(payment_id)
  WHERE status = 'APPROVED';

The partial unique index is the last line of defence: whatever bug exists in the retry logic, the database refuses to record two approvals for one payment. Client-facing idempotency (the same checkout request twice returns the same payment) works as described in idempotency keys, and the event publication follows the outbox pattern.

Routing: choosing a provider

Routing turns a payment into an ordered candidate list. Keep it deterministic and explainable: hard rules first (which providers support this method, currency and country), then a score. Health comes from your own measurements of recent attempts, not from provider status pages, which lag.

def route(payment, providers, health):
    """Return providers to try, best first. Pure: no I/O, easy to test and replay."""
    eligible = [p for p in providers
                if payment.method in p.methods
                and payment.currency in p.currencies
                and health[p.name].circuit != "OPEN"]

    def score(p):
        h = health[p.name]
        auth = h.auth_rate(payment.card_brand, payment.issuer_country)  # last 15 min
        fee = p.fee_minor(payment)                                      # expected cost
        latency_penalty = 0.0 if h.p95_ms < 1500 else 0.05
        # value of an approval minus cost; an unapproved payment earns nothing
        return auth * payment.amount_minor - fee - latency_penalty * payment.amount_minor

    return sorted(eligible, key=score, reverse=True)

Two cautions. First, auth-rate estimates from small samples are noisy; use a minimum sample size and blend with a long-run prior, or the router will chase noise. Second, always send a small share of traffic to every eligible provider, so that you keep measuring the ones you are not currently favouring, and a recovered provider is noticed.

The circuit state comes from a breaker per provider: open it when the rate of timeouts and 5xx responses over a short window crosses a threshold, send a trickle of probe traffic while half-open, and close it when those succeed. Declines do not count as failures; a provider that declines a stolen card is working.

The unknown outcome

Every attempt ends in one of three ways. Approved: done. Declined: the provider answered no. Unknown: you sent the request and did not get a definitive answer, because of a timeout, a dropped connection, or a 5xx after the request may have reached the card network. Unknown is the state that causes double charges, and the rule for it is absolute: never start a new attempt while any attempt is unknown.

An unknown attempt is resolved by asking the provider what happened, using the same idempotency key or your reference, until it answers definitively. If it cannot answer within your checkout budget, the payment stays pending and the customer sees "we are confirming your payment", not a failure and not a retry.

def resolve_unknown(attempt, adapter, max_wait_s=900):
    """Run by the recovery worker for attempts stuck in UNKNOWN."""
    for delay in backoff(start=2, factor=2, cap=120, total=max_wait_s):
        result = adapter.lookup(provider_key=attempt.provider_key)
        if result.found and result.status in ("approved", "captured"):
            return mark(attempt, "APPROVED", provider_ref=result.ref)
        if result.found and result.status in ("declined", "failed"):
            return mark(attempt, "DECLINED_HARD" if result.hard else "DECLINED_SOFT")
        if not result.found and age(attempt) > adapter.not_found_is_final_after:
            # the PSP says it never saw the request, and enough time has passed
            # that a late arrival is no longer plausible for this provider
            return mark(attempt, "DECLINED_SOFT", decline_code="never_received")
        sleep(delay)
    alert("payment attempt unresolved", attempt.attempt_id)   # human follow-up

Notice the not_found_is_final_after setting. "Not found" from a provider seconds after a timeout does not prove the request was lost; it may still be in flight. Agree per provider on how long it takes before not-found is final, and some teams additionally void or refund any approval that appears afterwards. Webhooks help, since providers usually send one on completion, but treat them as a hint that triggers a lookup, because they can arrive late, out of order, or twice.

Cascading and failover, worked through

Worked example. A customer pays EUR 84.00. The router returns [A, B]. The orchestrator creates attempt 1 on A and sends it. A's latency has been rising; after 8 seconds the call times out. Attempt 1 is UNKNOWN, so the orchestrator does not touch B. The recovery worker looks up attempt 1 about 2, 6 and 14 seconds after the timeout (sleeping 2, 4 and 8 seconds between tries): A returns not found twice, then approved. The payment succeeds on A, and the customer, who saw a spinner and then a confirmation, was charged once. Had the orchestrator cascaded to B on timeout, the customer would have been charged twice and you would be issuing a refund and an apology.

Now a different case: A answers quickly with a decline. Whether to cascade depends on the decline type:

OutcomeExample reasonsCascade to next provider?
Hard declinestolen or lost card, closed account, invalid numberno; retrying elsewhere can breach network rules and annoys the issuer
Soft decline, issuerinsufficient funds, do not honourrarely; same issuer decides regardless of acquirer
Soft decline, acquirer or providerprocessing error, acquirer-side issueyes, once, if the router has another acquirer
Authentication requiredissuer wants 3-D Secureno; return REQUIRES_ACTION to the customer
Unknowntimeout, connection reset, 5xxnever until resolved

Card networks publish rules limiting how often a declined card may be retried and charge fees for excessive retries; the specific limits and advice codes change, so encode them as configuration your payments team owns rather than as constants. Each adapter must map its provider's response codes to these categories, and that mapping is the code most worth reviewing and testing.

Tokens, ledger and reconciliation

Multi-provider routing only works if any provider can charge the card. If each PSP tokenises the card in its own vault, your saved cards are locked to that provider. The options are: run your own token vault (a small, isolated, PCI DSS-scoped service that stores card numbers and forwards them to providers), use a third-party vault that forwards to many PSPs, or use network tokens issued by the card schemes, which are provider-portable and update automatically when a card is reissued. Whichever you pick, keep card data out of the orchestrator and every other service: they handle only a vault token, which keeps them out of the most demanding PCI scope.

Routing freedom ends at approval. Everything that happens to a payment afterwards (capture, partial capture, void, refund, and the dispute that may arrive weeks later) must go to the provider that approved it, referencing that attempt's provider_ref. You cannot refund on B a charge that A approved; B has never heard of it. So follow-up operations look up the approved attempt and use its adapter, never the router. This is also why attempts are permanent records rather than scratch state: a refund requested in March for a payment approved in January needs the provider, the provider reference and the captured amount, and if that provider's contract has ended you still need the adapter, or a documented manual process, for the refund window. When you retire a provider, stop routing new payments to it first, and keep its adapter running for refunds and disputes until the last of its payments is out of the dispute window.

Money then flows into the ledger from events, never from adapter responses directly. An approved attempt emits an event through the outbox; the ledger posts the double-entry lines; reconciliation compares each provider's settlement report against the ledger daily and flags anything on one side only. With several providers, reconciliation is per provider, and the unresolved-attempt alerts above should be zero before the day closes.

Failure modes and trade-offs

FailureWhat goes wrongDefence
Cascade on timeoutcustomer charged twiceUNKNOWN blocks new attempts; lookup first
Router chases noisy auth ratestraffic flaps between providersminimum samples, priors, exploration share
Adapter mis-maps a declinehard declines retried, fees, issuer blockscontract tests per code; review mapping changes
Webhook before API responsestate set twice or out of orderstate transitions guarded by version; webhooks trigger lookups
Provider-locked tokensfailover impossible for saved cardsown or third-party vault, network tokens
Late approval after not-foundcharge with no orderper-provider finality window; auto-void stragglers

The trade-off throughout is between conversion and safety. Aggressive cascading recovers some soft declines and increases double-charge and fee risk; a strict UNKNOWN rule occasionally leaves a customer waiting. Measure both: auth rate per provider and route, time-to-resolution of unknown attempts, refunds caused by duplicates, and cost per approved payment. Stripe's platform architecture shows how one provider solves the same problems internally.

What to do next

  1. Split your model into payments and attempts, with a per-attempt provider idempotency key and a unique index allowing one approval per payment.
  2. Add the rule that no attempt starts while another is UNKNOWN, and a recovery worker that resolves unknowns by lookup.
  3. Agree with each provider how long before "not found" is final, and auto-void late approvals.
  4. Write the decline-code mapping per adapter, with contract tests for every code you have seen in production.
  5. Build the router as a pure function with circuit breakers and an exploration share; log every routing decision.
  6. Move saved cards to a provider-neutral vault or network tokens before relying on failover.
  7. Reconcile each provider's settlement report against the ledger daily.
Key takeaway: A multi-provider payment system is safe when one component owns state and retries, every try is a separate attempt with its own provider idempotency key, the database allows only one approval per payment, and no new attempt starts while an earlier one has an unknown outcome. Route with a pure, measured function, cascade only on declines that another acquirer could change, keep cards in a provider-neutral vault, and reconcile every provider daily.