Stripe is often described as a payments company, but architecturally it is an API company. Its product is a contract: objects, state machines and delivery guarantees that let merchants move money without building connections to card networks and banks. Understanding it means understanding two layers: the public contract you integrate against (PaymentIntents, idempotency keys, webhooks, API versions), and the internal machinery Stripe has written about publicly, such as its Ledger and its DocDB database platform.

This article explains each piece of the contract from first principles, shows the code an integrator should write, and walks one payment from click to fulfilment, including a timeout and a duplicate webhook. For a vendor-neutral payment design with your own processor connections, read Designing a Payment System alongside this page.

Advertisement

Two views of one platform

Browser / appPayment Element, 3DSMerchant serversecret key, order DBclient_secretStripe API edgeauth, version, idempotencyconfirmPOST + Idempotency-KeyPayment objectsPaymentIntent state machineCard networks, banksauthorize, capture, settlerouteEvents + webhookssigned, retried, unorderedstate changePOST eventLedgerdouble-entry, clearingfund flowsDocDB shardsproxy + chunk routingread/writeThe public contract (left) is what you integrate against; Ledger and DocDB (right) are what Stripe has published about the inside
Stripe from the outside in. Merchants call the API with a secret key and an idempotency key, browsers confirm with a client secret, state changes become signed events, and money movement is recorded in a double-entry ledger.

Your server holds a secret key and creates objects. The browser or app holds only a publishable key and a per-payment client secret, and Stripe's client libraries collect card details and run 3D Secure, so card data never touches your servers and most of your stack stays out of the strictest PCI DSS scope.

Every API request passes through the same front door, which authenticates the key, picks the API version and applies idempotency. Payment objects then talk to networks and banks, every state change emits an event pushed to your webhooks, and each movement of funds is recorded as balanced ledger entries.

The PaymentIntent: one object per payment, many attempts

requires_payment_methodrequires_confirmationrequires_actione.g. 3D Secureattach PMconfirmprocessingasync methodsrequires_capturemanual captureauthorizedhandledsucceededfulfil the ordercanceledterminalsettledcapturecanceldeclined: retryA failed attempt does not create a failed state: the PaymentIntent returns to requires_payment_method
PaymentIntent statuses as documented by Stripe. Declines loop back to requires_payment_method, so one PaymentIntent tracks every attempt to collect one amount.

A single synchronous charge call breaks once payments need customer interaction mid-flight, such as 3D Secure, or settle over days, as bank debits do. The PaymentIntent is a state machine representing the intent to collect one amount, recording every attempt.

The documented statuses are requires_payment_method, requires_confirmation, requires_action, processing, requires_capture, succeeded and canceled. The design choice worth copying is that there is no failed status. When a card is declined, the PaymentIntent returns to requires_payment_method, so the customer can try another card against the same object, and you never create a second payment for the same order by accident. succeeded is the only signal that funds are in your account and you can fulfil. requires_capture appears when you authorize now and capture later, as a hotel or marketplace might. You can cancel before processing or succeeded, and for some bank debit methods during processing.

Stripe recommends creating the PaymentIntent as soon as you know the amount, so every attempt is recorded. A typical server-side creation looks like this (Python SDK):

import os
import stripe

stripe.api_key = os.environ["STRIPE_SECRET_KEY"]

def start_checkout(order):
    # One key per business operation, stored with the order BEFORE the call,
    # so every retry of this operation reuses it.
    if order.pi_idempotency_key is None:
        order.pi_idempotency_key = f"order-{order.id}-pi-v1"
        db.save(order)

    intent = stripe.PaymentIntent.create(
        amount=order.total_minor_units,          # 4200 = $42.00
        currency="usd",
        metadata={"order_id": str(order.id)},
        idempotency_key=order.pi_idempotency_key,
    )
    order.payment_intent_id = intent.id
    db.save(order)
    return intent.client_secret                  # to the browser, never the secret key

The idempotency key is derived from the business operation and stored before the call, so a crash cannot produce a second PaymentIntent on retry. The amount is in integer minor units, never floats. (Stripe now recommends the higher-level Checkout Sessions API for most new integrations; the lifecycle above still applies underneath.)

Advertisement

Idempotency as an API primitive

Networks fail in the worst place: after the server has charged the card and before the client sees the response. The client cannot tell whether to retry. Stripe's answer is the Idempotency-Key header, and its documented semantics are precise. The server saves the status code and body of the first request made with a key, regardless of whether it succeeded or failed, and replays that result to every retry with the same key, including a 500 error. Keys can be up to 255 characters, should be random or derived from your own operation ID, and should not contain personal data. Keys may be pruned once they are at least 24 hours old, after which reusing one creates a new request. All POST requests accept keys; GET and DELETE are idempotent already.

Two edge cases are deliberate. Reusing a key with different parameters returns an error instead of replaying, which catches keys reused for the wrong operation. And results are saved only once an endpoint begins executing, so a request that fails validation or collides with a concurrent one using the same key is not stored and can be retried. Building such a layer yourself is covered in idempotency architecture.

Replaying a stored 500 surprises people, but a 500 may follow side effects such as a network authorization, so re-executing could double them. To recover, GET the object's state, then decide whether to start a new operation with a new key.

Events and webhooks: at-least-once, unordered, signed

Many outcomes arrive after the API call: a bank debit clears days later, a customer finishes 3D Secure in another tab, a dispute opens weeks later. Stripe records each change as an Event and pushes it to your endpoints under a documented contract:

  • At-least-once. The same event can arrive more than once; store processed event IDs.
  • Unordered. Order is not guaranteed and created has one-second resolution; re-fetch the object for current state.
  • Retried. Live mode retries failures for up to three days with exponential backoff; non-2xx, redirects and timeouts all count, so return 200 fast.
  • Signed. The Stripe-Signature header carries t= and v1= values: HMAC-SHA256 over timestamp, a dot and the raw body, keyed by the whsec_ secret.
  • Versioned. Payloads follow the account API version when the event was created.

A receiver following the manual verification steps (official libraries do the same, with a five-minute default tolerance):

import hmac, hashlib, time

def verify(raw_body: bytes, header: str, secret: str, tolerance=300):
    parts = [kv.split("=", 1) for kv in header.split(",")]
    ts = int(next(v for k, v in parts if k == "t"))
    sigs = [v for k, v in parts if k == "v1"]            # ignore v0 and unknown schemes
    signed = f"{ts}.".encode() + raw_body                  # the RAW bytes, not re-serialised JSON
    expected = hmac.new(secret.encode(), signed, hashlib.sha256).hexdigest()
    if not any(hmac.compare_digest(expected, s) for s in sigs):
        raise ValueError("bad signature")
    if abs(time.time() - ts) > tolerance:
        raise ValueError("stale timestamp")                # replay protection

@app.post("/stripe/webhook")
def webhook():
    verify(request.get_data(), request.headers["Stripe-Signature"], WHSEC)
    event = json.loads(request.get_data())
    inbox.insert_ignore(event_id=event["id"], payload=event)   # dedupe by event id
    return "", 200                                             # ack fast, work async

Verify the raw body: frameworks that re-serialise JSON break the signature. A rolled secret can stay valid for up to 24 hours, with one signature per active secret, hence accepting any matching v1. The inbox turns at-least-once delivery into effectively-once processing, and a worker handles each event:

def process(event):
    pi_id = event["data"]["object"]["id"]
    pi = stripe.PaymentIntent.retrieve(pi_id)      # re-read: events can arrive out of order
    with db.transaction():
        order = Order.lock_by_payment_intent(pi_id)
        if pi.status == "succeeded" and order.state != "paid":
            order.mark_paid(amount=pi.amount_received)
            outbox.add("order.paid", order.id)      # fulfilment via an outbox, same transaction

Writing the fulfilment message in the same transaction as the order update is the transactional outbox; without it a crash loses or duplicates fulfilment.

API versioning: never break an integration

Integrations written years ago still run because versions are dates. Each account is pinned to the version current at its first request; a request can override it with the Stripe-Version header, and webhook endpoints can choose a version for their payloads. Breaking changes ship only as new versions.

Stripe has written that, internally, current code produces the newest representation and a chain of small version-change modules transforms responses backwards, one version at a time, to the caller's pinned version. A breaking change costs one self-contained transform, not a fork. Copy that for your own APIs: make the version an explicit rendering input and test each transform in isolation.

As an integrator, pin the version in configuration, read the changelog, test with the per-request header in a sandbox, then move the account default. Old events keep their old shape forever.

Inside the platform: Ledger and double-entry money movement

A payment triggers a chain of internal movements: funds at a network, fees, a merchant balance, a payout, perhaps a refund. Stripe has described Ledger, its immutable, auditable system of record for this, based on double-entry bookkeeping so all money is accounted for. Stripe reports that Ledger sees five billion events a day, and that 99.99% of its dollar volume is fully ingested and verified within four days.

Ledger detects errors by modelling producers as fund flows in which money moves between discrete states, such as "charge submitted" and "funds received". When a flow completes, the intermediate balances return to zero; Stripe calls this clearing and measures the fraction of the ledger zeroed out at steady state. A balance that never clears points at a missing settlement file, a producer bug or a partner error. Its data-quality checks ask: did the flow clear, did data arrive on time, and is it complete? Apply the same idea at small scale: model flows as accounts that should net to zero, and alert on those that do not.

Inside the platform: DocDB and moving data without downtime

Stripe has also written about DocDB, its database-as-a-service built as an extension of MongoDB Community plus in-house services, which it reports serves more than five million queries per second from product applications with five nines of uptime. Applications do not connect to shards directly. They go through a proxy layer, and a chunk metadata service maps ranges of data (chunks) to the shards that hold them.

That indirection makes the Data Movement Platform possible. It moves chunks between shards to split hot shards, merge underused ones and upgrade engine versions, without application changes. A migration registers the target, bulk-imports a snapshot, replicates ongoing writes asynchronously until the target catches up, checks correctness, and then performs a short, versioned traffic switch: the routing metadata moves to the target, and the source rejects requests carrying the old routing version, so a stale proxy cannot write to the wrong place. This is the same shape as most online migrations: copy, catch up, verify, then switch atomically, with a fence that makes the switch safe.

Worked example: one $42 order

Follow a customer buying a $42.00 item, with two things going wrong along the way.

  1. The server creates a PaymentIntent for 4200 cents with key order-981-pi-v1. The response is lost to a timeout; the retry with the same key replays the same pi_... ID.
  2. The browser confirms with the client secret; the bank demands 3D Secure, so the status becomes requires_action until the customer completes it.
  3. The card is declined and the PaymentIntent returns to requires_payment_method. A second card on the same page is authorized and, with automatic capture, the status becomes succeeded.
  4. The success event arrives first, then again after a deploy briefly returned 502s; the inbox ignores the duplicate. The late payment_intent.payment_failed is processed, but the worker re-fetches, sees succeeded and leaves the order paid.
  5. The redirect races the webhook, so the success page shows "confirming" until the worker marks the order paid. It never fulfils from the redirect.

Each step has one guard: the idempotency key, the state machine, event-ID dedupe, re-fetching, and the webhook (not the redirect) as source of truth. The multi-step version, with holds, captures and payouts, is a saga.

Failure modes and trade-offs

FailureWhat goes wrongDefence
Random key per retryEach retry creates a new PaymentIntentDerive the key from the operation and store it first
Fulfil on redirectCustomer closes tab and is never fulfilled, or a forged return URL fulfilsFulfil only on a verified webhook or a server-side status check
Parse body before verifyingEvery signature failsVerify the raw bytes, then parse
Slow webhook handlerTimeouts trigger retries and floodsInsert into an inbox, return 200, work async
Assume event orderA late failure event overwrites successRe-fetch the object; make transitions monotonic

The trade-offs are deliberate: at-least-once, unordered delivery is cheaper and more available than ordered exactly-once and pushes dedupe to receivers, and pinned versions mean Stripe carries compatibility code so merchants do not. Compare the bank-to-bank design of UPI.

What to do next

  1. Audit every Stripe POST in your code: each must send an idempotency key derived from a stored business operation ID.
  2. Make payment_intent.succeeded (or a server-side retrieve) the only trigger for fulfilment, and make the redirect page read state, not set it.
  3. Put webhook handling behind signature verification on the raw body, an inbox table keyed by event ID, and an async worker that re-fetches objects.
  4. Use an outbox for side effects that follow a payment state change.
  5. Pin your API version in configuration, and add a sandbox test that runs with the next version's Stripe-Version header.
  6. Store all amounts as integer minor units with an explicit currency, and reconcile daily against Stripe's balance transactions using the clearing idea: flows should net to zero.
  7. Run stripe listen and stripe trigger locally to rehearse duplicate, delayed and failed events before production does it for you.
Key takeaway: Stripe's architecture is a contract built to survive unreliable networks. PaymentIntents give one object per payment that loops on failure instead of forking. Idempotency keys replay the first outcome, even a 500. Webhooks are at-least-once, unordered, signed and versioned, so receivers verify raw bytes, dedupe by event ID and re-fetch state. Internally, Stripe has described a double-entry Ledger that checks money flows clear to zero, and a proxy-routed DocDB that moves data between shards without downtime. Copy these patterns in your own payment systems.