Stripe is often described as a payments company, but architecturally it is an API company. Its product is a contract: objects, state machines and delivery guarantees that let merchants move money without building connections to card networks and banks. Understanding it means understanding two layers: the public contract you integrate against (PaymentIntents, idempotency keys, webhooks, API versions), and the internal machinery Stripe has written about publicly, such as its Ledger and its DocDB database platform.
This article explains each piece of the contract from first principles, shows the code an integrator should write, and walks one payment from click to fulfilment, including a timeout and a duplicate webhook. For a vendor-neutral payment design with your own processor connections, read Designing a Payment System alongside this page.
Two views of one platform
Your server holds a secret key and creates objects. The browser or app holds only a publishable key and a per-payment client secret, and Stripe's client libraries collect card details and run 3D Secure, so card data never touches your servers and most of your stack stays out of the strictest PCI DSS scope.
Every API request passes through the same front door, which authenticates the key, picks the API version and applies idempotency. Payment objects then talk to networks and banks, every state change emits an event pushed to your webhooks, and each movement of funds is recorded as balanced ledger entries.
The PaymentIntent: one object per payment, many attempts
A single synchronous charge call breaks once payments need customer interaction mid-flight, such as 3D Secure, or settle over days, as bank debits do. The PaymentIntent is a state machine representing the intent to collect one amount, recording every attempt.
The documented statuses are requires_payment_method, requires_confirmation, requires_action, processing, requires_capture, succeeded and canceled. The design choice worth copying is that there is no failed status. When a card is declined, the PaymentIntent returns to requires_payment_method, so the customer can try another card against the same object, and you never create a second payment for the same order by accident. succeeded is the only signal that funds are in your account and you can fulfil. requires_capture appears when you authorize now and capture later, as a hotel or marketplace might. You can cancel before processing or succeeded, and for some bank debit methods during processing.
Stripe recommends creating the PaymentIntent as soon as you know the amount, so every attempt is recorded. A typical server-side creation looks like this (Python SDK):
import os
import stripe
stripe.api_key = os.environ["STRIPE_SECRET_KEY"]
def start_checkout(order):
# One key per business operation, stored with the order BEFORE the call,
# so every retry of this operation reuses it.
if order.pi_idempotency_key is None:
order.pi_idempotency_key = f"order-{order.id}-pi-v1"
db.save(order)
intent = stripe.PaymentIntent.create(
amount=order.total_minor_units, # 4200 = $42.00
currency="usd",
metadata={"order_id": str(order.id)},
idempotency_key=order.pi_idempotency_key,
)
order.payment_intent_id = intent.id
db.save(order)
return intent.client_secret # to the browser, never the secret keyThe idempotency key is derived from the business operation and stored before the call, so a crash cannot produce a second PaymentIntent on retry. The amount is in integer minor units, never floats. (Stripe now recommends the higher-level Checkout Sessions API for most new integrations; the lifecycle above still applies underneath.)
Idempotency as an API primitive
Networks fail in the worst place: after the server has charged the card and before the client sees the response. The client cannot tell whether to retry. Stripe's answer is the Idempotency-Key header, and its documented semantics are precise. The server saves the status code and body of the first request made with a key, regardless of whether it succeeded or failed, and replays that result to every retry with the same key, including a 500 error. Keys can be up to 255 characters, should be random or derived from your own operation ID, and should not contain personal data. Keys may be pruned once they are at least 24 hours old, after which reusing one creates a new request. All POST requests accept keys; GET and DELETE are idempotent already.
Two edge cases are deliberate. Reusing a key with different parameters returns an error instead of replaying, which catches keys reused for the wrong operation. And results are saved only once an endpoint begins executing, so a request that fails validation or collides with a concurrent one using the same key is not stored and can be retried. Building such a layer yourself is covered in idempotency architecture.
Replaying a stored 500 surprises people, but a 500 may follow side effects such as a network authorization, so re-executing could double them. To recover, GET the object's state, then decide whether to start a new operation with a new key.
Events and webhooks: at-least-once, unordered, signed
Many outcomes arrive after the API call: a bank debit clears days later, a customer finishes 3D Secure in another tab, a dispute opens weeks later. Stripe records each change as an Event and pushes it to your endpoints under a documented contract:
- At-least-once. The same event can arrive more than once; store processed event IDs.
- Unordered. Order is not guaranteed and
createdhas one-second resolution; re-fetch the object for current state. - Retried. Live mode retries failures for up to three days with exponential backoff; non-2xx, redirects and timeouts all count, so return 200 fast.
- Signed. The
Stripe-Signatureheader carriest=andv1=values: HMAC-SHA256 over timestamp, a dot and the raw body, keyed by thewhsec_secret. - Versioned. Payloads follow the account API version when the event was created.
A receiver following the manual verification steps (official libraries do the same, with a five-minute default tolerance):
import hmac, hashlib, time
def verify(raw_body: bytes, header: str, secret: str, tolerance=300):
parts = [kv.split("=", 1) for kv in header.split(",")]
ts = int(next(v for k, v in parts if k == "t"))
sigs = [v for k, v in parts if k == "v1"] # ignore v0 and unknown schemes
signed = f"{ts}.".encode() + raw_body # the RAW bytes, not re-serialised JSON
expected = hmac.new(secret.encode(), signed, hashlib.sha256).hexdigest()
if not any(hmac.compare_digest(expected, s) for s in sigs):
raise ValueError("bad signature")
if abs(time.time() - ts) > tolerance:
raise ValueError("stale timestamp") # replay protection
@app.post("/stripe/webhook")
def webhook():
verify(request.get_data(), request.headers["Stripe-Signature"], WHSEC)
event = json.loads(request.get_data())
inbox.insert_ignore(event_id=event["id"], payload=event) # dedupe by event id
return "", 200 # ack fast, work asyncVerify the raw body: frameworks that re-serialise JSON break the signature. A rolled secret can stay valid for up to 24 hours, with one signature per active secret, hence accepting any matching v1. The inbox turns at-least-once delivery into effectively-once processing, and a worker handles each event:
def process(event):
pi_id = event["data"]["object"]["id"]
pi = stripe.PaymentIntent.retrieve(pi_id) # re-read: events can arrive out of order
with db.transaction():
order = Order.lock_by_payment_intent(pi_id)
if pi.status == "succeeded" and order.state != "paid":
order.mark_paid(amount=pi.amount_received)
outbox.add("order.paid", order.id) # fulfilment via an outbox, same transactionWriting the fulfilment message in the same transaction as the order update is the transactional outbox; without it a crash loses or duplicates fulfilment.
API versioning: never break an integration
Integrations written years ago still run because versions are dates. Each account is pinned to the version current at its first request; a request can override it with the Stripe-Version header, and webhook endpoints can choose a version for their payloads. Breaking changes ship only as new versions.
Stripe has written that, internally, current code produces the newest representation and a chain of small version-change modules transforms responses backwards, one version at a time, to the caller's pinned version. A breaking change costs one self-contained transform, not a fork. Copy that for your own APIs: make the version an explicit rendering input and test each transform in isolation.
As an integrator, pin the version in configuration, read the changelog, test with the per-request header in a sandbox, then move the account default. Old events keep their old shape forever.
Inside the platform: Ledger and double-entry money movement
A payment triggers a chain of internal movements: funds at a network, fees, a merchant balance, a payout, perhaps a refund. Stripe has described Ledger, its immutable, auditable system of record for this, based on double-entry bookkeeping so all money is accounted for. Stripe reports that Ledger sees five billion events a day, and that 99.99% of its dollar volume is fully ingested and verified within four days.
Ledger detects errors by modelling producers as fund flows in which money moves between discrete states, such as "charge submitted" and "funds received". When a flow completes, the intermediate balances return to zero; Stripe calls this clearing and measures the fraction of the ledger zeroed out at steady state. A balance that never clears points at a missing settlement file, a producer bug or a partner error. Its data-quality checks ask: did the flow clear, did data arrive on time, and is it complete? Apply the same idea at small scale: model flows as accounts that should net to zero, and alert on those that do not.
Inside the platform: DocDB and moving data without downtime
Stripe has also written about DocDB, its database-as-a-service built as an extension of MongoDB Community plus in-house services, which it reports serves more than five million queries per second from product applications with five nines of uptime. Applications do not connect to shards directly. They go through a proxy layer, and a chunk metadata service maps ranges of data (chunks) to the shards that hold them.
That indirection makes the Data Movement Platform possible. It moves chunks between shards to split hot shards, merge underused ones and upgrade engine versions, without application changes. A migration registers the target, bulk-imports a snapshot, replicates ongoing writes asynchronously until the target catches up, checks correctness, and then performs a short, versioned traffic switch: the routing metadata moves to the target, and the source rejects requests carrying the old routing version, so a stale proxy cannot write to the wrong place. This is the same shape as most online migrations: copy, catch up, verify, then switch atomically, with a fence that makes the switch safe.
Worked example: one $42 order
Follow a customer buying a $42.00 item, with two things going wrong along the way.
- The server creates a PaymentIntent for 4200 cents with key
order-981-pi-v1. The response is lost to a timeout; the retry with the same key replays the samepi_...ID. - The browser confirms with the client secret; the bank demands 3D Secure, so the status becomes
requires_actionuntil the customer completes it. - The card is declined and the PaymentIntent returns to
requires_payment_method. A second card on the same page is authorized and, with automatic capture, the status becomessucceeded. - The success event arrives first, then again after a deploy briefly returned 502s; the inbox ignores the duplicate. The late
payment_intent.payment_failedis processed, but the worker re-fetches, seessucceededand leaves the order paid. - The redirect races the webhook, so the success page shows "confirming" until the worker marks the order paid. It never fulfils from the redirect.
Each step has one guard: the idempotency key, the state machine, event-ID dedupe, re-fetching, and the webhook (not the redirect) as source of truth. The multi-step version, with holds, captures and payouts, is a saga.
Failure modes and trade-offs
| Failure | What goes wrong | Defence |
|---|---|---|
| Random key per retry | Each retry creates a new PaymentIntent | Derive the key from the operation and store it first |
| Fulfil on redirect | Customer closes tab and is never fulfilled, or a forged return URL fulfils | Fulfil only on a verified webhook or a server-side status check |
| Parse body before verifying | Every signature fails | Verify the raw bytes, then parse |
| Slow webhook handler | Timeouts trigger retries and floods | Insert into an inbox, return 200, work async |
| Assume event order | A late failure event overwrites success | Re-fetch the object; make transitions monotonic |
The trade-offs are deliberate: at-least-once, unordered delivery is cheaper and more available than ordered exactly-once and pushes dedupe to receivers, and pinned versions mean Stripe carries compatibility code so merchants do not. Compare the bank-to-bank design of UPI.
What to do next
- Audit every Stripe POST in your code: each must send an idempotency key derived from a stored business operation ID.
- Make
payment_intent.succeeded(or a server-side retrieve) the only trigger for fulfilment, and make the redirect page read state, not set it. - Put webhook handling behind signature verification on the raw body, an inbox table keyed by event ID, and an async worker that re-fetches objects.
- Use an outbox for side effects that follow a payment state change.
- Pin your API version in configuration, and add a sandbox test that runs with the next version's
Stripe-Versionheader. - Store all amounts as integer minor units with an explicit currency, and reconcile daily against Stripe's balance transactions using the clearing idea: flows should net to zero.
- Run
stripe listenandstripe triggerlocally to rehearse duplicate, delayed and failed events before production does it for you.