Robinhood made stock trading a two-tap action on a phone. Behind the tap is a chain of systems that has to be faster than the user's patience, correct to the cent, and still standing when millions of people open the app at the same moment. This article designs that chain. It is a reference design for a retail broker in the spirit of Robinhood, not a leaked diagram of Robinhood's internals; where it cites what actually happened, those are publicly reported facts and are labelled as such.

One distinction shapes everything: Robinhood is a broker, not an exchange. It does not match buyers with sellers in its own order book; it accepts customers' orders, checks them, sends them to market makers and venues that execute them, and then records, clears and settles the result. If you want the matching side, that is a different system. Here we follow one order from the app to settlement, then look at market data, peak load and the two incidents that every engineer in this space studies.

Advertisement

What a retail broker actually has to do

Break the job into five responsibilities. Intake: authenticate the user, validate the order, and make sure a repeated tap does not become two orders. Pre-trade risk: check buying power, position limits, trading permissions (options, margin), restricted symbols and pattern-day-trader rules, and reserve the money. Routing and execution: send the order to an execution partner over FIX and process the execution reports that come back. Post-trade: update positions and a double-entry ledger, notify the user, and hand trades to clearing. Market data: show live prices to millions of screens.

These have very different shapes. Orders are a low-volume, must-not-lose, must-not-duplicate path. Market data is an enormous, lossy-by-design path where only the latest price matters. Mixing the two on the same infrastructure is the classic way a quote storm takes order entry down with it.

Reference design for a retail broker: the app never talks to a venue; every order passes risk and a ledger firstMobile appidempotency keyAPI gatewayauth, rate limitsOrder servicestate machineRisk and buying powerreserve fundscheckRouterFIX sessionsacceptedMarket makers, venuesexecuteNewOrderSingleExecutionReportEvent logpartitioned by accountfillsLedger and positionsdouble entryClearingNSCC, T+1Market data feedsquotes, tradesFan-outconflate per symbolwebsocketMarket data is a separate, lossy-by-design pathOrders are a lossless path with exactly-once effects
Orders flow left to right through risk, routing and the ledger; market data flows through a separate fan-out tier that is allowed to drop intermediate quotes.

Intake and idempotency

Mobile networks retry. Users double-tap. Load balancers resend after a timeout. The app therefore generates a client key, a UUID per order intent, and sends it with the request. The order service stores it under a unique constraint per account, so the second arrival of the same key returns the existing order instead of creating a new one. This is the same idea as idempotency keys in payments, and it is non-negotiable: a duplicated buy is real money.

Validation happens before anything durable: symbol exists and is tradable, quantity or notional is positive, order type is allowed for this account, the market is open or the order is eligible for extended hours. Rejecting early is cheap; rejecting after reserving funds needs a compensating release.

Advertisement

Buying power: reserve, then route

Pre-trade risk is a race. A user with 1,000 dollars can submit two 800-dollar buys a millisecond apart, and both reads of the balance will say yes. The fix is to treat buying power as a reservation in the same transaction that creates the order, with a conditional update that only succeeds if enough is available. The transaction also writes an outbox row, so the message that triggers routing is published if and only if the order was accepted.

# One transaction: create the order, reserve funds, write the outbox row.
# The unique key on (account_id, client_key) makes a retried tap a no-op.
def place_buy(tx, o):
    new = tx.execute(
        "INSERT INTO orders (id, account_id, client_key, symbol, notional, state) "
        "VALUES (%s, %s, %s, %s, %s, 'new') "
        "ON CONFLICT (account_id, client_key) DO NOTHING RETURNING id",
        (o.id, o.account_id, o.client_key, o.symbol, o.notional)).fetchone()
    if new is None:                                   # duplicate tap
        return tx.execute("SELECT * FROM orders WHERE account_id = %s AND client_key = %s",
                          (o.account_id, o.client_key)).fetchone()
    reserved = tx.execute(
        "UPDATE buying_power SET available = available - %s, reserved = reserved + %s "
        "WHERE account_id = %s AND available >= %s",
        (o.notional, o.notional, o.account_id, o.notional)).rowcount
    if reserved == 0:                                 # insufficient funds
        tx.execute("UPDATE orders SET state = 'rejected' WHERE id = %s", (o.id,))
        return "rejected"
    tx.execute("INSERT INTO outbox (topic, key, payload) VALUES ('orders.accepted', %s, %s)",
               (o.account_id, o.to_json()))
    tx.execute("UPDATE orders SET state = 'accepted' WHERE id = %s", (o.id,))
    return "accepted"

The outbox pattern matters here because the alternative, committing the order and then publishing to Kafka, has a gap: a crash between the two leaves money reserved for an order nobody will route. A relay reads the outbox and publishes to the event log, keyed by account id so every event for one account lands on the same partition and is processed in order. When the order ends, the reservation is converted into a ledger entry for the filled amount and the remainder is released.

The order state machine

An order is a state machine whose transitions are driven by two sources: the user (cancel) and the execution partner (fills, rejects, cancel confirmations). Writing the allowed transitions down as data prevents a whole class of bugs, and handling duplicates by execution id makes replays safe.

from enum import Enum

class S(Enum):
    NEW = "new"; ACCEPTED = "accepted"; ROUTED = "routed"
    PARTIAL = "partially_filled"; FILLED = "filled"
    CANCEL_PENDING = "cancel_pending"; CANCELLED = "cancelled"; REJECTED = "rejected"

ALLOWED = {
    S.NEW: {S.ACCEPTED, S.REJECTED},
    S.ACCEPTED: {S.ROUTED, S.CANCELLED, S.REJECTED},
    S.ROUTED: {S.PARTIAL, S.FILLED, S.CANCEL_PENDING, S.REJECTED},
    S.PARTIAL: {S.PARTIAL, S.FILLED, S.CANCEL_PENDING},
    S.CANCEL_PENDING: {S.CANCELLED, S.PARTIAL, S.FILLED},  # a fill can beat the cancel
    S.FILLED: set(), S.CANCELLED: set(), S.REJECTED: set(),
}

def transition(order, new_state, exec_id=None):
    if exec_id is not None and exec_id in order.seen_exec_ids:
        return order                      # duplicate ExecutionReport: ignore
    if new_state not in ALLOWED[order.state]:
        raise ValueError(f"{order.id}: {order.state} -> {new_state} not allowed")
    order.state = new_state
    if exec_id is not None:
        order.seen_exec_ids.add(exec_id)
    return order

Note the edge from cancel-pending to filled. A cancel is a request, not a command; the venue may already have executed the order. The UI must say cancel requested until the partner confirms, and the ledger must accept a fill that arrives after the user pressed cancel. Treating cancel as immediate is one of the most common bugs in trading front ends.

Routing and execution over FIX

Brokers talk to execution partners using the FIX protocol: long-lived sessions with sequence numbers, a NewOrderSingle message (type D) carrying a client order id, and ExecutionReport messages (type 8) carrying the order status, each fill's quantity and price, and an execution id. The router keeps several sessions per partner, chooses a destination by policy (price improvement history, fill rates, partner health), and persists every outbound and inbound message before acting on it. On reconnect, FIX resend requests recover any gap in sequence numbers, and the execution-id check above makes duplicates harmless.

Robinhood publicly routes most equity orders to market makers and receives payment for order flow, which it discloses in its regulatory routing reports. Architecturally that means a small number of high-volume partner connections, a routing policy with an audit trail for best execution reviews, and fallbacks if a partner degrades. Fractional-share orders need extra logic, because venues trade whole shares; one common design fills the fractional part against the broker's own inventory and routes whole-share remainders.

Post-trade: ledger, positions, clearing

Every fill becomes a double-entry ledger transaction: cash moves from the customer's account to the trading account, shares move the other way, and fees, if any, are separate lines. Positions are a projection of the ledger, not an independently updated number; the reconciliation job compares the projection with the clearing house's view every night, and any break is a ticket. The same discipline as in payment system design applies: immutable entries, corrections as new entries, never in-place edits.

US equities settle on T+1 since 28 May 2024: trades executed today settle on the next business day through the clearing house. Until then the clearing house carries the risk that a broker fails to pay, and it collects deposits from brokers sized by the risk of their unsettled positions, which rises with volatility and concentration. That deposit is the hinge of the most famous incident in this space.

Incident one: the clearing deposit, January 2021

As publicly reported, at about 3:30 a.m. Pacific time on 28 January 2021 Robinhood received a deposit demand from the National Securities Clearing Corporation of roughly 3 billion dollars, about an order of magnitude more than usual, driven by concentrated buying in a handful of volatile stocks. The requirement was negotiated down to 1.4 billion and later to about 700 million, and the firm restricted customers to closing positions in those symbols.

The engineering lessons are concrete. Clearing exposure is a production metric: compute an estimate of the deposit intraday from open positions and the clearing house's published methodology, and alert treasury before the clearing house does. Build per-symbol controls into the order path before you need them, such as close-only, raised margin and position caps, as feature flags that risk staff can flip without a deploy. And design the user experience of a restriction, because shipping it under pressure produced the worst possible version.

Market data fan-out

A quote feed for US equities produces a very large number of updates per second at the open, but a person can read a few per second at most. The fan-out tier normalises feeds, publishes them on a partitioned log or in-memory bus keyed by symbol, and maintains websocket connections to clients, each subscribed to the symbols on screen. Between bus and socket sits a conflator that keeps only the latest quote per symbol and flushes on a timer.

import asyncio

class Conflator:
    # Keep only the latest quote per symbol; slow clients get fewer, fresher updates.
    def __init__(self, send, interval=0.25):
        self.latest, self.dirty, self.send, self.interval = {}, set(), send, interval

    def on_quote(self, symbol, quote):
        self.latest[symbol] = quote
        self.dirty.add(symbol)

    async def run(self):
        while True:
            await asyncio.sleep(self.interval)
            batch, self.dirty = self.dirty, set()
            if batch:
                await self.send({s: self.latest[s] for s in batch})

Conflation turns a burst of a hundred updates for one symbol into one message per interval, which bounds per-client bandwidth no matter how hot the market gets. The price shown in the app is therefore indicative; the order path never trusts it, and market orders execute at whatever the venue provides. Robinhood's engineers also open-sourced Faust, a Python stream-processing library modelled on Kafka Streams, which they described as used for real-time pipelines such as risk and fraud checks; whatever a team chooses, the stream processors for analytics must sit off the order path.

Incident two: the opening bell, March 2020

On 2 March 2020, a day of record volatility and volume, Robinhood was down for the whole trading day. The company's public explanation was unprecedented load that caused a thundering herd effect, which in turn triggered a failure of its DNS system. The general pattern is worth internalising: a dependency every service touches but nobody load-tests, hit by synchronised retries, fails, and every service fails with it.

Defences: pre-scale before 9:30 Eastern on days with heavy pre-market activity; give clients exponential backoff with jitter so reconnects spread out; cache DNS answers in-process with a stale-if-error policy; shed load by priority, with cancels and order entry above charts and news; and put rate limiters at the gateway that degrade gracefully instead of failing closed. Run a game day that replays the open at several times last year's peak.

Worked example: one 500-dollar buy

  1. 09:30:00.050: the app sends a market buy for 500 dollars of XYZ with client key k1. The gateway authenticates and forwards.
  2. 09:30:00.060: the order service validates, inserts the order, reserves 500 dollars and writes the outbox row in one transaction. The response tells the app the order is accepted.
  3. 09:30:00.070: the relay publishes orders.accepted to the account's partition; the router converts the notional order to the partner's format and sends a NewOrderSingle.
  4. 09:30:00.140: an ExecutionReport arrives: filled, 2.731 shares at 183.08 dollars, execution id e9. The state machine moves to filled, the ledger posts the cash and share legs for 499.99 dollars, and the reserve's remaining cent is released.
  5. 09:30:00.200: a push notification goes out; the positions projection updates; the next trading day the trade settles and the nightly reconciliation matches the clearing record.
  6. If the app had retried at 09:30:00.300 with k1, it would have received the same order back, and no second reservation would exist.

Trade-offs

Synchronous risk checks add latency to every order but are the only safe place to stop an over-spend; asynchronous checks are acceptable only for slower controls such as fraud scoring that can freeze an account afterwards. Partitioning by account gives ordering per customer but makes per-symbol controls a separate lookup, so keep restricted-symbol state in a fast replicated cache. Routing to a few market makers keeps connectivity simple but concentrates dependency; a second destination for every order type is cheap insurance. For a general treatment of event-driven order flows see event-driven order design.

What to do next

  1. Write your order state machine as an explicit transition table and reject any transition not in it.
  2. Put a unique client key on every order request and test double-tap, retry and timeout paths end to end.
  3. Make buying-power reservation and order creation one transaction, and publish through an outbox.
  4. Persist every FIX message before acting on it and dedupe execution reports by execution id.
  5. Compute an intraday clearing-deposit estimate and alert on it; build close-only and position-cap flags per symbol.
  6. Separate market data from the order path physically, conflate per symbol, and never price an order from the client's displayed quote.
  7. Load-test shared dependencies such as DNS, auth and config at several times last year's opening peak, with jittered client retries.
  8. Reconcile ledger positions against the clearing house nightly and treat every break as an incident.
Key takeaway: A retail broker is two systems: a lossless order path that must never duplicate or lose money, and a lossy market-data path that only needs the latest price. Idempotent intake, reservation in the same transaction as the order, an explicit state machine, deduplicated execution reports and a double-entry ledger make the first correct; conflation and isolation make the second cheap. The 2020 outage and the 2021 clearing deposit show that shared dependencies and clearing exposure are architecture concerns, not operations footnotes.