A food delivery order looks simple from the phone: pick a restaurant, pay, wait, eat. Behind it are three customers with conflicting interests. The consumer wants food fast and hot, the merchant wants orders paced to its kitchen, and the Dasher (DoorDash's name for a courier) wants well-paid trips with little waiting. The platform has to take payment exactly once, tell the kitchen at the right moment, pick a Dasher whose arrival matches the food being ready, and keep all three parties informed while anything can fail.

This article separates two things that are often blurred. First, what DoorDash has actually published on its engineering blog about its architecture: a move from a Python monolith to Kotlin gRPC services, a checkout flow rebuilt on Cassandra, Kafka, Cadence and gRPC, Cadence used as a safety net for event-driven delivery creation, and a dispatch system called DeepRed. Second, how you would build each piece yourself, with code, failure modes and a worked dinner-rush example. Anything DoorDash has not published, such as its volumes or latencies, is left out rather than guessed.

The moving parts

Consumer appbrowse, pay, trackMerchant tabletaccept, mark readyDasher appoffers, locationAPI gatewaygRPC to servicesCheckout serviceKotlin, idempotentOrder serviceorder state machineDelivery servicedelivery lifecycleDispatch (DeepRed)ML plus optimisationPaymentsauthorise, captureKafkaorder and delivery eventsCadencedurable workflowsassignPrediction servicesprep time, travel time, acceptanceIllustrative layout built from the components DoorDash has written about; service boundaries are an assumption.
Three apps, a gateway, core Kotlin services, Kafka for events, Cadence for durable workflows and prediction services feeding dispatch.

What DoorDash has published

What DoorDash has published. These are the claims this article relies on, all from posts on DoorDash's engineering blog.

  • From monolith to services. DoorDash started as a Python 2 / Django monolith and moved to a service-oriented architecture of gRPC microservices written in Kotlin. Its engineers wrote that they compared Kotlin, Java, Go, Rust and Python 3 before choosing Kotlin.
  • Checkout first. In 2020 the consumer checkout flow was extracted from the monolith into a Kotlin service using Apache Cassandra, Apache Kafka, Cadence and gRPC, with the stated aim of better performance, reliability and scalability.
  • Workflows as a fallback. A post from 21 October 2022, Building Reliable Workflows: Cadence as a Fallback for Event-Driven Processing, describes how the Drive business (white-label delivery for other merchants), after moving to a Kotlin service, made its delivery-creation flow resilient by running Cadence alongside the event-driven path.
  • DeepRed. Dispatch is handled by a system called DeepRed. Next-Generation Optimization for Dasher Dispatch at DoorDash (28 February 2020) covers matching Dashers to deliveries; Using ML and Optimization to Solve DoorDash's Dispatch Problem (17 August 2021) covers combining ML predictions with optimisation, simulation and experiments; Scaling a routing algorithm using multithreading and ruin-and-recreate (30 November 2021) covers the routing algorithm.

Everything after this section is how you would build it: a design consistent with those choices, not a description of DoorDash's internal code.

The order and the delivery are two state machines

Start with the order as a state machine, because every service hangs off it. A workable set of states is CREATED -> PAYMENT_AUTHORISED -> SENT_TO_MERCHANT -> CONFIRMED -> DASHER_ASSIGNED -> PICKED_UP -> DELIVERED, with CANCELLED and REFUNDED reachable from most of them. The delivery has its own lifecycle (unassigned, offered, assigned, at store, picked up, at customer, dropped off) because one order can have several delivery attempts: a Dasher can unassign, and a new one is found.

Keep the two apart. The order service owns money and merchant communication; the delivery service owns the physical job. Each publishes events when its state changes. Every transition is a conditional write (move from X to Y only if the current state is X), which turns duplicate or late events into no-ops instead of corruption.

The timing problem is what makes food different from rides. A ride starts when the driver arrives. Food has a second clock: the kitchen. Send the Dasher too early and they wait at the counter unpaid; too late and the food sits under a heat lamp. So dispatch needs a prep-time prediction per order, and the merchant's confirm and ready signals feed back into it.

Checkout: charge exactly once

Checkout is the right first extraction for the same reason it is risky: it is where money moves. The core requirement is that a retried request never charges twice. The client generates an idempotency key once per checkout attempt and sends it on every retry. The service stores the key with the result before returning, so a retry gets the original answer.

// Illustrative Kotlin, not DoorDash code. One row per idempotency key in a
// table keyed by key; insertIfAbsent is a lightweight transaction (IF NOT EXISTS)
// on Cassandra, or INSERT ... ON CONFLICT DO NOTHING on PostgreSQL.
suspend fun checkout(req: CheckoutRequest): CheckoutResponse {
    val claimed = idempotency.insertIfAbsent(req.idempotencyKey, status = "IN_PROGRESS")
    if (!claimed) return idempotency.awaitResult(req.idempotencyKey)   // retry: same answer

    val quote = pricing.quote(req.cart, req.address)          // fees, tax, promos
    val order = orders.create(req.consumerId, req.cart, quote, state = "CREATED")
    val auth = payments.authorise(order.id, quote.total, key = req.idempotencyKey)
    if (!auth.ok) {
        orders.transition(order.id, from = "CREATED", to = "CANCELLED")
        return idempotency.complete(req.idempotencyKey, CheckoutResponse.declined(auth.reason))
    }
    orders.transition(order.id, from = "CREATED", to = "PAYMENT_AUTHORISED")
    events.publish("order.authorised", order.id)   // via an outbox, not a bare send
    return idempotency.complete(req.idempotencyKey, CheckoutResponse.ok(order.id))
}

Three details matter. Authorise, do not capture, at checkout: the final amount can change (substitutions, tips, a cancelled item), so capture after delivery. Pass the idempotency key on to the payment provider so its side of the call is also safe to retry. And never publish the event with a separate network call after the database write: if the process dies between the two, the order exists and nobody hears about it. Write the event in the same transaction and relay it, as described in the transactional outbox pattern.

Events for speed, workflows for certainty

Once checkout succeeds, the rest of the pipeline is naturally event-driven: authorised order, send to merchant; merchant confirms, create delivery; delivery created, ask dispatch. Kafka fits this, with events keyed by order ID so all events for one order land in the same partition and stay ordered; Kafka partitioning covers why the key choice matters.

Pure event chains have a weakness DoorDash's 2022 post addresses: a step can silently not happen. A consumer crashes after reading but before acting, a message is dropped by a bug, a downstream call fails and the error is logged and swallowed. Nothing is waiting for the delivery, so nothing notices. A durable workflow engine such as Cadence fixes that by holding a timer for each order: if the expected state has not been reached by a deadline, the workflow performs the step itself.

// Illustrative Cadence-style workflow, started at checkout for every order.
// The fast path is still Kafka; the workflow only acts if the fast path stalled.
class DeliveryCreationWorkflowImpl : DeliveryCreationWorkflow {
    private val acts = Workflow.newActivityStub(DeliveryActivities::class.java, retryOptions)

    override fun ensureDelivery(orderId: String) {
        Workflow.sleep(Duration.ofSeconds(30))              // durable timer, survives restarts
        if (acts.deliveryExists(orderId)) return              // event path did its job
        acts.createDeliveryIdempotently(orderId)              // same idempotent call the consumer makes
        acts.emitMetric("delivery_created_by_fallback", orderId)
    }
}

The fallback only works if both paths call the same idempotent operation, because sometimes both will run. Count how often the fallback fires: a rising count is an early alarm that the event path is degrading. The trade-off is cost, since every order now has a workflow, timers and history in the workflow store.

Dispatch: predictions plus optimisation

Dispatch decides which Dasher takes which delivery and when they should head to the store. DoorDash's 2021 post frames it as ML plus optimisation: models predict the quantities the decision depends on, and an optimiser chooses assignments over all open deliveries and available Dashers at once rather than greedily one order at a time. The mechanics of batched matching and offer loops are covered in designing Uber-style dispatch, and kitchen timing and order batching in the Grab architecture article; here is how the food-specific pieces fit.

PredictionWhy dispatch needs itMain signals
Food ready timeDasher should arrive just as food is readymerchant history, order size, time of day, live queue
Travel time to store and customercost of each candidate pairingroad network, traffic, vehicle type, parking
Offer acceptancea likely decline wastes minutespay, distance, Dasher history
Time at store and at doorstacked orders compound delaysstore layout, apartment versus house

The routing post describes the ruin-and-recreate principle for building routes when one Dasher carries several orders: take a good solution, remove part of it, re-insert the removed stops in the best positions, keep the result if it scores better, and repeat. It is simple, parallelises well across threads, and handles constraints such as pickup before drop-off and food-ready times naturally because every insertion is checked against them.

import random

def ruin_and_recreate(routes, cost, feasible, iterations=2000, ruin_fraction=0.2):
    """routes: dict dasher -> ordered list of stops (pickups and drop-offs)."""
    best = current = routes
    for _ in range(iterations):
        removed, partial = ruin(current, ruin_fraction)       # pull out ~20% of orders, both stops each
        candidate = recreate(partial, removed, cost, feasible) # cheapest feasible insertion, one order at a time
        if candidate and cost(candidate) < cost(current):      # or accept slightly worse ones early, to escape local minima
            current = candidate
            if cost(current) < cost(best):
                best = current
    return best

Worked example: one Friday dinner order

Worked example, with illustrative numbers. It is 19:05 on a Friday. A consumer orders two pizzas. Checkout authorises the card with key k-7f3; the app's first request times out on a weak connection and retries, and the retry returns the same order ID because the key row already exists. The authorised event reaches the merchant integration, the tablet rings, and the merchant confirms at 19:06. The ready-time model predicts 19:24 for this store at this hour.

The delivery is created from the confirm event. The workflow's 30-second timer finds it present and exits. Dispatch runs every few seconds over the zone. A Dasher 6 minutes from the store is free now, so assigning them immediately means 12 minutes waiting at the counter. A second Dasher, 9 minutes away and finishing a drop-off at 19:13, arrives at 19:22, close to the ready time. The optimiser prefers the second, and also considers adding a nearby order ready at 19:26 to that route if that delays the pizzas by under a set limit.

At 19:15 the merchant marks an item unavailable. The order service recalculates the total and the capture amount drops; nothing needs refunding because nothing was captured. At 19:41 the Dasher marks delivered, the order captures the final amount, and payout and rating events go out.

Failure modes

  • Double charge. A retry without an idempotency key creates a second order. Fix: client-generated keys and key-scoped payment calls.
  • Lost step. An event is consumed but never acted on and the order sits in CONFIRMED with no delivery. Fix: a workflow or a sweeper that finds orders stuck past a deadline.
  • Both paths fire. The fallback and the consumer both create a delivery and two Dashers arrive. Fix: one idempotent create keyed by order ID, enforced by a unique constraint.
  • Prep-time model drift. A merchant changes staffing and food is ready 10 minutes later than predicted; Dashers wait and decline. Fix: per-merchant error monitoring and fast feedback from ready signals.
  • Zone-wide stall. The optimiser for a dense zone exceeds its time budget at peak. Fix: a hard time limit that returns the best solution found so far, and a greedy fallback.
  • Partial migration. During monolith extraction, the old and new paths both write order state. Fix: one owner per field at any moment, switched by a flag, with shadow reads to compare.

Trade-offs

Extracting services bought DoorDash independent scaling and deploys, at the price of distributed failure modes that a single database transaction used to hide: idempotency keys, outboxes and workflow fallbacks are the bill. A durable workflow engine adds a strong guarantee that every order reaches a terminal state, but it is another stateful system to run. Global optimisation in dispatch beats greedy matching on wait times and Dasher utilisation but needs good predictions; with poor ones, it optimises the wrong objective confidently. If you are smaller, a modular monolith with an outbox and a sweeper job gets you most of the reliability; see microservices design before splitting.

What to do next

  1. Draw your order and delivery state machines separately and make every transition a conditional write.
  2. Add client-generated idempotency keys to checkout and pass them through to the payment provider.
  3. Publish order events through an outbox, keyed by order ID.
  4. Add a deadline check (workflow timer or sweeper) for every step that must happen, and alert on how often it fires.
  5. Log predicted versus actual food-ready time per merchant before building any dispatch optimiser.
  6. Start dispatch with a greedy baseline, then try batched optimisation in simulation and an A/B test.
Key takeaway: DoorDash has published a move from a Python Django monolith to Kotlin gRPC services, a checkout service built on Cassandra, Kafka, Cadence and gRPC, Cadence used as a fallback so event-driven delivery creation cannot silently stall, and DeepRed, which combines ML predictions with optimisation and ruin-and-recreate routing. To build something similar, keep orders and deliveries as separate state machines with conditional transitions, make checkout idempotent end to end, publish events through an outbox, back every required step with a deadline check, and invest in food-ready predictions before optimising dispatch.