Grab began as a taxi-hailing app in Malaysia and grew into a super app: ride hailing, food and grocery delivery, parcel delivery, a wallet, lending and insurance, inside one Android and iOS app, which Grab says is available in eight Southeast Asian countries. The interesting engineering question is not how to build a ride-hailing service, which the Uber dispatch design covers, but how to run many marketplaces that share drivers, users, maps and money without each one rebuilding them or breaking the others.
This article walks through that architecture layer by layer: shared platform services, location ingestion, allocation that deliberately does not pick the nearest driver, food dispatch timing and batching, the wallet ledger, the event backbone and the constraints of the region. Where Grab has published how it works, the article says so and dates the source. Where it has not, the design is presented as reasoning about the problem, with hypothetical numbers, so you can apply it to your own system.
The shape of a super app
A super app is a set of marketplaces that share four expensive things: a user base with one identity, a partner fleet that can switch between rides and deliveries, a map with live positions, and a wallet. Sharing them is the business advantage. A user who rides can be offered food; a driver idle between rides can carry a parcel; one balance pays for all of it.
That pushes the architecture into two layers. Vertical services own their own workflows: a ride booking has a different state machine from a food order, and teams need to ship them independently. Platform services own what is shared and must be generic enough that a new vertical can plug in without changing them. The hard design work sits at the boundary: a platform API that is too narrow forces verticals to fork it, and one that is too broad turns the platform team into the bottleneck for every launch.
Grab has published parts of its platform. It has described Grab-Kit, a framework for building its Go microservices that was inspired by Go-Kit, generates server and client scaffolding and standardises communication between services. Coban is its real-time data streaming platform, providing self-served Kafka topics, Flink and change-data-capture pipelines and Kafka Connect connectors. Catwalk serves its machine-learning models, and Grab has described migrating it onto NVIDIA Triton Inference Server. Trident, covered below, is its rules engine.
Location: the highest-volume write path
Every online driver and delivery partner sends position updates every few seconds. That stream is the input to allocation, ETAs, surge, fraud checks and the live map the customer watches, and it is the highest-volume write path in the system. Work it through with hypothetical numbers: 400,000 online partners reporting every 4 seconds is 100,000 updates per second, almost all of which overwrite a previous position.
So the latest position belongs in memory, keyed by partner and indexed spatially, while the history goes to the event stream for analytics and trip reconstruction. Grab has said it locates drivers and passengers with geohashes, which divide the map into a hierarchy of cells, each named by a base-32 string whose prefix is its parent cell. A 6-character geohash is a cell about 1.2 km by 0.6 km at the equator; 7 characters is about 153 m square. Finding nearby drivers means reading the rider's cell and its eight neighbours, then filtering by true distance:
# Write path: called for every location ping.
def on_ping(partner_id, lat, lon, ts, status):
cell = geohash.encode(lat, lon, precision=6)
old = positions.get(partner_id)
if old and old.cell != cell:
cells[old.cell].discard(partner_id) # moved to a new cell
cells[cell].add(partner_id)
positions[partner_id] = Pos(lat, lon, ts, cell, status)
events.publish("partner.location", partner_id, (lat, lon, ts)) # history, async
# Read path: candidates for one booking.
def nearby(lat, lon, radius_m, max_age_s=15):
home = geohash.encode(lat, lon, precision=6)
out = []
for cell in [home] + geohash.neighbors(home): # 3x3 block
for pid in cells[cell]:
p = positions[pid]
if now() - p.ts > max_age_s or p.status != "available":
continue # stale or busy
if haversine_m(lat, lon, p.lat, p.lon) <= radius_m:
out.append(pid)
return outShard the index by cell prefix so that a city lives on a few nodes and a query touches one or two. Busy cells in a city centre are much hotter than rural ones, so allow a hot shard to split by a longer prefix. Treat stale pings as absent: a phone in a tunnel should not receive a booking.
Allocation is not nearest-driver
The naive design sends a booking to the nearest available driver. Grab has written publicly that it does not do this. Its allocation weighs estimated time of arrival rather than straight-line distance, driver trip preferences, fraud prevention, vehicle suitability, regional traffic rules and order batching. A driver 300 m away across a river may have a worse ETA than one 900 m away on the same road, and a motorbike is the wrong vehicle for a large grocery order.
Two structural choices follow. First, allocation is a scoring and assignment problem, not a lookup. Second, it works better in small batches than one booking at a time: collect the bookings and available partners in an area over a short window, score every pair and solve the assignment for the whole window. That trades a couple of seconds of latency for better global matches. The details of batched matching are in the Uber dispatch design; the super-app twist is that rides, food and parcels compete for the same partners.
def score(booking, partner):
eta = eta_service.pickup_eta(partner.pos, booking.pickup) # road network, live traffic
if eta is None or not vehicle_fits(partner.vehicle, booking):
return None # infeasible pair
if fraud.risky_pair(partner.id, booking.user_id):
return None
s = -eta.seconds
s -= 120 * partner.recent_declines # likely to decline again
s += 60 if booking.vertical in partner.preferred_verticals else 0
return s
def allocate_window(bookings, partners):
pairs = {(b.id, p.id): score(b, p) for b in bookings for p in partners}
feasible = {k: v for k, v in pairs.items() if v is not None}
return max_weight_matching(feasible) # Hungarian or a greedy approximationThe weights are illustrative. In practice they come from experiments and models; the structure is the point. Keep the scoring function in one place shared by all verticals, so a change in how declines are penalised does not apply to rides and silently not to parcels.
Food: timing dispatch to the kitchen, and batching
Food delivery is a three-sided marketplace: customer, merchant and delivery partner. Its key difference from rides is that the item is not ready when the order arrives. Grab has described working out when a partner should head to the restaurant from the partner's travel time and the food's preparation time, so the partner arrives as the food is ready. Send the partner too early and they wait unpaid at the counter, blocking other orders; too late and the food goes cold.
The order is therefore a longer state machine, and dispatch is scheduled rather than immediate:
PLACED -> MERCHANT_ACCEPTED -> PREPARING -> PARTNER_ASSIGNED -> AT_MERCHANT
-> PICKED_UP -> AT_CUSTOMER -> DELIVERED
any state before PICKED_UP -> CANCELLED (refund rules depend on the state)
dispatch_at = ready_at_estimate - pickup_eta_estimate - safety_marginBatching is the second difference. When there are not enough partners for every order, Grab's system can give one partner two or more orders whose pickups or drop-offs are close, and its published factors include item type, weight and preparation time, not just proximity. Batching raises partner earnings per hour and fleet throughput at the cost of a later delivery for at least one customer. Grab's Saver option makes the trade explicit: a lower delivery fee for a longer delivery time, which gives the dispatcher room to batch. Bound every batch by the promised delivery time of each order in it, and recompute the plan when a merchant runs late.
Money: one wallet, many flows, one ledger
Payments are where a super app is least forgiving. One wallet pays for rides, food and in-store purchases, partners are paid out, and in markets where many riders pay cash a driver collects money that partly belongs to the platform. Every one of those flows must balance.
The standard answer is an append-only double-entry ledger: every movement is a set of entries that sum to zero across accounts, and balances are derived from entries. A cash ride, for example, records the fare as cash held by the driver, and the commission as an amount the driver owes the platform, which is then netted against their next card or wallet earnings. The ledger never updates a balance in place, so any balance can be rebuilt and audited.
def settle_cash_trip(trip):
txn = Txn(idempotency_key=f"trip:{trip.id}:settle")
txn.entry("driver:{d}:cash_held".format(d=trip.driver), +trip.fare)
txn.entry("customer:{c}:paid_cash".format(c=trip.customer), -trip.fare)
txn.entry("driver:{d}:payable".format(d=trip.driver), -trip.commission)
txn.entry("platform:commission", +trip.commission)
ledger.post(txn) # rejects if entries do not sum to zero; no-op if key seenTwo properties make this survive retries on flaky networks. Every posting carries an idempotency key, so a client or service that retries after a timeout cannot charge twice; idempotency keys covers the mechanics. Workflows that span services, such as charging a wallet and then confirming a food order, run as sagas with compensating steps rather than distributed transactions. If the merchant rejects the order, the compensation is a refund posting, not a deleted row. The broader design is in designing a payment system.
The event backbone and the rules engine
Shared services communicate through events far more than through synchronous calls. A completed trip publishes an event; the ledger, rewards, driver incentives, analytics and fraud systems each consume it on their own schedule. This keeps the booking path short: confirming a ride should not wait for loyalty points to be awarded.
Trident is Grab's in-house real-time if-this-then-that engine. Campaign authors configure which event triggers which action under which conditions: send a food promotion after a ride, award points after several bookings, compensate a customer for a late delivery. Grab's 2021 description put it at more than 2,000 events per second and tens of millions of actions per day. Centralising those rules keeps cross-vertical promotions out of the vertical services, which would otherwise each grow their own copy.
The trap with event-driven design is losing an event between the database write and the publish. Write the event to an outbox table in the same transaction as the state change and publish from it, as described in the outbox pattern. Consumers must be idempotent, because at-least-once delivery means duplicates will arrive.
Designing for the region
- Devices and networks. Many users have low-end Android phones on patchy mobile data. Keep payloads small, make every mutating call idempotent so the client can retry blindly, and let the server drive screens so a new feature does not need an app update.
- Cash and local payment methods. Cash plus many local wallets and bank rails means the payment service is an adapter layer over many providers, each with its own failure behaviour.
- Vehicles and roads. Motorbikes, cars and taxis travel differently, and addresses are often imprecise. Grab built its own map product, GrabMaps, and ETA quality depends on it.
- Many countries. Fares, regulations, tax and data rules differ by country, so per-country configuration must be data, not code branches.
Failure modes and trade-offs
| Failure | What happens | Mitigation |
|---|---|---|
| Location index node lost | Bookings in a city find no drivers | Replicas per shard; rebuild from the ping stream within seconds |
| Allocation slow at peak | Customers see long searching screens | Shrink the window, fall back to greedy matching, shed low-priority verticals first |
| Merchant prep estimate wrong | Partners wait at counters or food goes cold | Feed actual ready times back into the model; re-plan dispatch on delay |
| Duplicate payment request | Double charge | Idempotency keys on every posting; reconcile ledger against providers daily |
| Event consumer lags | Rewards or incentives arrive late | Alert on consumer lag; never block the booking path on it |
| Shared platform outage | Every vertical fails at once | Per-vertical bulkheads, rate limits and degraded modes |
The last row is the central trade-off of a super app. Sharing platform services is what makes it one product, and it is also a shared blast radius. Isolate capacity per vertical inside each shared service, so a food promotion that floods the system cannot stop rides being allocated.
What to do next
- Draw your own system's vertical and platform layers, and list every capability more than one vertical needs.
- For each shared service, write down its API contract and who can change it; fix any that a vertical has forked.
- Separate latest location from location history; size the in-memory index from pings per second and pick a geohash precision from your search radius.
- Replace nearest-driver logic with a scoring function and a short batch window, and measure ETA accuracy before tuning weights.
- For any item with preparation time, schedule dispatch from ready time minus pickup ETA, and log actual ready times.
- Move money onto a double-entry ledger with idempotency keys, and run cross-service workflows as sagas with compensating postings.
- Publish domain events through an outbox, make consumers idempotent and centralise cross-vertical rules in one engine.
- Add per-vertical bulkheads to every shared service and test what happens when one vertical floods it.