Alipay started as an escrow service for online shopping and grew into a super app: payments, transfers, wealth products, bills, transport, mini programs from thousands of third parties, all behind one login and one balance. The part of that story most worth studying is not the app screen. It is how Ant Group rebuilt the back end so that a system moving money for a very large user base could keep growing without a single database or data centre becoming the ceiling, and so that losing a data centre affects a slice of users for minutes instead of everyone for hours.

The answer was unitization, which Ant calls LDC, the logical data centre. This article explains it from first principles: why a sharded database alone is not enough, how RZone, GZone and CZone divide the system, how a request finds its unit, what happens when a payment involves two users in different units, how failover works at unit granularity, and what it costs. It closes with failure modes and a checklist for applying the pattern to your own system.

A note on sources. Ant has published the zone model in its SOFAStack documentation on Alibaba Cloud and in OceanBase talks; both are summarised here. Internal details such as exact unit counts in production, peak throughput and data-centre layout change over time and are not claimed. Where the article describes how a component would work, it is a reasoned design, labelled as such, not a disclosure.

Why sharding was not enough

A conventional scaled-out payment system shards its databases by user and runs stateless application servers behind a load balancer. That gets you storage capacity, but three limits remain. First, any application server can talk to any shard, so the number of database connections grows with servers times shards, and connection counts become a hard ceiling. Second, the failure domain is still the whole system: a bad deploy, a saturated cache cluster or a data-centre outage hits every user, because every request may touch every tier. Third, disaster recovery is coarse. Failing over means moving everything from one site to another, and nobody can rehearse that often enough to trust it.

Unitization changes the unit of scaling from the tier to a vertical slice. A unit contains everything needed to serve a subset of users: application instances for each service, their caches, message queues and a database shard. A request for a user enters that user's unit and, for most operations, never leaves it. Capacity grows by adding units, a failure is contained to the units it hits, and disaster recovery means moving a few units, which can be practised routinely.

The LDC zone model

Not everything can be sliced by user. The SOFAStack description defines three zone types. An RZone (regional zone) holds data and services that partition cleanly by a key such as user ID. The published example divides users into 100 logical units, so each represents about 1% of traffic, and several units are deployed in each physical RZone. A GZone (global zone) holds what cannot be split, such as configuration and some customer information, and there is only one of it for the whole system. A CZone (city zone) is a read mirror of GZone data placed in each city, so that RZones read global data locally instead of crossing cities on every request.

Gateway and routeruser ID to unituid 00-49uid 50-99City ARZone 00-24apps + DBRZone 25-49apps + DBCity BRZone 50-74apps + DBRZone 75-99apps + DBGZoneshared data, one writerCZone in City Aread mirror of GZoneCZone in City Bread mirror of GZonereplicatereplicatelocal readslocal readsEach RZone owns a slice of users end to end; GZone holds what cannot be split; CZones mirror it per city.
The LDC zone model: user-partitioned RZones grouped by city, a single writable GZone, and per-city CZone mirrors for local reads.

The distinction between logical units and physical zones is what makes migration cheap. The 100 logical units are fixed; how they map to physical RZones in particular data centres is configuration. Moving units 25 to 49 from City A to City B is a change to the routing table plus data movement, not a re-shard. It is the same idea as consistent hashing with virtual nodes, applied to entire application stacks.

The CZone pattern deserves a precise statement. Writes to global data still go to the single GZone; CZones receive them asynchronously. That is acceptable only for data that is read often, written rarely and tolerant of brief staleness, such as merchant configuration or feature switches. Anything that needs read-after-write consistency across cities, such as an account balance, must live in an RZone and be owned by exactly one unit.

Routing a request to its unit

Every layer must agree on which unit owns a user. The routing key is derived from the user ID, for example two digits of the ID, giving a number from 0 to 99. The mapping from that number to a physical zone is held in a routing table that every layer subscribes to. The following sketch shows the logic at each hop; it is a reasoned design, not Ant's code.

ROUTES = config.subscribe("ldc.routes")   # {unit: zone}, versioned, pushed on change

def unit_of(user_id: str) -> int:
    return int(user_id[-4:-2])            # stable two-digit slice, 0..99

def handle(request):
    unit = unit_of(request.user_id)
    owner = ROUTES[unit]
    if owner != LOCAL_ZONE:
        # Wrong zone: forward once, never serve. Stale clients and DNS land here.
        return forward(request, zone=owner, hop_limit=1)
    with db_shard(unit) as tx:            # shard lives in this zone
        return service.process(request, tx)

def call_downstream(service, user_id, payload):
    # RPC layer applies the same rule, so internal calls stay in-unit.
    return rpc.invoke(service, payload, zone=ROUTES[unit_of(user_id)])

Three rules make this work. Routing happens as early as possible, at the gateway, so most requests enter the right zone directly. Every later layer re-checks, so a stale entry in one layer is corrected by the next instead of writing to the wrong shard. And the database layer refuses writes for units it does not own; this is the last line of defence during a migration, when the routing table and the data may briefly disagree. Without that refusal, two zones can both accept writes for the same user, which in a payment system is a reconciliation incident.

Worked example: a transfer across units

Most user operations touch one user, but money moves between two. Suppose Alice is in unit 07 in City A and Bob is in unit 63 in City B, and Alice sends Bob 100 yuan. No single local transaction can cover both balances. The general approach, consistent with Ant's published work on distributed transactions, is to commit Alice's side locally together with a durable intent, and complete Bob's side through a reliable, idempotent step.

  1. The request is routed to unit 07. In one local transaction, the transfer service debits Alice (or freezes the amount), inserts a transfer record with a unique transfer ID in state PENDING, and writes an outbox message addressed to unit 63.
  2. A relay delivers the message to unit 63. There, in one local transaction, the service checks that the transfer ID has not been applied, credits Bob, and records the ID as applied.
  3. Unit 63 acknowledges. Unit 07 marks the transfer COMPLETED. If unit 63 rejects the credit, for example because Bob's account is frozen, unit 07 runs a compensating credit back to Alice and marks the transfer FAILED.
  4. A reconciliation job compares transfer records on both sides and alerts on any transfer that is PENDING for longer than its service level.

This is the saga pattern with an outbox, described in general terms in Saga pattern. Ant has also described try-confirm-cancel style protocols for cases that need resources reserved on both sides before either commits; the freeze in step 1 is the try. What matters for the architecture is that cross-unit work is explicit, asynchronous and idempotent, and that its volume is a design metric. If a feature makes most requests cross units, it is in the wrong zone or keyed by the wrong ID. Contrast this with the single-ledger view in Designing a payment system, which unitization deliberately splits.

Data placement and the global zone

Unit-local data still needs replication for durability. OceanBase, Ant's distributed relational database, keeps multiple replicas of each partition and uses Paxos so that a write commits once a majority acknowledges it. Placing a partition's replicas across data centres, for example two in the home city and one elsewhere, lets the unit survive a data-centre loss without losing committed writes, at the cost of commit latency on the replication path. The practical rule is that the leader for a unit's partitions sits in the same zone as the unit's application tier; a leader in another city turns every statement into a long round trip.

GZone data is the scaling risk. It has one writer, every RZone depends on it, and its mirrors lag. Keep it small, keep its write rate low, give it its own capacity plan, and make RZones survive a GZone outage by serving from the CZone mirror in read-only mode for global data. If a new feature wants to write global data on the hot path, push back: either the data can be keyed by user and moved to RZones, or it can be written asynchronously.

Failover as a routing change

Failover in this design is a routing change. If a data centre hosting units 25 to 49 fails, operators update the routing table to point those units at a standby zone whose database replicas are already in sync, and every layer picks up the new version. Because each unit is about 1% of users, operators can move traffic in small steps, watch error rates, and continue. The same mechanism drains a zone for maintenance, shifts load during a sale peak, or canaries a release on one unit before the rest.

The hard part is the moment of change. Old routes and new routes coexist briefly, so the database refusal rule from the routing section matters more than anywhere else. Replica promotion must be confirmed before the routing switch, not after. And in-flight cross-unit messages addressed to a moved unit must be redelivered to its new home, which is why messages target logical units, not physical hosts. Teams that run unitized systems rehearse these switches regularly, because a failover path that is not exercised does not work when needed.

The super-app layer on top

The super-app layer sits on top. Ant's mobile platform, mPaaS, provides the client framework, release and gateway tooling, and Alipay hosts third-party mini programs in a sandboxed container that separates rendering from logic and exposes capabilities such as payment through controlled APIs. From the back end's point of view, a mini program is one more caller whose requests carry a user identity and therefore route to that user's unit. The comparison with WeChat's mini program model is covered in WeChat super-app architecture, and a different regional super app in Grab super-app architecture.

Failure modes

  • Split-brain writes during migration. Two zones accept writes for the same unit. Guard every write with an ownership check at the database layer and switch routes only after replica promotion.
  • Hot units. A merchant or celebrity account concentrates traffic in one unit. Keep the routing key independent of business popularity, and give hot entities their own handling, such as sub-accounts for a large merchant.
  • GZone dependency creep. New features read or write global data on the critical path until GZone is the bottleneck again. Review every new global dependency.
  • Stale CZone reads. A merchant changes configuration and some cities serve the old value. Version global data and tolerate lag explicitly, or move the data to an RZone.
  • Stuck cross-unit transfers. A relay outage leaves transfers PENDING. Alert on age, make every step idempotent, and reconcile both sides.
  • Untested failover. Routing, replica placement and message redelivery drift apart. Drill unit moves on a schedule.

Trade-offs

Unitization buys horizontal growth, small blast radius and practised disaster recovery. It costs a great deal of engineering. Every service must be unit-aware, every data model must declare its zone, cross-user operations become distributed protocols, and analytics that need a global view must assemble it from all units. Most companies do not need it: a well-sharded database with a regional standby serves many millions of users. Consider it when one of the original three limits actually binds: connection counts, whole-system blast radius, or failover that cannot be rehearsed. For a national-scale comparison built on a shared switch instead, read UPI payments architecture.

What to do next

  1. List your data entities and classify each as user-partitionable, global and rarely written, or global and frequently written.
  2. Pick a routing key and a fixed number of logical units, and write down how units map to physical zones.
  3. Add a unit ownership check to every write path, starting with the database access layer.
  4. Measure the share of requests that cross units, and redesign features where it is high.
  5. Convert cross-unit money movement to an outbox plus idempotent consumer with reconciliation.
  6. Move one unit between zones in a test environment, including in-flight messages, and time it.
  7. Shrink the global zone: move user-keyed data out and make global reads survive a GZone outage.
Key takeaway: Alipay scales by unitization: users are split into fixed logical units, each served end to end by its own slice of applications and data in an RZone, while unsplittable data lives in a single GZone mirrored per city by CZones. Every layer routes by user ID and the database refuses writes for units it does not own. Cross-unit payments become outbox-driven, idempotent, reconciled steps, and failover becomes a rehearsable routing change at about 1% granularity.