WeChat is the standard example of a super-app: one client that combines messaging, a social feed, payments and a platform for third-party mini programs, used by well over a billion people. Running all of that in one app changes the backend problem. Every feature shares one network connection, one identity, one set of servers that can be overloaded together, and one trust boundary that third-party code must not cross.
Much of WeChat's internal design is not public, so this article is careful about sources. It first sets out what WeChat engineers have published in papers and open-source projects, then builds a reference design around those pieces for the parts that are not public, labelled as such. The aim is a design you could defend in a review and build with ordinary tools. The connection tier is covered in depth in designing a real-time chat system.
What is published and what is inferred
Several pieces of WeChat's backend are documented by its engineers. The DAGOR paper, presented at ACM SoCC 2018 under the title 'Overload Control for Scaling WeChat Microservices', describes the backend as more than 3,000 services on over 20,000 machines and gives its overload-control algorithm in detail. PaxosStore, published at VLDB 2017, describes the Paxos-based storage layer used across many WeChat services. Tencent has open-sourced Mars, the cross-platform client networking and logging component; libco, the C++ coroutine library used to write synchronous-looking server code; and PhxPaxos, a Paxos library. WeChat engineers have also written publicly about seqsvr, a service that hands out per-user sequence numbers for message sync. The mini-program login flow is documented on the official developer site.
What is not public in detail: exact service boundaries, schemas, capacity numbers per service, group fan-out strategy at every size, and the payment system's internal ledger design. Where this article describes those, it is a reference design, not a description of WeChat.
The shape of the system
The client keeps one long-lived connection to an access layer. Mars handles connection management on the client: reconnection with backoff, switching between networks, heartbeats tuned to mobile carriers' idle timeouts, and short-connection fallback. Every feature multiplexes over that one connection. This is cheaper on battery and radio than one connection per feature, and it gives the server one place to authenticate the device.
Behind the access layer sit stateless logic services that call each other over RPC, and below them storage services. The access layer holds connection state and nothing else, so it can be scaled and drained separately. Logic services are written as many small services, which is why overload control has to be designed for a call graph rather than a single server.
Messaging: sync by sequence number, not push
The central idea in WeChat's messaging, as its engineers have described in writing about seqsvr, is that each user has a monotonically increasing sequence number. Every item delivered to that user, whether a chat message, a contact change or a notification, is written to the user's inbox with the next number. The client remembers the highest number it has applied. To catch up, it sends that number and receives everything above it.
This turns delivery into a pull with a cursor. Pushes become small notifications that say 'your sequence has advanced', after which the client syncs. Lost pushes are harmless, because the next sync covers the gap. A phone that was offline for a week sends one sync request. Multiple devices each keep their own cursor against the same inbox. The server never has to track which message reached which device.
The published seqsvr design keeps allocation fast by handing out numbers from memory and persisting only an upper bound, raised in steps. After a crash the allocator restarts from the persisted bound, so some numbers are skipped but none is ever issued twice. Gaps are fine for a cursor; repeats would not be. That property, monotonic but not dense, is the one to copy.
# Reference design: per-user inbox with a sequence cursor
def deliver(user_id, item):
seq = seqsvc.next(user_id) # monotonic per user, may skip, never repeats
inbox.put(user_id, seq, item) # replicated write; acknowledged only when durable
notify.wake(user_id, seq) # best-effort; losing it is safe
def sync(user_id, client_seq, limit=200):
items = inbox.range(user_id, start=client_seq + 1, limit=limit)
new_cursor = items[-1].seq if items else client_seq
return items, new_cursor # client applies items, then stores new_cursor
Groups and the fan-out choice
A group message must reach every member's inbox. There are two ways. Write fan-out copies the message, or a pointer to it, into each member's inbox at send time, so sync stays a single range read per user. Read fan-out stores the message once in a group log, and each member's sync must also read every group they belong to.
Write fan-out makes reads cheap and uniform, at the cost of work proportional to group size on every send. It is practical when groups have a hard size cap, and WeChat groups are capped at a few hundred members. Public accounts and channels with millions of followers are a different product and need read fan-out or a hybrid. The Moments feed, visible only to friends, has the same choice between pushing each post into friends' timelines and pulling at read time. Treat any public statement of which one WeChat uses as secondhand; the trade-off is what matters for your design.
Worked example: one message to a 500-member group
Assume write fan-out of a pointer. The sender's client sends one message. The message service stores the body once, then for each of 500 members allocates a sequence number and writes a small inbox entry of perhaps 100 bytes containing the message id. That is 500 sequence allocations and 500 replicated writes, about 50 KB of inbox data plus the body.
Now consider a peak where 100,000 such group messages are sent per second. That becomes 50 million inbox writes per second, plus up to 50 million wake-ups, most of which are coalesced because recipients are offline or already syncing. This is why sequence allocation must be in-memory and batched, why inbox writes must be cheap appends, and why notifications must be best-effort. It also shows why the group cap is an architectural decision, not a product whim: doubling it doubles the write load of every message in the largest groups.
Storage: replicated before acknowledged
An acknowledged message must never disappear, even if a data centre is lost. PaxosStore addresses this by replicating each write through Paxos to replicas in different data centres before acknowledging it, and the paper describes running different storage engines for different data shapes beneath one consensus layer. The point to take away is the contract: the sender's tick means a quorum has the data.
In a reference design you would get the same contract from a consensus-replicated store such as one built on Raft or Paxos, keyed by user id so that one user's inbox lives in one replica group, and partitioned so that groups can be moved for balance. Sharding strategy is covered in database sharding.
Overload control: DAGOR
In a call graph of thousands of services, overload in one service cascades. DAGOR's design, as published, has four parts.
- Detection by queuing time. Each service measures the average time requests wait in its pending queue, not CPU. The paper uses a 20 ms threshold with a 500 ms default request timeout, checked every second or every 2,000 requests, whichever comes first.
- Business priority. Each request carries a priority set at the entry service according to the action, so that, for example, login and payment rank above lower-value features. Downstream calls inherit it.
- User priority. Within a business priority, users get one of 128 priorities from a hash of the user id that changes every hour. When shedding, a service drops whole users rather than random requests, so admitted users get a complete experience and the burden rotates fairly.
- Collaborative shedding. Services piggyback their current admission level on responses, so callers reject requests locally before sending them to an overloaded service. This avoids the 'subsequent overload' problem in which a request that calls the same overloaded service several times almost always fails partway.
When overloaded, a service lowers its admission level to shed about 5% of load; when healthy, it raises it to admit about 1% more. The asymmetry makes it back off quickly and recover cautiously.
# Reference sketch of DAGOR-style admission at one service
level = (MAX_BUSINESS, MAX_USER) # admit requests at or above this compound level
def admit(req):
return (req.business_prio, req.user_prio) <= level # lower tuple = more important
def on_window_end(avg_queue_ms, histogram):
global level
if avg_queue_ms > 20:
level = histogram.level_that_sheds(fraction=0.05)
else:
level = histogram.level_that_admits_more(fraction=0.01)Rate limiting at the edge complements this but cannot replace it, because the edge does not know which internal service is struggling. See rate limiter architecture for the edge side.
Mini programs: running third-party code safely
Mini programs are third-party apps that run inside WeChat. The official framework documentation splits each one into a logic layer, which runs the developer's JavaScript, and a view layer, which renders pages from WXML and WXSS templates. The two communicate through the host app, and the developer updates the view by sending data, not by touching the DOM. This keeps third-party code away from the rendering surface and lets the host mediate every capability call.
Identity is the most security-sensitive flow. The client calls wx.login to get a short-lived code and sends it to the developer's server. The server calls auth.code2Session with its app id and app secret and receives the user's openid, which is specific to that app, and a session_key; it may also receive a unionid that is stable across one developer's apps when the documented conditions are met. The documentation is explicit that the session key must never be sent to the client.
# Developer backend: exchange a wx.login code for an app-scoped identity
def login(js_code):
r = http.get("https://api.weixin.qq.com/sns/jscode2session", params={
"appid": APP_ID, "secret": APP_SECRET,
"js_code": js_code, "grant_type": "authorization_code"})
data = r.json()
if "openid" not in data:
raise AuthError(data)
store_session(data["openid"], data["session_key"]) # server side only
return issue_own_token(data["openid"]) # your session, not WeChat's keyFor a reference design of any mini-program platform, the same principles apply: per-app user identifiers so apps cannot correlate users across developers, code review and signing before publication, capability calls mediated by the host with explicit user consent, and per-app quotas so one popular app cannot exhaust shared services.
Payments inside the app
Payments follow the standard pattern for a provider integrated into a client. The merchant's server creates an order with the payment provider and receives a prepay identifier; the client uses it to open the payment sheet, where the user authenticates; the provider later sends a signed asynchronous notification to the merchant's server. The notification, or an explicit order query, is the source of truth. The client's 'success' callback is a hint only, because the app may be killed before it fires.
Merchants must verify the notification's signature, process it idempotently because it can be delivered more than once, and reconcile against the provider's statements daily. Inside the super-app, payments get the highest business priority so that overload sheds chat stickers before it sheds checkouts. General payment-system design is in designing a payment system.
Failure modes
- Reconnect storms. A network event drops millions of connections, and they all reconnect and sync at once. Jittered backoff in the client and admission control at the access layer are both required.
- Sequence repeats. An allocator that restarts below a number it already issued makes clients skip messages permanently. Persist the upper bound before using numbers below it.
- Peak-day hot keys. Festival traffic such as red-packet sends concentrates on a few groups and accounts. Pre-scale, cap per-group rates and keep red-packet grabbing on a separate, priority-managed path.
- Third-party abuse. A mini program that loops on an API or harvests identifiers. Per-app quotas, app-scoped ids and the ability to suspend an app quickly.
- Partial cascades. One slow storage shard raises queuing time in every caller. Queuing-time detection and piggybacked admission levels contain it.
Trade-offs
| Decision | Benefit | Cost |
|---|---|---|
| One connection for all features | Battery, one auth point, simpler client | Access tier is critical for every feature |
| Sync by sequence cursor | Lost pushes harmless, multi-device simple | Allocator must be fast and never repeat |
| Write fan-out with group cap | Cheap, uniform reads | Send cost proportional to group size |
| Consensus-replicated inboxes | No acknowledged message lost | Cross-data-centre write latency |
| Priority-based shedding | Important actions survive overload | Some users are refused, by design |
What to do next
- Write down the delivery contract: when is a message acknowledged, and what storage guarantee backs that acknowledgement.
- Implement per-user inboxes with a monotonic, gap-tolerant sequence and a pull-based sync; make notifications best-effort.
- Choose a group size cap from fan-out arithmetic at peak load, and a separate read-fan-out design for broadcast channels.
- Assign business priorities to every entry API, propagate them on every internal call, and measure queuing time in every service.
- Add piggybacked admission levels so callers shed before calling an overloaded service.
- For third-party code, use app-scoped user ids, a server-side code exchange that keeps session keys off the client, and per-app quotas.
- Treat payment notifications as the source of truth and process them idempotently; reconcile daily.
- Rehearse a mass reconnect and a festival peak in a load test before they happen in production.