Slack looks like a chat app, but its hard problems are those of a workplace system of record. Messages belong to channels, channels belong to workspaces and organisations, threads hang off individual messages, and every user expects an exact unread count, search over years of history and delivery in well under a second. A single channel can have tens of thousands of members, and the same user may be logged in from a laptop, a phone and a browser tab at once.
This article explains how Slack is built, as far as Slack's engineering team has described it publicly: the channel, gateway, admin and presence servers that move messages in real time, consistent hashing of channels and how failures are absorbed, the original workspace-sharded MySQL estate and why it moved to Vitess with channel-level sharding, the Flannel edge cache that made client boot cheap, and the data model of channels and threads exposed by the public API. Where the article fills a gap Slack has not documented, it says so and labels the design as a reconstruction. It ends with a worked fan-out example, failure modes and a checklist.
Channels, messages and threads
Start with the nouns, because the architecture follows from them. A workspace (called a team in the API) is a company or group; an Enterprise Grid organisation groups many workspaces; and Slack Connect lets a single channel be shared between separate organisations. A conversation is a public or private channel, a direct message or a group direct message, and each has an ID. A message is identified by the pair of its conversation ID and its ts, a timestamp string with microsecond precision that is unique within the conversation. The ts doubles as the sort key, the pagination cursor and the target for edits, reactions and pins.
A thread is not a separate object. Replies are ordinary messages in the same conversation that carry a thread_ts equal to the parent's ts. The parent carries summary fields such as the reply count and the latest reply. A reply posted with reply_broadcast set also appears in the main channel timeline. The Web API reflects this directly: conversations.history returns the channel timeline, and conversations.replies takes a channel and a parent ts and returns the thread.
Two consequences shape everything else. First, a thread always lives with its channel, so any storage layout that keeps a channel together also keeps its threads together. Second, per-user state, which channels you belong to, your last-read ts in each, which threads you follow, is separate from the messages and is read on every client start. That is the data that made booting large workspaces expensive, and it is why Slack built an edge cache for it.
Channels, messages and threads
Start with the nouns, because the architecture follows from them. A workspace (called a team in the API) is a company or group; an Enterprise Grid organisation groups many workspaces; and Slack Connect lets a single channel be shared between separate organisations. A conversation is a public or private channel, a direct message or a group direct message, and each has an ID. A message is identified by the pair of its conversation ID and its ts, a timestamp string with microsecond precision that is unique within the conversation. The ts doubles as the sort key, the pagination cursor and the target for edits, reactions and pins.
A thread is not a separate object. Replies are ordinary messages in the same conversation that carry a thread_ts equal to the parent's ts. The parent carries summary fields such as the reply count and the latest reply. A reply posted with reply_broadcast set also appears in the main channel timeline. The Web API reflects this directly: conversations.history returns the channel timeline, and conversations.replies takes a channel and a parent ts and returns the thread.
Two consequences shape everything else. First, a thread always lives with its channel, so any storage layout that keeps a channel together also keeps its threads together. Second, per-user state, which channels you belong to, your last-read ts in each, which threads you follow, is separate from the messages and is read on every client start. That is the data that made booting large workspaces expensive, and it is why Slack built an edge cache for it.
The real-time tier
Slack's engineering post on real-time messaging describes four server roles, all keeping state in memory.
- Channel servers (CS) are stateful and hold recent history for the channels they own. Channels are assigned to channel servers by consistent hashing on the channel ID, so one workspace's channels are spread across the whole fleet rather than concentrated on one machine. Slack quotes around 16 million channels served per host at peak.
- Gateway servers (GS) are stateful too. They hold each connected user's information and websocket subscriptions, sit between clients and channel servers, and are deployed in multiple geographic regions so users connect nearby.
- Admin servers (AS) are stateless and connect the Webapp backend, Slack's main application tier, to the channel servers.
- Presence servers (PS) track who is online. Users, not channels, are hashed to presence servers.
A fifth component, the Consistent Hash Ring Managers (CHARMs), owns the ring of channel servers and replaces an unhealthy one in under 20 seconds, according to the same post. Clients reach gateways through Envoy at the nearest edge region.
Connecting goes like this: the client gets a token and websocket connection details from the Webapp, opens the websocket to the nearest region, and the gateway fetches the user's channel list from the Webapp and asynchronously subscribes to the channel servers that own those channels. Sending goes the other way: the client calls the Webapp API, the Webapp hands the message to an admin server, the admin server finds the owning channel server through the ring, and the channel server pushes the event to every gateway that has a subscriber, which writes it to each user's sockets. Slack reports roughly 500 ms delivery worldwide. Transient events such as typing indicators take a shorter path, client to gateway to channel server and back out, with no persistence.
Two-stage fan-out and the hash ring
The useful property is that the expensive fan-out happens in two stages. The channel server sends one copy per gateway with a subscriber, and each gateway makes the per-user copies for the sockets it holds.
# Illustrative model of the published real-time path; names are ours.
import bisect, hashlib
def h(key):
return int.from_bytes(hashlib.sha1(key.encode()).digest()[:8], "big")
class Ring:
"""Consistent hash ring of channel servers, with virtual nodes."""
def __init__(self, servers, vnodes=64):
self.points = sorted((h(f"{s}#{v}"), s) for s in servers for v in range(vnodes))
self.keys = [p for p, _ in self.points]
def owner(self, channel_id):
i = bisect.bisect(self.keys, h(channel_id)) % len(self.points)
return self.points[i][1]
class ChannelServer:
def __init__(self):
self.subscribers = {} # channel_id -> set of gateway ids
def subscribe(self, channel_id, gateway_id):
self.subscribers.setdefault(channel_id, set()).add(gateway_id)
def publish(self, channel_id, event, send):
for gw in self.subscribers.get(channel_id, ()):
send(gw, event) # one copy per gateway, not per user
class Gateway:
def __init__(self):
self.sockets = {} # channel_id -> set of open websockets
def deliver(self, channel_id, event):
for ws in self.sockets.get(channel_id, ()):
ws.write(event) # per-user copies happen here, near the userVirtual nodes keep ownership balanced, and the ring means that when a channel server is replaced only its share of channels moves. Gateways learn of the new owner and resubscribe. What the ring does not solve is a single enormous channel: all of its events still pass through one channel server. In practice that is bounded by the number of gateways rather than the number of members, which is the point of the two-stage design.
Storage: from workspace shards to channel shards
Slack's original storage was MySQL sharded by workspace: each shard held every workspace on it with all their messages and channels, and each shard was a pair of MySQL instances in different datacenters replicating to each other asynchronously, both taking writes. Slack's own account of why it stopped working lists the problems: the largest customers outgrew a single shard's hardware, their load could not be spread while thousands of other shards sat mostly idle, one shard failure took Slack down for every customer on it, and Enterprise Grid and Slack Connect broke the assumption that a team's data belongs together. A shared channel spans organisations; under workspace sharding, whose shard owns it?
The fix was Vitess, which puts a routing layer and sharding metadata in front of MySQL and lets each table choose its own sharding key. Messages could then be sharded by channel ID instead of by workspace. Slack reported that by December 2020 99% of its MySQL traffic ran through Vitess, at a peak of about 2.3 million queries per second (2 million reads, 300,000 writes), with median latency of 2 ms and p99 of 11 ms. Features such as mentions, reactions and sessions moved over table by table.
Channel-ID sharding is the natural fit for the data model described earlier: a channel's history and all of its threads live on one shard, history pagination is a range scan on (channel_id, ts), and a Slack Connect channel has exactly one home. The cost is that per-user views, such as "all my unread threads" or "all mentions of me", span many channels and therefore many shards, so they need their own tables keyed by user and maintained asynchronously.
-- Reconstruction consistent with the public API, not Slack's actual schema.
-- Sharded by channel_id, so a channel's history and its threads live together.
CREATE TABLE messages (
channel_id VARBINARY(16) NOT NULL,
ts DECIMAL(16,6) NOT NULL, -- unique within the channel; the message's ID
user_id VARBINARY(16) NOT NULL,
thread_ts DECIMAL(16,6) NULL, -- NULL if no replies; own ts on a parent; parent ts on replies
reply_broadcast BOOLEAN NOT NULL DEFAULT FALSE,
client_msg_id CHAR(36) NULL, -- client-generated, for idempotent retries
body MEDIUMBLOB NOT NULL,
edited_ts DECIMAL(16,6) NULL,
deleted BOOLEAN NOT NULL DEFAULT FALSE,
PRIMARY KEY (channel_id, ts),
KEY thread (channel_id, thread_ts, ts),
UNIQUE KEY dedupe (channel_id, user_id, client_msg_id)
);
-- Per-parent summary, updated in the same transaction as each reply.
CREATE TABLE thread_summary (
channel_id VARBINARY(16) NOT NULL,
thread_ts DECIMAL(16,6) NOT NULL,
reply_count INT NOT NULL,
latest_reply DECIMAL(16,6) NOT NULL,
PRIMARY KEY (channel_id, thread_ts)
);The schema above is a reconstruction that satisfies the public API, not Slack's real one. Slack's API does expose a client_msg_id on messages; how it is enforced server-side is not published.
Booting clients and asynchronous work
Booting a client used to mean downloading a model of the whole workspace: every user, every channel, every membership. For a workspace with hundreds of thousands of members that payload was enormous and slow. Slack's answer was Flannel, a globally distributed edge cache that serves the workspace model lazily. The client boots with a much smaller payload and asks Flannel for users and channels as it needs them. Slack has described Flannel answering over a million queries per second, and also the cost: cache coherency with the source of truth and duplicated business logic at the edge became new problems to manage.
Work that does not need to happen before the sender sees their message goes onto an asynchronous job queue: push notifications to offline users, link unfurls, search indexing, and similar tasks. This keeps the synchronous send path to persistence plus real-time fan-out, and it means a burst of notifications cannot slow message delivery.
Worked example: a 50,000-member announcement
Take an announcement posted to a company-wide channel with 50,000 members, of whom 15,000 are connected, spread over 200 gateway servers. The numbers are illustrative; the structure follows the published design.
- The client sends the message to the Webapp with a fresh
client_msg_id. The Webapp writes one row to the shard that owns the channel ID, receives thetsand returns it to the sender. - The Webapp passes the event to an admin server, which looks up the channel's owner on the ring and forwards it.
- The channel server sends one copy to each of the 200 subscribed gateways, not 15,000 copies.
- Each gateway writes the event to its local sockets for that channel, about 75 per gateway on average, so the 15,000 socket writes are spread across the fleet and happen close to users.
- Mentions and notification rules for the 35,000 offline members become jobs. If 10% of them have mobile push enabled for this channel, that is 3,500 push jobs, processed at the queue's pace, not the sender's.
- Now 400 people reply in the thread over ten minutes. Each reply is a message row with
thread_tsset plus an update to the parent summary. Clients showing the channel update the reply count; only clients showing the thread, and users following it, need the full reply. That difference is why threads cut load compared with 400 top-level messages: most of the 15,000 connected clients only repaint a counter.
Failure modes
- Channel server loss. In-memory state is gone, and CHARMs replace the server in under 20 seconds. Because messages are persisted before fan-out, nothing durable is lost; transient events are. Clients must reconcile by fetching history after the last
tsthey saw, which is why a channel-scoped, sortable ID matters. - Region failure. Slack drains the region and clients reconnect to the nearest healthy one. The reconnect wave is the hazard: every reconnecting client re-fetches its channel list and resubscribes. Jittered backoff and the edge cache absorb most of it.
- Hot shard or hot channel. Channel-ID sharding spreads workspaces but a single viral channel still lands on one shard and one channel server. Caching recent history and the two-stage fan-out bound it; there is no way to split one channel's write order across shards without giving up per-channel ordering.
- Duplicate sends. Mobile clients retry on flaky networks. Without a dedupe key, a retry creates a second message with a new
ts. - Cache incoherence. An edge cache that serves membership can show a user a channel they were just removed from. Permission checks must happen at the source of truth on every read of message content, with the cache trusted only for display hints.
What to do next
Related reading on this site: designing a real-time chat system for the websocket tier, Discord's architecture and WhatsApp's architecture for contrasting designs, database sharding for shard maps and resharding, and publish-subscribe systems.
- Write down your system's ownership object, the thing whose events must be ordered, and check that your storage shards by it.
- Make message IDs scoped to that object and sortable, and build reconnect logic that fetches everything after the last ID a client saw.
- Add a client-generated idempotency key to every send and enforce it with a unique index on the owning shard.
- Measure fan-out at each stage separately: copies per channel server and socket writes per gateway, and alert on the largest channel's share.
- Model threads as replies in the parent's partition with a summary row, and send full replies only to clients viewing or following the thread.
- Move notifications, unfurls and indexing onto a queue and alert on its backlog age, not just its length.
- Run a game day that kills one stateful server and one region, and measure reconnect load and time to full delivery.