A feature flag looks like an if statement with a remote control. At the scale of one service that is all it is. At the scale of a company it becomes a distributed configuration system that sits on the hot path of every request, changes production behaviour without a deploy, and can take the whole product down faster than any code change. Flags gate risky launches, run experiments, act as kill switches during incidents and hold entitlements for paying customers, and each of those uses pulls the design in a different direction.
This article designs the system the way you would in a design review. We start from requirements and the numbers that shape them, define a data model with explicit versioning, write down evaluation semantics precisely enough that two SDKs in different languages give the same answer, build a propagation path that keeps working when the control plane is down, and then walk through failure modes and a real rollout. The operational side of delivery, including SDK wiring and flag lifecycle policy, is covered in the feature flag delivery guide; here the focus is on the design decisions and why they are made.
Requirements and the numbers that shape the design
Functional requirements: typed variations; targeting by attributes, segments and percentage; separate environments; approvals and scheduling for production changes; a full audit trail; and exposure events for experiment analysis.
Non-functional requirements are where the design is decided. Evaluation happens inside request handlers, often dozens of times per request, so it must cost microseconds and must never make a network call. A change must reach every running process quickly, because a kill switch that takes ten minutes is not a kill switch; a reasonable target is p99 under five seconds worldwide. Evaluation must keep working if the flag service, its database or the whole control plane region is down. And every process must be able to say which configuration version it is using, so incidents can be reconstructed.
Some estimates for a mid-sized company: 20,000 flags across 10 environments, 3,000 services, 60,000 running SDK instances, and a peak of roughly 50 million evaluations per second in aggregate. The full rule set for one environment is typically a few megabytes of JSON. Changes are rare by comparison: perhaps a few thousand per day. That ratio, millions of reads per second against a handful of writes per minute, is the whole argument for the architecture that follows: push the entire rule set to every process and evaluate locally.
Architecture: slow audited writes, local reads
The write side is an ordinary transactional service. The read side is a distribution network whose only job is to get the latest versioned snapshot into memory everywhere.
A change is written to Postgres in one transaction together with an outbox row and an audit row; the outbox pattern guarantees that a committed change is always published and an uncommitted one never is. A relay reads the outbox in commit order and assigns each change a monotonically increasing environment version. Two consumers read that feed. The snapshot builder writes a full, immutable snapshot per environment and version to object storage behind a CDN. The relay fleet keeps the latest snapshot in memory and streams deltas to connected SDKs over server-sent events, a fan-out problem shaped much like publish-subscribe. Server SDKs hold the rules in memory and evaluate locally. Browsers and mobile apps talk to an edge evaluator instead, for reasons covered below.
The data model and versioning
Model a flag as a key plus a list of variations, and model each environment's configuration of that flag separately, because production and staging diverge constantly. Rules are ordered and reference segments by key, so a segment can be reused by hundreds of flags. Every flag environment carries its own version for optimistic concurrency, and every published snapshot carries the environment version.
{
"env": "production", "version": 1842,
"flags": {
"new-checkout": {
"version": 37, "on": true, "salt": "a91f03",
"variations": [false, true],
"offVariation": 0,
"prerequisites": [{"flag": "payments-v2", "variation": 1}],
"targets": [{"variation": 1, "values": ["user-qa-1", "user-qa-2"]}],
"rules": [
{"id": "r1", "clauses": [{"attr": "segment", "op": "in", "values": ["employees"]}],
"serve": {"variation": 1}},
{"id": "r2", "clauses": [{"attr": "country", "op": "in", "values": ["NZ", "IE"]}],
"serve": {"rollout": [{"variation": 0, "weight": 7500},
{"variation": 1, "weight": 2500}], "bucketBy": "userId"}}
],
"fallthrough": {"variation": 0}
}
},
"segments": {"employees": {"version": 9, "included": ["..."], "rules": []}}
}Two details matter later. The per-flag salt means the same user lands in unrelated buckets for unrelated flags, so the 10 percent who get one experiment are not always the same 10 percent who get every other. And weights are expressed in basis points out of 10,000, which gives rollout percentages to two decimal places without floating point in the evaluation path.
Evaluation semantics, precisely
Flag evaluation must be a pure function of the snapshot and the context: no clock, no randomness, no I/O. Write it as a specification, then implement it identically in every SDK and test every implementation against a shared suite of golden cases. The order below is a common and defensible one.
import hashlib
BUCKETS = 10_000
def bucket(flag_key, salt, value):
# Deterministic: same (flag, salt, value) -> same bucket in every language.
digest = hashlib.sha256(f"{flag_key}.{salt}.{value}".encode()).digest()
return int.from_bytes(digest[:8], "big") % BUCKETS
def evaluate(snapshot, key, ctx, depth=0):
flag = snapshot["flags"].get(key)
if flag is None:
return None, "FLAG_NOT_FOUND" # caller's default wins
if depth > 5:
return flag["offVariation"], "PREREQ_CYCLE"
if not flag["on"]:
return flag["offVariation"], "OFF"
for pre in flag.get("prerequisites", []):
got, _ = evaluate(snapshot, pre["flag"], ctx, depth + 1)
if got != pre["variation"]:
return flag["offVariation"], "PREREQ_FAILED"
for t in flag.get("targets", []):
if ctx.get("userId") in t["values"]:
return t["variation"], "TARGET_MATCH"
for rule in flag.get("rules", []):
if all(match(snapshot, c, ctx) for c in rule["clauses"]):
return serve(flag, key, rule["serve"], ctx), "RULE_MATCH:" + rule["id"]
return serve(flag, key, flag["fallthrough"], ctx), "FALLTHROUGH"
def serve(flag, key, spec, ctx):
if "variation" in spec:
return spec["variation"]
b = bucket(key, flag["salt"], ctx.get(spec.get("bucketBy", "userId"), ""))
running = 0
for part in spec["rollout"]:
running += part["weight"]
if b < running:
return part["variation"]
return spec["rollout"][-1]["variation"]Three properties fall out of this. Stickiness: a user's bucket never changes, so moving a rollout from 25 to 50 percent only adds users; nobody who had the feature loses it, as long as you only move the boundary between the two slices, as the cumulative loop in serve does when you change the weights. Explainability: the reason string (RULE_MATCH:r2) goes into exposure events and debug logs, which turns "why did this user see the new checkout?" into a lookup. Safety: a missing flag returns the caller's coded default, so a deleted flag degrades to known behaviour rather than an exception. Bucket by a stable identifier. Bucketing by session id makes users flip between variations, and bucketing a B2B feature by user rather than account splits one customer's team across both experiences.
Propagation: snapshot, stream, and gap detection
An SDK starts by loading a snapshot, then applies deltas in version order. If it ever sees a gap, it throws away its incremental state and reloads a snapshot. That single rule keeps every process convergent without the relay needing per-client state.
class FlagStore:
def __init__(self, cdn, relay, env, cache_path):
self.cdn, self.relay, self.env, self.cache_path = cdn, relay, env, cache_path
self.snapshot, self.version = None, -1
def start(self):
try:
self.install(self.cdn.get_latest(self.env)) # immutable, cacheable
except Exception:
self.install(load_json(self.cache_path)) # last good copy on disk
self.relay.subscribe(self.env, since=self.version, on_event=self.on_delta)
def on_delta(self, d):
if d["version"] <= self.version:
return # duplicate, ignore
if d["version"] != self.version + 1:
self.install(self.cdn.get_version(self.env, d["version"])) # gap: resync
return
new = apply_patch(self.snapshot, d["patch"]) # copy-on-write
self.snapshot, self.version = new, d["version"] # one atomic swap
def install(self, snap):
self.snapshot, self.version = snap, snap["version"]
save_json(self.cache_path, snap)
metrics.gauge("flags.config_version", self.version)Readers never lock: evaluation reads whatever snapshot reference is current, and an update replaces the reference in one assignment. Every process exports its config version as a metric, so a dashboard showing the version distribution across the fleet answers "has the kill switch landed everywhere?" directly. Relays reconnect clients with jittered back-off after a relay restart, otherwise 60,000 SDKs reconnect in the same second. The snapshot on the CDN is the bootstrap path and the resync path; putting it on a CDN means a mass restart during an incident reads from edge caches rather than from the flag service.
The write path: concurrency, approvals and audit
Writes use optimistic concurrency. The client sends the flag version it edited; the service rejects the write with a conflict if the stored version moved. Without this, two people editing targeting rules at the same time silently lose one change. Store changes as semantic patches ("add rule r3 at position 2") rather than whole documents, which makes diffs readable in review and makes the audit log useful.
Production changes go through change requests: a proposed patch, a required approver who is not the author, and an optional schedule. The kill switch path is the deliberate exception. Turning a flag off, or serving its off variation, should need no approval, because the person responding to an incident at 3 a.m. must not wait for a reviewer. Log it loudly instead. Scheduled changes, such as a launch at 09:00 in a given time zone, are executed by the flag service writing an ordinary versioned change at that time; the SDKs never evaluate clocks.
Every audit row records actor, time, environment, flag, old and new version, the patch and a reason. Show it on the same timeline as deploys: flags are changes that no deploy log will show.
Client-side flags are a different product
Never ship the full rule set to browsers and mobile apps: rules reveal unreleased features, segment lists can contain personal data, and anyone can modify the client. Instead, the client sends its context to an edge evaluator, which runs the same code and returns only that user's variations, cached for offline use. Client-side flags are therefore advisory. Anything that matters for security or billing, such as entitlements, must be enforced again on the server.
Failure modes
| Failure | What users see | Design response |
|---|---|---|
| Flag service or database down | Nothing, if designed well | SDKs keep last good snapshot; writes queue in the UI; kill switch has a break-glass path |
| Relay fleet down | Changes stop propagating | SDKs fall back to polling the CDN snapshot; alert on fleet version skew |
| Bad targeting change | Wrong users get a feature | Approval, diff preview with estimated audience size, one-click revert to prior version |
| Cold start with no network | Service boots with defaults | Snapshot baked into the image or persisted on disk; coded defaults must be safe |
| SDK evaluation drift between languages | Same user sees different variations | Shared golden test suite run in every SDK's CI |
| Exposure event flood | Analytics lag, cost spikes | Sample and deduplicate per user per flag per hour in the SDK |
| Flag debt | Nobody knows if a flag can be removed | Owner and expiry required at creation; detector flags ones serving one variation for 30 days |
Worked example: rolling out a new checkout
The team ships the new checkout behind new-checkout, off in production. Day 1, the employees rule serves it; internal users find two layout bugs. Day 3, an approved change request adds a 1 percent rollout. Guardrail metrics, payment errors and checkout p95 latency, are compared between the two groups rather than against yesterday, which removes day-of-week noise.
Day 5, the rollout moves to 5 percent and then 25 percent. At 25 percent the payment error rate in the treatment group rises from 0.4 to 0.9 percent for one card network. The on-call engineer flips the flag's rule to serve the old checkout; the fleet version dashboard shows 99 percent of processes on the new version within three seconds and all of them within eight. Because bucketing is sticky, after the fix is deployed the rollout resumes at 25 percent with exactly the same users, so the before and after comparison is clean. Day 12 reaches 100 percent, and day 40 removes the flag from the code and archives it, closing the ticket the stale-flag detector opened.
Trade-offs to decide explicitly
- Build or buy. The evaluation engine is small; SDKs in six languages with identical semantics, the relay fleet, approvals and experiment analysis are not. Standardise on a vendor-neutral SDK interface such as OpenFeature, a CNCF project, so the backend can change later.
- Streaming or polling. Streaming gives seconds of latency but long-lived connections to operate; polling is simpler and slower.
- One flag system or several. Release toggles, experiments, operational switches and entitlements have different lifetimes and owners. Keep one engine, but tag the flag type and apply different policies: release flags expire, entitlements never do and live closer to billing.
- Rich targeting or cheap evaluation. Regex clauses and huge segment lists slow evaluation and bloat snapshots; cap rule complexity.
- Automatic rollback. Tying rollouts to guardrail metrics and burn-rate alerts catches regressions faster than people do, but a noisy metric will roll back good launches. Start with alert-and-page, then automate for the metrics you trust.
What to do next
- Write the evaluation specification and a golden test suite of contexts and expected variations before writing any SDK.
- Store flags with per-environment versions and write every change, outbox row and audit row in one transaction.
- Publish immutable per-version snapshots to object storage and stream deltas with monotonic versions; resync on any gap.
- Persist the last good snapshot locally in every SDK and make coded defaults safe for a cold start with no network.
- Export each process's config version as a metric and build the fleet version-skew dashboard before the first incident.
- Give kill switches a no-approval path, require owners and expiry dates on release flags, and run a stale-flag detector weekly.