A feature flag is a runtime decision point: code that can take one of several paths depending on configuration that changes without a deploy. Flags let you merge unfinished work, release to one percent of users, run experiments and switch off a misbehaving dependency in seconds. They also accumulate. A codebase with three hundred flags, half of them fully rolled out years ago, has three hundred untested branches and nobody sure which ones are safe to delete.
Most writing about flags is about the flag service: the control plane, the rule language, how rules reach servers. That is covered here in designing a feature flag system and feature flag delivery architecture. This article is about the other half, the part your application team owns: how flags enter code, how they are typed, evaluated, defaulted, observed, tested and removed. It uses OpenFeature, the vendor-neutral CNCF specification for flag evaluation, and its Python SDK, with names checked against the specification and SDK source.
The layers between a call site and a rule
Each layer in the diagram exists to keep a kind of change local. The call site asks a business question, such as whether the new pricing engine is on for this request. The typed registry turns that into a flag key, a type and a default, in one file. The OpenFeature client attaches the evaluation context and runs hooks. The provider is the adapter for whichever flag system you use, and usually evaluates rules from a local cache that the control plane keeps up to date.
The payoff is that switching flag vendors changes one provider registration, renaming a flag changes one registry entry, and a reviewer can find every flag in the codebase by reading one module.
Kinds of flags and their contracts
Flags differ in how long they live, who changes them and what the safe default is. Writing the kind down per flag decides most of the questions later sections raise.
| Kind | Purpose | Lifetime | Safe default when evaluation fails |
|---|---|---|---|
| Release | Hide unfinished or newly shipped code | Days to weeks; delete after full rollout | Off: the old path is the proven one |
| Kill switch (ops) | Turn off a costly or failing dependency fast | Long-lived, by design | On: an outage of the flag system must not disable the feature |
| Experiment | Assign variants and measure | Length of the experiment | Control variant, and log no exposure |
| Permission / entitlement | Gate features by plan or customer | Long-lived | Usually off; often better served by an entitlement service |
The kill-switch row is the one teams get wrong. A kill switch is phrased as "disable recommendations" and defaults to false, so when the flag system is unreachable the feature keeps running. If you phrase it as "recommendations enabled" with a default of false, a flag outage switches the feature off everywhere at once.
OpenFeature in a few calls
OpenFeature separates the API your code calls from the provider that answers. You register a provider once at startup, get a client, and ask for typed values with a default. Each evaluation can return details as well as the value: the variant, a reason, and an error code if something went wrong.
# pip install openfeature-sdk
from openfeature import api
from openfeature.evaluation_context import EvaluationContext
from openfeature.provider.in_memory_provider import InMemoryFlag, InMemoryProvider
api.set_provider(InMemoryProvider({
"new-pricing-engine": InMemoryFlag("off", {"on": True, "off": False}),
}))
client = api.get_client()
ctx = EvaluationContext(targeting_key="user-4821",
attributes={"plan": "pro", "country": "DE"})
enabled = client.get_boolean_value("new-pricing-engine", False, ctx)
details = client.get_boolean_details("new-pricing-engine", False, ctx)
print(details.value, details.variant, details.reason, details.error_code)The specification defines reasons including STATIC, DEFAULT, TARGETING_MATCH, SPLIT, CACHED, DISABLED, UNKNOWN, STALE and ERROR, and error codes including PROVIDER_NOT_READY, FLAG_NOT_FOUND, PARSE_ERROR, TYPE_MISMATCH, TARGETING_KEY_MISSING, INVALID_CONTEXT, PROVIDER_FATAL and GENERAL. Evaluation does not raise on these errors: it returns your default and reports the error code. That design keeps a flag failure from crashing a request, and it means the default you pass is a production decision, not a placeholder.
Building the evaluation context
The evaluation context is what targeting rules see. The targeting key is the identity used for percentage rollouts, so it must be stable for the thing you are rolling out to: a user id for user-facing features, an account id when a whole organisation should flip together, a device id for logged-out traffic. A key that changes between requests, such as a session id for a feature meant to be sticky per user, gives a user both experiences on alternate page loads.
Build the context once per request, in middleware, from data you already have, and pass it through. Include only attributes rules actually use. The context may leave your process when a provider evaluates remotely or logs it, so email addresses and other personal data should not be in it unless a rule needs them and your provider agreement covers them.
A typed flag registry
String keys scattered across the codebase are how flag debt starts: a typo silently evaluates a non-existent flag and returns the default forever, and nobody can list which flags exist. Instead, declare every flag once, with its type, default, kind, owner and an expiry date, and make call sites use the declarations.
from dataclasses import dataclass
from datetime import date
@dataclass(frozen=True)
class BoolFlag:
key: str
default: bool
kind: str # "release" | "kill_switch" | "experiment" | "permission"
owner: str
expires: date | None
NEW_PRICING_ENGINE = BoolFlag("new-pricing-engine", False, "release",
"team-pricing", date(2026, 11, 15))
DISABLE_RECOMMENDATIONS = BoolFlag("disable-recommendations", False, "kill_switch",
"team-discovery", None)
class Flags:
# evaluated once per request so one request never sees two answers
def __init__(self, client, ctx):
self._client, self._ctx, self._seen = client, ctx, {}
def is_on(self, flag: BoolFlag) -> bool:
if flag.key not in self._seen:
self._seen[flag.key] = self._client.get_boolean_value(
flag.key, flag.default, self._ctx)
return self._seen[flag.key]The per-request memo in Flags is deliberate. If a rule changes in the middle of a request, a handler that evaluates the same flag twice can take the new path in one function and the old path in another, writing half-migrated data. Evaluating once and reusing the answer makes each request internally consistent, and it keeps flag checks out of hot loops.
Startup, readiness and defaults
A provider that fetches rules over the network is not ready the instant the process starts. Evaluations before it is ready return the default with PROVIDER_NOT_READY. For most services the right behaviour is to wait for readiness with a short bound during startup, then serve with defaults if the bound expires, and report it. A service that refuses to start because the flag system is down has turned an optional dependency into a hard one.
Once running, providers typically keep serving the last rules they received if the control plane becomes unreachable; check what yours does and whether it reports STALE or CACHED reasons so you can alert on it. Then audit the defaults in the registry against the table above. A release flag defaulting to on means a flag outage releases unfinished code to everyone.
Hooks for exposure and health telemetry
Hooks run around every evaluation without touching call sites. Two are worth having from the start: an exposure hook that records which variant a user actually received, which experiment analysis depends on (see A/B experimentation platform architecture), and a health hook that counts evaluations by reason and error code.
from openfeature.hook import Hook
class FlagTelemetry(Hook):
def after(self, hook_context, details, hints):
metrics.incr("flag.eval", tags={"flag": details.flag_key,
"variant": str(details.variant),
"reason": str(details.reason)})
if hook_context.flag_key.startswith("exp-"):
exposures.log(key=details.flag_key, variant=details.variant,
subject=hook_context.evaluation_context.targeting_key)
def error(self, hook_context, exception, hints):
metrics.incr("flag.error", tags={"flag": hook_context.flag_key,
"error": type(exception).__name__})
api.add_hooks([FlagTelemetry()])Reason counts become the input for flag clean-up later: a flag whose evaluations have returned the same variant for every caller for a month is no longer making a decision.
Testing flagged code
Every flag doubles the number of paths through the code it guards, and you cannot test every combination of many flags. Test each flag's branches independently with all other flags at their defaults, plus the specific combinations you know interact. Use the in-memory provider so tests never talk to a real flag service.
import pytest
from openfeature import api
from openfeature.provider.in_memory_provider import InMemoryFlag, InMemoryProvider
@pytest.fixture(params=["on", "off"])
def pricing_flag(request):
api.set_provider(InMemoryProvider({
"new-pricing-engine": InMemoryFlag(request.param, {"on": True, "off": False}),
}))
return request.param == "on"
def test_quote_is_consistent(pricing_flag):
quote = price_order(order_fixture())
assert quote.total > 0
assert quote.engine == ("v2" if pricing_flag else "v1")
def test_registry_defaults_are_safe():
for f in ALL_FLAGS:
if f.kind == "kill_switch":
assert f.key.startswith("disable-") and f.default is False
if f.kind == "release":
assert f.default is False and f.expires is not None
Finding and removing stale flags
Flag debt is cheapest to remove while the author still remembers the code. Three signals catch it. The registry's expiry date: CI fails, or opens a ticket, when a release flag is past its date. Usage: a static scan finds registry entries no longer referenced and references to keys not in the registry. Behaviour: telemetry shows flags whose every evaluation returned the same variant for 30 days.
import ast, pathlib, sys
from datetime import date
def referenced_flags(src_root, registry_file="flags_registry.py"):
used = set()
for path in pathlib.Path(src_root).rglob("*.py"):
if path.name == registry_file: # declarations are not uses
continue
for node in ast.walk(ast.parse(path.read_text(encoding="utf-8"))):
if (isinstance(node, ast.Name) and isinstance(node.ctx, ast.Load)
and node.id.isupper()):
used.add(node.id)
return used
def check(registry, src_root):
used = referenced_flags(src_root)
problems = [f"{n} declared but unused" for n in registry if n not in used]
problems += [f"{n} expired {f.expires}" for n, f in registry.items()
if f.expires and f.expires < date.today()]
if problems:
print("\n".join(problems))
sys.exit(1)Removal is a code change, not a configuration change: roll the flag to its final state, wait long enough to be sure, delete the losing branch and the registry entry, deploy, and only then archive the flag in the control plane. Archiving first makes the code fall back to its default, which for a fully released feature is usually the old path.
Worked example: replacing a pricing engine
The pricing team adds NEW_PRICING_ENGINE as a release flag expiring in six weeks, plus DISABLE_PRICING_CACHE as a kill switch for the new engine's cache. The quote handler evaluates the flag once and calls either engine. Tests run both branches. The rollout goes to internal accounts, then 5 percent by account id, watched through the canary process in canary deployment. At 25 percent the cache misbehaves; on-call flips the kill switch, the engine keeps running uncached, and no deploy is needed. At 100 percent for two weeks, the stale-flag check notices every evaluation returns on, the old engine and the flag are deleted in one pull request, and the flag is archived after deploy.
Failure modes
| Failure | What happens | Defence |
|---|---|---|
| Typo in a string key | FLAG_NOT_FOUND, default forever | Typed registry; alert on FLAG_NOT_FOUND |
| Unstable targeting key | Users flip between variants | Key on user or account id, built once in middleware |
| Flag read twice per request | Half-old, half-new behaviour | Evaluate once per request and reuse |
| Kill switch with an inverted default | Flag outage disables the feature | Name it disable-..., default false, test it |
| Archived before code removal | Code silently reverts to old path | Remove code first, archive last |
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| API layer | OpenFeature: portable, small indirection cost | Vendor SDK directly: every feature, harder to leave |
| Evaluation timing | Once per request: consistent, ignores mid-request changes | Per call: fastest reaction, inconsistent requests |
| Startup | Wait for readiness: correct first requests, slower boot | Serve defaults immediately: fast boot, first requests may differ |
What to do next
- List every flag key in your codebase and move them into one typed registry with kind, owner and expiry.
- Check each kill switch's naming and default so a flag-system outage leaves features running.
- Put OpenFeature between call sites and your vendor SDK, and build the evaluation context once in middleware.
- Add a telemetry hook that counts evaluations by reason and error code, and alert on FLAG_NOT_FOUND and PROVIDER_NOT_READY.
- Use InMemoryProvider in tests and run both branches of every release flag.
- Add the stale-flag check to CI and remove one expired flag this week to prove the procedure.