A feature flag is a runtime decision point: code that can take one of several paths depending on configuration that changes without a deploy. Flags let you merge unfinished work, release to one percent of users, run experiments and switch off a misbehaving dependency in seconds. They also accumulate. A codebase with three hundred flags, half of them fully rolled out years ago, has three hundred untested branches and nobody sure which ones are safe to delete.

Most writing about flags is about the flag service: the control plane, the rule language, how rules reach servers. That is covered here in designing a feature flag system and feature flag delivery architecture. This article is about the other half, the part your application team owns: how flags enter code, how they are typed, evaluated, defaulted, observed, tested and removed. It uses OpenFeature, the vendor-neutral CNCF specification for flag evaluation, and its Python SDK, with names checked against the specification and SDK source.

Advertisement

The layers between a call site and a rule

Each layer in the diagram exists to keep a kind of change local. The call site asks a business question, such as whether the new pricing engine is on for this request. The typed registry turns that into a flag key, a type and a default, in one file. The OpenFeature client attaches the evaluation context and runs hooks. The provider is the adapter for whichever flag system you use, and usually evaluates rules from a local cache that the control plane keeps up to date.

The payoff is that switching flag vendors changes one provider registration, renaming a flag changes one registry entry, and a reviewer can find every flag in the codebase by reading one module.

Feature flags inside a service: call sites never see a vendor SDK or a string keyCall siteif flags.new_pricingTyped registrykey, type, default, ownerOpenFeature clientcontext, hooksProvidervendor or in-houseRules cachelocal evaluationHooksexposure, metricsRequest contexttargeting key, attrsControl planerules, auditstreamAnalyticsexposures, reasonsCI: registry tests, stale-flag scanexpiry dates, unused keys, both branchesThe provider and control plane are swappable; the registry and call sites are what your team maintains
The call site depends only on the registry. The provider can be swapped without touching business code, and CI checks the registry rather than grepping for strings.

Kinds of flags and their contracts

Flags differ in how long they live, who changes them and what the safe default is. Writing the kind down per flag decides most of the questions later sections raise.

KindPurposeLifetimeSafe default when evaluation fails
ReleaseHide unfinished or newly shipped codeDays to weeks; delete after full rolloutOff: the old path is the proven one
Kill switch (ops)Turn off a costly or failing dependency fastLong-lived, by designOn: an outage of the flag system must not disable the feature
ExperimentAssign variants and measureLength of the experimentControl variant, and log no exposure
Permission / entitlementGate features by plan or customerLong-livedUsually off; often better served by an entitlement service

The kill-switch row is the one teams get wrong. A kill switch is phrased as "disable recommendations" and defaults to false, so when the flag system is unreachable the feature keeps running. If you phrase it as "recommendations enabled" with a default of false, a flag outage switches the feature off everywhere at once.

Advertisement

OpenFeature in a few calls

OpenFeature separates the API your code calls from the provider that answers. You register a provider once at startup, get a client, and ask for typed values with a default. Each evaluation can return details as well as the value: the variant, a reason, and an error code if something went wrong.

# pip install openfeature-sdk
from openfeature import api
from openfeature.evaluation_context import EvaluationContext
from openfeature.provider.in_memory_provider import InMemoryFlag, InMemoryProvider

api.set_provider(InMemoryProvider({
    "new-pricing-engine": InMemoryFlag("off", {"on": True, "off": False}),
}))
client = api.get_client()

ctx = EvaluationContext(targeting_key="user-4821",
                        attributes={"plan": "pro", "country": "DE"})

enabled = client.get_boolean_value("new-pricing-engine", False, ctx)

details = client.get_boolean_details("new-pricing-engine", False, ctx)
print(details.value, details.variant, details.reason, details.error_code)

The specification defines reasons including STATIC, DEFAULT, TARGETING_MATCH, SPLIT, CACHED, DISABLED, UNKNOWN, STALE and ERROR, and error codes including PROVIDER_NOT_READY, FLAG_NOT_FOUND, PARSE_ERROR, TYPE_MISMATCH, TARGETING_KEY_MISSING, INVALID_CONTEXT, PROVIDER_FATAL and GENERAL. Evaluation does not raise on these errors: it returns your default and reports the error code. That design keeps a flag failure from crashing a request, and it means the default you pass is a production decision, not a placeholder.

Building the evaluation context

The evaluation context is what targeting rules see. The targeting key is the identity used for percentage rollouts, so it must be stable for the thing you are rolling out to: a user id for user-facing features, an account id when a whole organisation should flip together, a device id for logged-out traffic. A key that changes between requests, such as a session id for a feature meant to be sticky per user, gives a user both experiences on alternate page loads.

Build the context once per request, in middleware, from data you already have, and pass it through. Include only attributes rules actually use. The context may leave your process when a provider evaluates remotely or logs it, so email addresses and other personal data should not be in it unless a rule needs them and your provider agreement covers them.

A typed flag registry

String keys scattered across the codebase are how flag debt starts: a typo silently evaluates a non-existent flag and returns the default forever, and nobody can list which flags exist. Instead, declare every flag once, with its type, default, kind, owner and an expiry date, and make call sites use the declarations.

from dataclasses import dataclass
from datetime import date

@dataclass(frozen=True)
class BoolFlag:
    key: str
    default: bool
    kind: str          # "release" | "kill_switch" | "experiment" | "permission"
    owner: str
    expires: date | None

NEW_PRICING_ENGINE = BoolFlag("new-pricing-engine", False, "release",
                              "team-pricing", date(2026, 11, 15))
DISABLE_RECOMMENDATIONS = BoolFlag("disable-recommendations", False, "kill_switch",
                                   "team-discovery", None)

class Flags:
    # evaluated once per request so one request never sees two answers
    def __init__(self, client, ctx):
        self._client, self._ctx, self._seen = client, ctx, {}
    def is_on(self, flag: BoolFlag) -> bool:
        if flag.key not in self._seen:
            self._seen[flag.key] = self._client.get_boolean_value(
                flag.key, flag.default, self._ctx)
        return self._seen[flag.key]

The per-request memo in Flags is deliberate. If a rule changes in the middle of a request, a handler that evaluates the same flag twice can take the new path in one function and the old path in another, writing half-migrated data. Evaluating once and reusing the answer makes each request internally consistent, and it keeps flag checks out of hot loops.

Startup, readiness and defaults

A provider that fetches rules over the network is not ready the instant the process starts. Evaluations before it is ready return the default with PROVIDER_NOT_READY. For most services the right behaviour is to wait for readiness with a short bound during startup, then serve with defaults if the bound expires, and report it. A service that refuses to start because the flag system is down has turned an optional dependency into a hard one.

Once running, providers typically keep serving the last rules they received if the control plane becomes unreachable; check what yours does and whether it reports STALE or CACHED reasons so you can alert on it. Then audit the defaults in the registry against the table above. A release flag defaulting to on means a flag outage releases unfinished code to everyone.

Hooks for exposure and health telemetry

Hooks run around every evaluation without touching call sites. Two are worth having from the start: an exposure hook that records which variant a user actually received, which experiment analysis depends on (see A/B experimentation platform architecture), and a health hook that counts evaluations by reason and error code.

from openfeature.hook import Hook

class FlagTelemetry(Hook):
    def after(self, hook_context, details, hints):
        metrics.incr("flag.eval", tags={"flag": details.flag_key,
                                         "variant": str(details.variant),
                                         "reason": str(details.reason)})
        if hook_context.flag_key.startswith("exp-"):
            exposures.log(key=details.flag_key, variant=details.variant,
                          subject=hook_context.evaluation_context.targeting_key)

    def error(self, hook_context, exception, hints):
        metrics.incr("flag.error", tags={"flag": hook_context.flag_key,
                                          "error": type(exception).__name__})

api.add_hooks([FlagTelemetry()])

Reason counts become the input for flag clean-up later: a flag whose evaluations have returned the same variant for every caller for a month is no longer making a decision.

Testing flagged code

Every flag doubles the number of paths through the code it guards, and you cannot test every combination of many flags. Test each flag's branches independently with all other flags at their defaults, plus the specific combinations you know interact. Use the in-memory provider so tests never talk to a real flag service.

import pytest
from openfeature import api
from openfeature.provider.in_memory_provider import InMemoryFlag, InMemoryProvider

@pytest.fixture(params=["on", "off"])
def pricing_flag(request):
    api.set_provider(InMemoryProvider({
        "new-pricing-engine": InMemoryFlag(request.param, {"on": True, "off": False}),
    }))
    return request.param == "on"

def test_quote_is_consistent(pricing_flag):
    quote = price_order(order_fixture())
    assert quote.total > 0
    assert quote.engine == ("v2" if pricing_flag else "v1")

def test_registry_defaults_are_safe():
    for f in ALL_FLAGS:
        if f.kind == "kill_switch":
            assert f.key.startswith("disable-") and f.default is False
        if f.kind == "release":
            assert f.default is False and f.expires is not None

Finding and removing stale flags

Flag debt is cheapest to remove while the author still remembers the code. Three signals catch it. The registry's expiry date: CI fails, or opens a ticket, when a release flag is past its date. Usage: a static scan finds registry entries no longer referenced and references to keys not in the registry. Behaviour: telemetry shows flags whose every evaluation returned the same variant for 30 days.

import ast, pathlib, sys
from datetime import date

def referenced_flags(src_root, registry_file="flags_registry.py"):
    used = set()
    for path in pathlib.Path(src_root).rglob("*.py"):
        if path.name == registry_file:       # declarations are not uses
            continue
        for node in ast.walk(ast.parse(path.read_text(encoding="utf-8"))):
            if (isinstance(node, ast.Name) and isinstance(node.ctx, ast.Load)
                    and node.id.isupper()):
                used.add(node.id)
    return used

def check(registry, src_root):
    used = referenced_flags(src_root)
    problems = [f"{n} declared but unused" for n in registry if n not in used]
    problems += [f"{n} expired {f.expires}" for n, f in registry.items()
                 if f.expires and f.expires < date.today()]
    if problems:
        print("\n".join(problems))
        sys.exit(1)

Removal is a code change, not a configuration change: roll the flag to its final state, wait long enough to be sure, delete the losing branch and the registry entry, deploy, and only then archive the flag in the control plane. Archiving first makes the code fall back to its default, which for a fully released feature is usually the old path.

Worked example: replacing a pricing engine

The pricing team adds NEW_PRICING_ENGINE as a release flag expiring in six weeks, plus DISABLE_PRICING_CACHE as a kill switch for the new engine's cache. The quote handler evaluates the flag once and calls either engine. Tests run both branches. The rollout goes to internal accounts, then 5 percent by account id, watched through the canary process in canary deployment. At 25 percent the cache misbehaves; on-call flips the kill switch, the engine keeps running uncached, and no deploy is needed. At 100 percent for two weeks, the stale-flag check notices every evaluation returns on, the old engine and the flag are deleted in one pull request, and the flag is archived after deploy.

Failure modes

FailureWhat happensDefence
Typo in a string keyFLAG_NOT_FOUND, default foreverTyped registry; alert on FLAG_NOT_FOUND
Unstable targeting keyUsers flip between variantsKey on user or account id, built once in middleware
Flag read twice per requestHalf-old, half-new behaviourEvaluate once per request and reuse
Kill switch with an inverted defaultFlag outage disables the featureName it disable-..., default false, test it
Archived before code removalCode silently reverts to old pathRemove code first, archive last

Trade-offs

DecisionOption AOption B
API layerOpenFeature: portable, small indirection costVendor SDK directly: every feature, harder to leave
Evaluation timingOnce per request: consistent, ignores mid-request changesPer call: fastest reaction, inconsistent requests
StartupWait for readiness: correct first requests, slower bootServe defaults immediately: fast boot, first requests may differ

What to do next

  1. List every flag key in your codebase and move them into one typed registry with kind, owner and expiry.
  2. Check each kill switch's naming and default so a flag-system outage leaves features running.
  3. Put OpenFeature between call sites and your vendor SDK, and build the evaluation context once in middleware.
  4. Add a telemetry hook that counts evaluations by reason and error code, and alert on FLAG_NOT_FOUND and PROVIDER_NOT_READY.
  5. Use InMemoryProvider in tests and run both branches of every release flag.
  6. Add the stale-flag check to CI and remove one expired flag this week to prove the procedure.
Key takeaway: The flag service decides what a rule says; your application decides whether flags stay safe and removable. Route every flag through a typed registry with kind, owner, default and expiry, evaluate through OpenFeature with a stable targeting key and once per request, choose defaults per flag kind so a flag outage fails in the safe direction, observe reasons and exposures through hooks, test both branches with an in-memory provider, and delete flags in code before archiving them.