A feature flag can make a legacy release safer, but only if the on-call engineer can predict what happens when the flag service fails. I would introduce flags at one existing decision point, leave the old path as the default, and establish a tested disable procedure before adding a shared flag framework. A broad instrumentation pass creates more switches than production can safely operate.
The first flag belongs at a request boundary, not across the codebase
Pick a change that already has a clear choice between old and new behavior: selecting a checkout handler, choosing a query implementation, or routing a job to a replacement worker. Put the flag immediately before that choice. Do not thread flag checks through every helper the new path calls, because a scattered flag makes it hard to determine which behavior a disable action will actually restore.
I agree with Add feature flags to legacy code without rewriting everything about keeping the insertion small, but I would make the first change smaller still: one entry point, one named owner, and one documented fallback. A new flag should not silently substitute for a missing configuration value; its safe default should select the existing path until an operator deliberately changes it.
Before editing that entry point, trace its side effects. If both paths can send an email, charge a card, or write to PostgreSQL, switching between them mid-request may duplicate work. Keep the decision stable for the entire request or job, and use an existing idempotency key where retries can repeat a side effect. For a background job, record the chosen path in the job payload or execution record if a retry might run after the flag changes. That small piece of persistence is more valuable than a sophisticated targeting rule when an operator is investigating a duplicate action.
Give the flag a contract in the repository: its name, old and new paths, default, on-call owner, expected lifetime, and disable command. Add a test that forces each path and another that exercises missing or malformed configuration. This is not a rewrite of legacy tests; it is a narrow check that the flag can return traffic to known behavior. If the old path is already unsafe, do not label it a fallback simply because it is familiar. In that case, the release needs a different recovery plan, such as stopping the affected job.
A local file can prove the seam, but it is not a fast kill switch
There are two reasonable first implementations. A mounted Kubernetes ConfigMap wins when the initial flag is changed rarely and the team already deploys through Kubernetes, because it adds no flag service to operate. It costs an update delay that must be measured in the cluster, and a ConfigMap mounted with subPath will not receive subsequent updates. OpenFeature with a flagd provider wins when operators need targeting and a centrally managed runtime change, because application code can evaluate a flag without owning the transport. It costs a provider deployment, connection monitoring, and an explicit policy for stale or unavailable evaluations.
The following Python 3.11 example demonstrates the local seam. It runs even when the file does not exist, returning the old path. In production, write a complete JSON file and replace it atomically; do not edit the file in place while requests may read it.
import hashlib
import json
from pathlib import Path
def enabled(account_id: str) -> bool:
try:
cfg = json.loads(Path("/etc/app/flags.json").read_text())
pct = int(cfg.get("new_checkout_percent", 0))
if not 0 <= pct <= 100: return False
digest = hashlib.sha256(account_id.encode("utf-8")).digest()
return int.from_bytes(digest[:8], "big") % 100 < pct
except (OSError, ValueError, TypeError, AttributeError):
return False
print(enabled("account-42"))
SHA-256 makes the sample account consistently fall into the same percentage bucket, which avoids moving a user between implementations on every request. The example reads a file on each evaluation to make failure behavior visible; that disk read may be too expensive on a hot path. Cache a validated snapshot in a real service, then measure how long an update takes to reach every process. Set a provisional 30-second propagation objective only if that delay is acceptable for this flag; it is an operational value to tune, not a property guaranteed by ConfigMap.
For a genuine emergency disable, test the full route from the operator’s action to the application’s next evaluation. If the route relies on a Kubernetes rollout, call it a deployment-time switch rather than promising an instant kill switch. If it relies on flagd or another remote provider, specify what happens when the provider is unreachable. Returning the old path is sensible for a reversible checkout implementation; it may be wrong for a flag that prevents a dangerous operation, because “old” and “safe” are not always synonyms.
A rollout is observable only when its two paths can be separated
Start with a cohort that is small enough to limit harm but large enough to produce useful traffic. An initial 1% exposure is a value to tune against request volume, not a universal safe percentage. Move to a larger cohort only after comparing new-path and old-path outcomes over the same period. A global success rate can hide a regression when the new path serves a small share of requests.
Instrument the decision at the boundary. A Prometheus counter such as checkout_requests_total{path=”new”,result=”error”} lets an operator compare errors by path; a latency histogram offers the corresponding view of slow requests. Keep labels bounded: use the path name, not account IDs or flag payloads, because high-cardinality series can make monitoring expensive and slow. In OpenTelemetry traces, record the selected path as a low-cardinality attribute so an incident responder can follow a failing request without guessing which implementation handled it.
Choose the observation window from actual traffic. A locally chosen five-minute window may be useful for a busy endpoint but nearly meaningless for a job that runs hourly. Prometheus documents 1m as its default scrape_interval; configure scrape_interval: 15s if faster sampling is needed, then account for scrape and alert-evaluation delay when setting the disable procedure. Faster scrapes alone do not create faster detection if the alert waits for a long window or the endpoint receives little traffic.
Write the stop condition before raising exposure. For example, decide what increase in new-path error rate warrants a pause, who can make that call, and whether an Argo Rollouts analysis will halt promotion or a human will change the flag. Do not automatically disable on a single low-volume error, because that rule can oscillate between paths without establishing which one is healthier. Do not rely solely on Grafana screenshots either; an operator needs a recorded flag value, timestamp, and outcome to reconstruct a rollout after the dashboard window has passed.
Exercise the disable path in a staging environment and once under a controlled production rollout. Record the elapsed time from changing the flag to observing only old-path decisions. That measured interval, rather than the flag vendor’s update claim or the configuration file’s modification time, is the number the on-call runbook should use.
Ownership must include the person who can stop traffic
I disagree with At 30 engineers, CI and code ownership fail before runtime on sequencing for this migration: I would establish runtime disablement before automating flag governance, because a bad flag change can affect users without waiting for the next pull request. GitHub CODEOWNERS can require a reviewer for changes to flag definitions, but it cannot by itself identify who is awake and authorized to turn one off during an incident.
Keep a short registry beside the code, with a link to the on-call rota and the precise control-plane action for each flag. Separate permission to change targeting from permission to edit application code. If a provider such as LaunchDarkly supplies an audit log, retain the actor, old value, new value, and time; if the team uses a ConfigMap, preserve equivalent change history through the GitOps repository and its deployment records. The requirement is the same in either case: an operator must be able to answer who changed production behavior without searching chat messages.
Use GitHub Actions to check what automation can check cheaply: every new flag has an owner, a default, a removal date, and tests for both branches. Keep authorization for production flag changes in the control plane rather than assuming CI approval covers it. A flag changed through an administrative interface bypasses the pull request that CODEOWNERS reviewed, so the on-call runbook must describe that separate route.
Schedule removal when the rollout starts, not after everyone has forgotten the old path. Once the new implementation has held the intended exposure through an agreed observation period, open the deletion change: remove the flag check, old branch, targeting rule, and path-specific alert only after confirming that no jobs or retries still depend on the old choice. A permanent flag with two aging implementations raises the cost of every future incident, because responders must keep reasoning about behavior that was meant to be temporary.
The first production change should be a disable drill
Choose one existing request boundary this week and write down its current behavior, owner, and side effects. Add a default-off decision there, then have someone other than its author disable it while a controlled cohort is active. Time the change until old-path requests appear in telemetry. If that drill cannot be completed confidently, improve the control path before adding another flag.


