intermediate 2 min answer

Your codebase has 340 feature flags. Nobody knows which are live. A recent incident was caused by an untested flag combination. How do you get out of this?

feature-flagstechnical-debtcomplexity
Show the full answer Hide the answer

What the interviewer is testing

Whether you understand that flags are branches in the code with combinatorial consequences, and whether you can propose a remediation that will actually be executed.

Why this became an incident

Every flag is a conditional. 340 flags is 2^340 theoretical paths, of which the team has tested a handful. The incident came from a combination nobody imagined being reachable — typically two flags that were each safe alone and interacted through shared state.

The root cause is not the number of flags but the absence of a lifecycle. Release flags are meant to be temporary and were never removed.

The remediation

Classify first, because the four kinds need different treatment:

Kind Lifetime Action
Release Days to weeks Remove after rollout — this is the debt
Experiment Duration of the test Remove when the test concludes
Operational (kill switches, load shed) Permanent Keep, document, test
Permission / entitlement Permanent Not really a flag — move to configuration or authorisation

Instrument to find the truth. Evaluate-and-log so you know which flags are actually consulted, what value they return, and whether any code path is dead. A large proportion will turn out to be permanently on or permanently off, and those are trivially removable.

Remove in batches with a deadline, prioritising flags that are stale and fully on or off — the lowest-risk, highest-count group.

Then prevent recurrence: every release flag is created with a mandatory expiry date, the build warns at expiry and fails some period after, and flag creation records an owner. This is the part that matters, because a one-off cleanup without it returns to 340 within two years.

What a strong answer adds

Recognising the availability risk: flags evaluated in the request path make the flag service a hard dependency, so it needs local caching with a safe default and must never fail the request. And noting that flag state should be visible in incident tooling — during an outage, "what changed" includes flag flips, which are invisible to deployment history.

Common weak answers

"Delete them all", which will break production. Adding tests for flag combinations, which does not scale and treats the symptom.