advanced 2 min answer

Eight product teams each call a shared entitlements service on every request. A four-person platform team owns it. It measured 99.95% availability over the last year and answers in 8 ms at p99. Leadership now wants feature-flag evaluation moved into the same service because it would be one place to change. What does that centralisation buy, and when does the bill arrive?

centralisationavailabilityfeature-flagsentitlementscoupling
Show the full answer Hide the answer

What is gained, quantified

One implementation of a rule is one place to fix it. A permission semantics change that costs 3 engineer-days per team costs 24 engineer-days across eight teams and 3 days centrally. An auditor's question becomes one query instead of eight investigations. For entitlements that argument is decisive, because a wrong answer is a security incident and eight slightly different implementations guarantee one of them is wrong.

What is paid

Availability composes in series. A dependency measured at 99.95% is about 4.4 hours a year, and every consumer in its synchronous path inherits that ceiling before adding its own failures. Eight products at 99.9% of their own behind a 99.95% dependency land near 99.85%, and all eight degrade in the same minute rather than independently. The 8 ms at p99 is also now on every request in the company.

The second payment is the queue. Four people become the serialisation point for eight roadmaps, and the honest measure of that is how long a consumer waits for a change it could have made itself.

Which side wins at these numbers

Centralise entitlements. Do not centralise flag evaluation into the same request path. The distinction is not importance, it is what a wrong or missing answer costs:

  • Entitlements have no safe default. Guessing "allowed" is a breach and guessing "denied" is an outage, so the decision must be authoritative and must be one implementation.
  • A feature flag has a safe default: the previous behaviour. So centralise the definition and distribute the evaluation — publish a signed snapshot of flag values that each service caches and evaluates locally, refreshed every few seconds. One place to change, no new dependency in the request path, and a flag-service outage freezes flags rather than failing requests.

What would have to change to flip it

If a flag gates money movement or a legal disclosure, its evaluation is a compliance decision and belongs with entitlements, authoritative and logged per evaluation. Conversely, if the entitlements answer can be cached for 30 seconds without breaking a contractual requirement, push it into a library or sidecar reading the same snapshot and the availability coupling disappears. The test is whether a stale answer is merely old or actually wrong.

When the bill arrives

At the platform team's first bad deploy, when eight products degrade together and each one's own dashboards look fine. Then at the first audit, when nobody can reconstruct which flag was active for which tenant on the day in question, because local evaluation was added without an evaluation log. Budget for both before centralising: a per-evaluation log with the snapshot version, and a documented degraded mode in every consumer that says what happens when the shared service is unreachable.

When this is the wrong answer

With two teams rather than eight, a shared service is premature: the duplicated rule costs 6 engineer-days a year and the coupling costs an availability ceiling plus a platform team you do not have. Duplicate deliberately, note the trigger, and centralise when the fourth consumer appears.