Review this design. A streaming platform resolves three things in an edge function on every playback request: the active feature-flag set written from CI and read 200000 times a second; the regional price shown in the paywall; and whether this account's subscription is still active. All three live in a globally replicated edge key-value store with roughly 10-second propagation, writable from any location with last-writer-wins. What would you remove, what would you change, and what would you leave alone?
Show the full answer Hide the answer
What is actually required
Three different requirements have been given one mechanism:
- Playback start must not wait on a cross-region read. True, and the reason the edge store is here.
- A cancelled or expired subscription must stop granting new playback quickly, with "quickly" being a number someone in finance or licensing can live with.
- A displayed price must match the price charged. Not "be fresh" - match.
What I would remove
The write-from-anywhere path. Each of these three has exactly one authoritative writer: CI for flags, the pricing service for prices, billing for subscription state. Multi-master writes buy nothing and introduce an unordered-write hazard: a price corrected at 10:00:01 from one location can be overwritten by a delayed retry of the old value from another, and last-writer-wins will call that resolution. Make the edge store read-only at the edge and write it from one place. Read local, write global, and the write path has one door.
I would also remove the subscription check from the edge store entirely, for the reason below.
The one change that matters
Replicate the policy, not the decision. The central authorisation service issues a short-lived signed playback token - five to fifteen minutes, scoped to one title - and the edge verifies the signature with no state at all. Revocation latency becomes the token lifetime: a number you choose, can state in a contract and can shorten under pressure, instead of "about 10 seconds, except when a location's replication is behind and nobody is watching that metric".
The rule underneath it: staleness is not symmetric. A stale deny is safe, because the client retries and the central service decides. A stale allow is lost revenue and a licensing breach. Arrange every money-bearing or rights-bearing decision so that the failure of your freshness mechanism can only fail closed.
What I would leave alone, though it looks odd
- Flags in the edge store with 10-second propagation. A value read 200,000 times a second and written a few times a day is the textbook case for replication, and 10 seconds is fine - unless a flag is also the kill switch for a bad rollout. If it is, it needs a stated propagation objective and a dashboard showing the flag-set version per location, because "we flipped it" and "it is in effect everywhere" are different events.
- Regional prices at the edge. Prices change on a schedule, not per request. Keep a version-stamped price blob, and have checkout re-validate the version server-side so a stale display is caught before the charge rather than after it. That turns a consistency problem into a validation problem, which is much cheaper.
When not to have an edge layer at all
If playback starts are 2,000 per second rather than 200,000, and the audience sits in two regions, a plain regional round trip beats this entire design. A 30 ms in-region call costs less than a replication system whose failure modes you must learn, monitor and explain. The edge earns its keep when request volume is high, the data is read far more than written, and staleness is provably safe in the direction it can fail. Two of these three fields fail that test as currently designed.
How I would argue it in the review
Not as "10 seconds is too slow". As a question with three answers: which of these decisions is allowed to be wrong for 10 seconds, and who pays when it is? Flags, the product team. Prices, finance, and only until checkout validates. Entitlement, the rights holder, and that one is not ours to spend.