advanced 3 min answer

The platform is removing v1 of its config API. Telemetry showed 412 calling services at announcement; after six months of a funded migration 9 remain - 4 have no owner in the catalogue, 3 are monthly batch jobs that touch the API only on the first of the month, and 2 belong to a team that says it will not move this quarter. Removal was promised for the end of the quarter. Sequence the removal so you keep the date and stay reversible at every step.

deprecationbrownoutsunsetreversibilityownership
Show the full answer Hide the answer

What is being tested

Whether you can land the last 2% of a deprecation. The first 98% is a communications and tooling exercise; the tail is an ownership and reversibility exercise, and it is where internal deprecations go to die. Python 2 is the public illustration: the end-of-life date moved from 2015 to 2020 because the tail, not the plan, set the date.

The sequence, each step reversible

  1. Split the fleet by identity, not by intention. Every caller already authenticates, so move the 9 known callers onto an explicit allowlist and make everything else receive 410 Gone at the edge. This is a config change, reversible in one deploy, and it stops the count going back up. Without it, a team re-introduces a v1 call in week three and your 9 becomes 11.
  2. Put a translation shim behind the same hostname. v1 requests are rewritten onto v2 by a small proxy the platform owns. The shim is the artifact that makes the date keepable, because removing the v1 implementation is now separable from removing the v1 endpoint.
  3. Brownout on the callers' clock, not yours. Two scheduled 30-minute windows returning 503 with Deprecation and Sunset headers (RFC 8594). One window must overlap the first of the month, or the 3 batch jobs never experience the brownout and will discover removal in production. A brownout reveals what telemetry cannot: which callers retry silently, which cache the last good response, and which have no alerting at all.
  4. Resolve the unowned 4 as a security item, not a platform item. A service with no owner and no deploy in nine months is a running credential nobody is maintaining. Route it to whoever owns credential hygiene: either it is adopted within two weeks or it is turned off. The platform should not be negotiating with absent owners.
  5. For the 2 unwilling callers, the platform writes the patch. At this point it is cheaper for one platform engineer to open two pull requests than to hold the migration open for a quarter. Offer the PR, with a stated date after which the allowlist entry expires.

Where data can diverge, and how you would know

While the shim is live, config can be written through both paths. Run a reconciliation job every 15 minutes comparing a checksum of each tenant's config as seen through v1 and v2, and alert on any non-zero diff. Without that job, a translation bug shows up weeks later as one tenant's traffic routed to the wrong cluster, and nobody connects it to the deprecation.

The point of no return

Not the brownout and not the endpoint removal. It is deleting the v1 storage schema or revoking the v1 credentials' permissions, because both are hard to undo under incident pressure. Keep the shim for one full quarter after the last caller leaves, and delete v1 code only after one complete monthly batch cycle has run clean.

How long it really takes

Plan the tail at the same duration as the bulk. 412 to 9 took six months; 9 to 0 takes six to ten weeks of mostly non-engineering work, and the single biggest accelerator is funding the migration out of the platform's budget rather than asking.

When this is the wrong answer

If one of the 9 is on a revenue-critical path where a 30-minute 503 is a customer-visible outage, do not brownout it; buy the shim permanently. A translation proxy costs on the order of one engineer-week a year to keep alive, which is cheaper than an incident and far cheaper than a deprecation that slips twice and teaches the organisation that platform dates are fiction.