practice

Rollback Path Independence

also called Platform-Independent Revert, Second Deploy Path

The requirement that the emergency revert route not depend on the platform surface it bypasses, so an outage of the developer platform cannot also freeze every team's ability to undo a change.

internal developer platformrollbackblast radiusincident responsedependency

A platform succeeds. Within two years every deploy goes through one portal, one pipeline service and one config API, and nobody thinks about how a change reaches production.

Then the platform's control surface fails - an expired certificate on the portal, a bad pipeline release, an identity outage that stops the runners authenticating. No product is broken and nobody can change anything, including the teams whose services are broken for unrelated reasons. The platform has become a single point of failure for remediation, which is strictly worse than being one for deployment.

Why it matters

Time to restore is dominated by time to act once the cause is known, so a platform outage that removes the revert route lengthens every concurrent incident at the same time. The exposure is correlated by construction: the more successful the platform, the more incidents it can freeze at once.

The normal path also accumulates dependencies invisibly. A route that began as "push and the pipeline runs" grows a portal, a policy engine, a config service and an artefact proxy - each addition sensible, none reviewed as a dependency of incident response.

Implementation patterns

  • Write down the dependency chain of a deploy and of a revert and mark the shared links. The exercise usually surfaces two or three nobody considered.
  • Keep one route using provider-native tooling and nothing of the platform's - a direct rollout undo, a Terraform apply from a laptop against pinned state, a registry pull by digest. Slower, less safe, and it works when nothing else does.
  • Store the artefact and its rendered configuration outside the platform's database, addressed by digest, because a revert that needs platform metadata to find what to revert to is not independent.
  • Pre-provision break-glass credentials on a separate identity path, short-lived and loudly alerted on use, since credentials in the platform's own secret manager fail with it.
  • Exercise it quarterly in a game day where the portal is blocked and a team reverts a real change; the first attempt usually finds a missing permission and a tool nobody has installed.
  • Keep the route narrow: revert only.

Industry example

The pattern has a documented sibling in observability: tooling that runs on the infrastructure it monitors is unavailable exactly when its output matters, which is why mature organisations keep a minimal independent signal path. The deploy equivalent appears in published incident reports from 2019 onwards, where control-plane and identity failures left engineers unable to push a fix while the system needing it ran normally. The response that survives in production is a deliberately minimal path with separate dependencies, covering the catastrophic case rather than duplicating the platform.

Failure scenarios

  • Break-glass credentials in the platform's own secret manager, discovered during the outage that took it down.
  • A documented route naming an internal CLI that itself calls the platform API.
  • Permissions that no longer exist, because a least-privilege cleanup removed a direct-access role nobody appeared to use.
  • Reverting to an artefact the platform cannot name, because the release-to-digest mapping lives only in its database.

Trade-offs

Choose Gains Pays
A narrow independent revert path Remediation survives a platform outage A second route to exercise and a bypass to audit
Full duplication of the deploy path Everything survives Two platforms and consistency problems
Accept the dependency Nothing extra to build Every platform outage is a fleet-wide change freeze

The cost is small and recurring and the saving is concentrated in rare expensive events - the shape of cost people defer.

When not to use it

If the deploy surface is already the provider's own tooling with a thin wrapper, there is no second path to build - removing the wrapper leaves the native route intact.

It is also unnecessary where a revert is not the remediation. Prefer independence for whichever mechanism your incidents actually use; if recovery is a feature-flag flip or an edge traffic shift, the flag service is the thing that must not depend on the platform.

Interview question

Q: Your platform is the only way to deploy and its availability objective is 99.9%. A director argues a second deploy path is wasted work because the platform is more reliable than most services. What is your response, and what would you build?

What a strong answer covers: that the risk is correlation rather than availability, since one platform outage freezes remediation for every team at once including incidents it did not cause · that 99.9% is about 43 minutes a month of potential company-wide change freeze · the scope, revert only, provider-native, artefacts by digest outside the platform's database, separately-rooted credentials · a quarterly exercise with a named owner, because an untested path is a document · and loud audit on every use.

Quick check

Quiz: The portal is down and no service is broken. What is the business impact? Every team has lost the ability to change production, so any unrelated incident runs for the platform's outage plus its own - which is why the revert path must not depend on the platform.

Flashcard: What is the test for whether your rollback path is independent? — Block the platform's control surface in a game day and revert a real change; if the route calls a platform API, needs credentials from its secret manager, or cannot name the artefact, it is not independent.