Rollback & Forward Fix
When reversing is genuinely possible, and designing so that it usually is.
5 to work through
-
intermediate
A deployment has caused a production problem. When should the team roll back and when should they fix forward?
2 min answer -
intermediate
Production is degraded after a release. The team tries to roll back and discovers a migration has already run. What do you do now, and what do you change afterwards?
2 min answer -
advanced
A background-job system has millions of queued jobs after an incident, and replaying them all at once will crash the databases. How should recovery be sequenced?
2 min answer -
advanced
A change causes a global incident. When should you roll back and when should you fix forward, and what property makes the choice easy?
2 min answer -
advanced
An order service and a billing service deploy independently. Billing shipped a version at 09:00 that requires a tax_basis field on every message. Orders shipped at 10:00 and began emitting it. At 10:40 the orders release is causing checkout errors and the on-call rolls orders back. What happens next, and what property would have made that rollback safe?
3 min answer
4 terms in this topic
Forward Fix
Recovering from a bad release by shipping a corrective change rather than reverting to the previous version.
conceptRecovery Load
The traffic generated by a system returning to service - reconnections, queued work, cold caches, retries - which is qualitatively different from ste…
conceptRestore-Only Rollback
The state of a system that cannot be returned to its previous version by deploying it, so the only reversal available is restoring from backup - whic…
conceptReverse Compatibility Window
The period during which every live consumer must still accept the previous version's output, because a rollback of one component is a forward-incompa…
Neighbouring topics
Delivery & Release Engineering
General material on getting a change from commit to production safely and often.
Pipeline Architecture
Stages, fan-out, caching, and the difference between a pipeline and a long script.
Build Reproducibility
Pinned inputs and hermetic builds, so one commit cannot produce two different artifacts.
Artifact Management
Immutable versioned outputs, promotion between repositories, and retention policy.
Environment Strategy
How many environments earn their cost, what each proves, and what none of them prove.
Branching Models
GitFlow, trunk and release branches as delivery constraints rather than Git preferences.
Continuous Integration Discipline
Integrating to the mainline daily, and the test speed and review culture that requires.
Deployment Strategies
Rolling, blue-green, canary and shadow, and the traffic and state each one assumes.
Progressive Delivery
Separating deploy from release, and exposing a change to users in controlled increments.
Database Migration Under CD
Expand-contract, backwards-compatible schema change, and migrations that cannot roll back.
GitOps
Declared desired state in version control, with a reconciler closing the gap continuously.
IaC Modules & Drift
Reusable infrastructure modules, state ownership, and detecting what changed out of band.
Policy as Code
Encoding standards as automated admission and plan-time checks instead of review comments.
Pipeline Secrets
Short-lived credentials, workload identity, and why the CI system is a prime target.
Supply-Chain Provenance
SBOMs, signed artifacts, attestation, and knowing what actually went into a build.
Deployment Gates
Automated verification between stages, and the difference between a gate and a delay.
Flow Metrics
Work in progress, flow time and flow efficiency — where a change waits rather than moves.
Change Management vs CD
Reconciling CAB-era controls with continuous delivery without pretending either away.
Multi-Region Rollout
Ordering regions, bake time, and stopping a bad change before it becomes global.