practice

Business Continuity

Keeping the business operating through disruption — a broader discipline than disaster recovery, and one that includes not using the system at all.

continuitybcpdrmanual-workaroundsdependencies

Definition

Business continuity planning asks how the organisation keeps functioning when something significant fails: a system, a supplier, a site, a workforce. Disaster recovery is the technical subset concerned with restoring systems.

What it includes that DR does not

  • Manual workarounds. How does the business operate for four hours without the system? Frequently the highest-value and cheapest continuity control, and frequently undocumented because it is unglamorous. Retail, healthcare and logistics all have paper procedures for exactly this reason.
  • Third-party failure. A payment provider, a delivery partner, a logistics system, a key SaaS application. Your continuity plan is incomplete if it assumes suppliers are available, and their availability is not something you control.
  • People. Key-person dependency, site loss, a workforce unable to travel.
  • Communication. Who tells customers, regulators and staff, through which channel, and using what system — noting that the usual channel may itself be down.
  • Prioritisation. Which business functions are restored first, decided in advance. Under pressure this argument cannot be had.

The business impact analysis

The foundation: for each business function, determine the impact of losing it over time — one hour, one day, one week. That produces the maximum tolerable outage, from which technical RTO and RPO are derived rather than invented.

Impact is not only revenue. Regulatory obligations, safety, contractual penalties and customer trust all count, and some of them dominate.

The output is a prioritised list, which is the point. Not everything can be restored first, and deciding the order in advance is the whole exercise.

What makes plans real

Rehearsal, including the parts that are not technical. The recurring finding across the industry is that the control plane fails: identity, secrets, deployment or communication tooling has a dependency that makes the planned response impossible. No document review produces that finding; only attempting it does.

Second: the plan must be accessible when the systems are down. A continuity plan stored in the system it covers is a recurring and entirely predictable failure.

Failure scenarios

  • A plan written for an audit and never rehearsed.
  • Manual workarounds undocumented, so staff improvise.
  • Third-party dependencies unmapped.
  • Contact lists out of date, discovered at 3am.
  • No prioritisation, so everything is restored simultaneously and badly.

Interview question

"Your primary system is unavailable for six hours. What must the business be able to do without it, and what have you prepared?"