intermediate 2 min answer

A communication platform is business-critical for its customers' own operations. What does business continuity require beyond technical disaster recovery?

business-continuitydependenciescommunicationpeoplezoomdesign
Show the full answer Hide the answer

Beyond technical recovery

Disaster recovery restores systems. Business continuity keeps the business operating, which includes several things no failover addresses.

1. Communication when the platform is down. For a communications company this is acute: the tools used to coordinate an incident may be the product that is failing. An out-of-band channel — a different provider, tested regularly — is a genuine requirement rather than a formality.

2. Customer communication. A status page hosted independently of the affected infrastructure, prepared templates, and a decided threshold for who communicates what. Deciding what to tell customers during an outage is not something to improvise while the outage is happening, and for a platform embedded in customers' operations, silence is itself damaging.

3. People. Key-person dependency is a continuity risk. If one engineer understands the media routing layer, that is a single point of failure with a notice period. Documentation, cross-training and rotation are continuity controls.

4. Third-party dependencies. Identity providers, payment processors, carriers, cloud regions and CDN providers. Continuity planning must include their failure, their business failure, and their contractual exit, since a supplier's outage is your outage from the customer's perspective.

5. Regulatory and contractual obligations during disruption. Notification deadlines, service credits, and reporting requirements that continue while the technical incident is being resolved.

6. Recovery of business processes, not just systems. Billing that did not run, support tickets that accumulated, compliance reports that were not filed. The technical recovery is the start of the business recovery.

What makes plans real

They have been exercised, unannounced, by the people who would actually execute them — not by the people who wrote them. And they include the failback, which is usually harder than the failover and almost never rehearsed.

The consistent finding from real exercises is that plans fail on unglamorous details: an out-of-date contact list, a runbook step referring to a system that was decommissioned, a person who has left, a credential nobody has, or an assumption that a dependency would be available.

The question that tests it

"Our primary region is unavailable and the person who wrote the runbook is unreachable. What happens?" If the answer depends on that person, the plan documents their knowledge rather than replacing the need for it.