practice

Failure Domain Annotation

also called Shared-Fate Labelling

Writing on a deployment diagram what each group of nodes shares - image, configuration source, pipeline, control plane - so the diagram answers what fails together rather than only where things run.

deployment-diagramfault-domaincorrelated-failureavailabilityoracle

An auditor asks what survives the loss of one availability zone. The diagram shows three zones, so the answer looks like "two thirds of capacity, no data loss". Three weeks later a configuration change reaches all three zones in 90 seconds and everything is down. The diagram was not wrong about geography and it was silent about fate, and fate is what the auditor was asking about.

A failure domain is the set of things that go down together. A deployment diagram showing only zones implies independence the system does not have, because the zones share a pipeline, a configuration source, an identity provider and a machine image.

Why it matters

Availability arithmetic is extremely sensitive to the domain count. Three genuinely independent 99.9% components in parallel fail together with probability on the order of one in a billion — a number nobody has ever achieved. Now add one shared configuration path: 200 pushes a month with 1 in 500 of them bad is 0.4 bad pushes a month, and at 30 minutes to detect and roll back that is about 12 minutes of downtime a month, or 99.97%. The parallel arithmetic promised nine nines and the shared path delivers four, because the shared term dominates every independent one. The diagram is the only place where a reviewer can see that before an incident does.

The second reason is operational. During an incident the first useful question is "what else is in this blast radius", and an annotated diagram answers it in seconds rather than by interrogating three platform teams.

Implementation patterns

  • Label every node group with what it shares: image, configuration source, deploy pipeline, control plane, credential, DNS zone. Each shared item is a failure domain of its own.
  • Use the finest grain the provider actually exposes. Oracle Cloud Infrastructure documents three fault domains inside each availability domain and uses them as anti-affinity placement for instances, so replicas of one workload do not land on the same hardware within a zone.
  • Mark each cross-boundary edge mandatory or incidental. Quorum writes and replication must cross; service-to-service reads usually need not, and that distinction is both a cost and a failure-mode fact.
  • Colour by failure domain rather than by owning team, since team colour is the annotation people reach for and it predicts nothing about correlated failure.
  • Derive it from infrastructure code where possible, and date it where not.

Industry example

Oracle Cloud Infrastructure's published model separates regions, availability domains and, inside each availability domain, three fault domains grouping hardware so instances can be placed anti-affine to each other. The useful part for a diagram is the vocabulary: it names a boundary below the zone that is schedulable, so an annotation against it is actionable rather than decorative.

Failure scenarios

  • One configuration push to all zones, the case in the opening, which no zone boundary contains.
  • A cross-zone quorum of two, where losing one zone loses write availability even though two thirds of capacity survives.
  • Identical machine images carrying one kernel defect into every zone at once, or a single CI runner on every recovery path.
  • A control plane homed in one zone, so losing it removes the ability to scale the other two.

Trade-offs

Choose Gains Pays
Annotate shared fate honest availability numbers · fast blast-radius answers a busier diagram · worse numbers that someone must own
Zones only clean picture · matches the provider's console implies independence that does not exist
Annotate everything shared complete unreadable past a few groups — rank by what one action can touch

When not to use it

On a context diagram, where it is the wrong altitude and clutters the one artefact that should stay readable. For a single-zone system, where the only honest annotation is "all of it", and a sentence says that better than a diagram. And where no workload placement decision exists — a fully managed serverless product gives you no scheduling control, so annotating domains you cannot influence produces anxiety rather than action.

Interview question

Q: Your deployment diagram shows three zones. An executive asks for the availability number. What do you need before you can answer, and what will you probably have to tell them?

What a strong answer covers: the number depends on correlated failure, not zone count, so first enumerate what the zones share — pipeline, configuration, image, control plane, identity. Then state availability as bounded by the worst shared path rather than by the parallel arithmetic, with the honest consequence that a fourth zone changes almost nothing while staging configuration changes a great deal. Name the cheapest structural fix: push configuration zone-by-zone with a wait, which converts the largest shared domain into three.

Quick check

Quiz: Three zones each at 99.9%, plus one configuration path taking 200 pushes a month of which 1 in 500 is bad and takes 30 minutes to roll back. What bounds your availability? — The configuration path, at roughly 12 minutes of downtime a month, because it is common to all three zones.

Flashcard: Why does zone count overstate independence? — Zones share the pipeline, the configuration source, the image and the control plane; those shared paths are single failure domains regardless of geography.