beginner 3 min answer Multiple choice

Forty minutes into an incident someone asks whether a particular batch host could have written to the production object store. The deployment diagram is current and shows every host, subnet, zone and region correctly. Which missing annotation stops it answering?

deployment-diagramsworkload-identityblast-radiusincident-responseannotation
Pick one
Show the full answer Hide the answer

What is being tested

Whether you know what question a deployment diagram is for. Most are drawn to show where software runs. The questions asked of them under pressure are what could this thing have touched and what fails with it — and neither is answerable from topology alone.

The mechanism

A host's reach is decided by two things the diagram usually omits: the identity it assumes (an IAM role, a service account, a certificate subject) and what that identity is allowed to do. Network position is necessary and not sufficient. A batch host sitting in a private subnet with no inbound route can still hold a role with write access to the production bucket, because object-store access goes through an API endpoint and an authorisation check, not through the subnet.

So two hosts drawn identically, in the same zone, on the same subnet, can have completely different authority. The diagram shows them as the same box, and during an incident that costs you the 20 minutes it takes someone to go and read the role policy. In production that reading happens while the clock is running and the answer is being guessed at in a call with thirty people on it.

What to write on the diagram

  • The identity on every node: batch-runner@prod rather than "Batch host". One short string per box.
  • The sensitive destinations each identity can reach, as a labelled edge: write to the primary database, write to the artefact bucket, read-only to the secrets store. Choose 3 or 4 edges, not every permission — a complete permission dump is a policy document and nobody reads it mid-incident.
  • Which credential path it uses — instance metadata, mounted token, long-lived key — because that decides whether revocation is instant or requires a rotation.
  • The scope boundary: account, project or subscription. Hosts in different accounts cannot reach each other by accident; hosts in one account usually can.

The rule: a deployment diagram should let a responder answer "could that have done this?" without opening a console. If it cannot, the annotation is missing, not the diagram.

Why the other options fail

  • Instance type and vCPU count. Useful for a capacity conversation and irrelevant to authority. A 2-vCPU host with an over-broad role is the dangerous one, and the size tells you nothing about it. This is the commonest thing people do add, because it is easy to generate.
  • Average request rate on each link. This is a traffic view. It helps with cross-zone cost and saturation, and it still cannot say what a host is permitted to write. Rates also change hourly, so the annotation rots faster than anything else on the page.
  • The build and ship pipeline. Worth knowing, and it answers a different question — what changed and what shares a release. It is the right annotation for "what failed together", not for "what could this reach".

When this is the wrong answer

On a diagram drawn for a capacity or cost review, identity is noise: the reader wants instance counts, zone split and link volumes. Draw two views rather than one crowded one — an operations view carrying identity and reachability, a capacity view carrying sizes and rates. One diagram that tries to serve both gets skimmed by both audiences.