practice

Incident Severity Levels

A small, agreed scale of incident severity that determines response, escalation and communication without requiring debate during the event.

The purpose is to remove a decision from the worst possible moment. A severity level is a shorthand for a response, not a judgement of importance.

A workable scale: SEV1 — critical business function unavailable or a data-loss or security event. All-hands response, executive notification, customer communication. SEV2 — significant degradation or a major feature broken; a subset of users severely affected. Paged response, incident commander, status page update. SEV3 — minor degradation with a workaround. Business-hours response, ticket tracked. SEV4 — no user impact but requires attention.

Three rules make it work:

Define severity by user impact, not by which component failed or how technically alarming it looks. A crashed instance with no user effect is not a SEV1.

Anyone may declare, and declaring high is free. Downgrading a SEV1 that turns out to be a SEV3 costs a few minutes of several people's time. Under-declaring costs an hour of the wrong response.

Each level carries a predefined response — who is paged, who is notified, whether customers are told, how often updates go out. If the level does not determine those, it is a label rather than a mechanism.