advanced 3 min answer

At 22:52 UTC on 21 October 2018 GitHub lost connectivity between its US East Coast network hub and its primary East Coast data centre for 43 seconds while failing optical equipment was replaced. The cascade produced 24 hours and 11 minutes of degraded service in which repository data was intact but webhooks were not delivered and Pages sites were not built. What does an incident shaped like that demand of a status report and what does a single overall indicator make impossible?

githubstatus-pageincident-communicationcomponentsdeveloper-platform
Show the full answer Hide the answer

The trigger and the shape of the impact

GitHub's published analysis records a 43-second loss of connectivity during maintenance on 100G optical equipment, automated database failover across the country as a result, and 24 hours and 11 minutes of degraded service while writes on both coasts were reconciled. The company stated that no user data was lost and that the impact was to website metadata held in MySQL — issues, pull requests — rather than to Git repository data.

The engineering detail that decides the communication problem is that the impact was uneven. Clone and push kept working for most of the window. Webhook delivery and Pages builds did not. Those are not degrees of the same thing; they are different products to different consumers.

Why one indicator cannot carry it

A consumer reads a status page to make a decision, and the decisions differ by what they use. A developer pushing code wants to know whether to keep working. A team whose CI is triggered by webhooks wants to know whether to switch to polling. A company whose marketing site is a Pages build wants to know whether to tell its own customers.

A single aggregate colour forces every one of those consumers to derive the same answer from the same pixel, which is wrong for almost all of them. Aggregation is also directionally biased: any rule that produces one colour from many components either understates the outage for the people fully broken or overstates it for the people unaffected. There is no setting of the rule that is honest to both.

The same argument applies to the vocabulary. A word like "degraded" spans everything from slightly slow to writes rejected. If a status vocabulary is not tied to measurable conditions published in advance, the word is chosen under pressure and means whatever the person typing it hoped. Define each level against a condition — error rate above a stated threshold, a named operation unavailable — before you need it.

The structural fix, and what GitHub actually shipped

On 11 December 2018 GitHub launched a new status site at githubstatus.com that lists individual component statuses rather than one overall indicator. Git operations were split out from API requests and Pages builds became trackable independently of Notifications. The new site also separated component status — a snapshot of one component — from incidents, which are tracked communications, and added subscriptions by email, SMS and webhook. The old status.github.com was deprecated on 28 February 2019.

The component list is the design decision. Derive it from the integration surfaces a consumer can depend on independently, not from the internal service graph: a consumer cannot map "metadata reconciliation service" onto anything they do. Six to ten rows that match the things people build against is the useful range. Thirty rows named after internal systems is a different failure — the reader cannot find their row.

Common weak answers

  • "Post more frequent updates." Cadence without granularity still leaves a CI owner unable to learn that webhooks specifically are down. Commit to the next update time as well, since that is a promise you can keep when a restoration estimate is not.
  • "Say it is degraded and explain in the incident text." Prose is not machine-readable, and a large share of consumers read status through an integration rather than with their eyes.
  • "Add a severity score." A number aggregates again, one layer up.

What a strong answer adds

The subscription mechanism matters as much as the page: a consumer who can subscribe to one component gets paged about the thing they depend on and ignores the rest, which is the only way a status page scales to a platform with many unrelated products.