1. Failure Modes advanced

    One instance in a fleet of 20 is failing 30% of its requests. Health checks pass, aggregate error rate is 1.5%, no alerts fire. How do you detect and handle this?

    2 min answer gray-failureobservabilityresilience
  2. Failure Modes advanced

    One instance in a fleet of fifty is returning correct responses very slowly. Health checks pass and it stays in rotation. How do you detect and handle this?

    2 min answer grey-failurehealth-checksdetectionmitigation
  3. Failure Modes advanced

    Roblox's 2021 outage lasted 73 hours after a service-discovery and key-value cluster degraded under contention. Trace the failure chain, and identify the three architectural properties that turned a degradation into a three-day outage.

    2 min answer robloxconsulservice-discoverycascading-failure
  4. Failure Modes advanced

    What happens when a globally distributed real-time communication service loses an entire region - and how should DNS, health checks, traffic steering, failover, connection draining and data consistency interact?

    3 min answer failure-modesregional-failurefailoverdns
  5. Failure Thinking intermediate

    A major platform launch is three months away. How would you run a pre-mortem, and why bother?

    2 min answer riskfacilitationmeta-skills
  6. Failure Thinking advanced

    Between August and early September 2025 three separate infrastructure bugs degraded Claude's output quality - one affected 16% of Sonnet 4 requests in the worst hour of 31 August - while error rates and latency stayed normal. What class of failure is this, and what would you have had to build beforehand to notice it?

    2 min answer anthropicsilent-failuredetectionquality-regression
  7. Failure Thinking advanced

    What does it mean to design by thinking about failure first, and what does that produce that requirement-driven design does not?

    2 min answer credfailure-modesdesignpremortem
  8. Failure Thinking advanced

    What does it mean to think in failure modes as a habit rather than as a checklist item, and which questions consistently produce findings?

    3 min answer failure-modesdesign-reviewresiliencequestions
  9. Failure Thinking advanced

    What questions characterise failure thinking, and which are most often skipped in design reviews?

    2 min answer failure-thinkingreviewdependenciesdegradation
  10. Fault Isolation intermediate

    A design platform's synchronous editing experience and its asynchronous export rendering share a compute cluster. Exports occasionally saturate it and the editor becomes unusable. What would you change?

    2 min answer fault-isolationbulkheadsworkload-separationcanva
  11. Fault Isolation advanced

    A financial platform's card authorisation path must survive dependency failures, deployments, overloaded downstreams and partial network failures. How should isolation, deadlines, breakers, bulkheads, caching, shedding and fallback combine?

    2 min answer brexrampauthorisationisolation
  12. Fault Isolation advanced

    A security vendor pushes a content update to millions of endpoints simultaneously and a malformed file crashes them all at kernel level. What should have been in place, and why is "it was data, not code" the wrong defence?

    3 min answer crowdstrikestaged-rolloutblast-radiuskernel