1. Systems Thinking advanced

    An organisation finds most incidents are caused not by individual service failures but by interactions between healthy services. What discipline addresses this, and what specifically should be done?

    2 min answer delhiverysystems-thinkingemergentdependencies
  2. Systems Thinking advanced

    An organisation finds most incidents are caused not by individual service failures but by unexpected interactions between healthy services. What does systems thinking add here?

    2 min answer systems-thinkingemergencefeedback-loopsdependencies
  3. Systems Thinking advanced

    The same class of incident has recurred four times in a year despite each postmortem producing actions that were completed. What is going wrong?

    1 min answer meta-skillssystems-thinkingincidentsroot-cause
  4. Systems Thinking advanced

    You add an approval gate to reduce production defects. Six months later defect rates are higher. Explain.

    2 min answer systems-thinkingfeedbackgovernance
  5. Technical Leadership advanced

    What changes when you move from making architectural decisions to being responsible for how an organisation makes them?

    2 min answer leadershipscalingdirectionrestraint
  6. Technical Leadership advanced

    You are the first architect at a 200-engineer company with no architecture function. What do you do in the first three months?

    3 min answer leadershipinfluenceassessmentgovernance
  7. Technical Leadership advanced

    You join a platform group and find four teams have each written their own retry-and-backoff library, with different defaults, and a fifth team has none. Everyone knows this is duplicated. Nobody has consolidated it. What is the diagnosis, and what do you change?

    3 min answer platformduplicationincentivesconways-law
  8. Trade-off Analysis advanced

    A SaaS vendor's documented recovery posture is point-in-time restore for any individual data store. Atlassian's April 2022 incident tested it at scale - sites for 775 customers were deleted at once, no customer lost more than five minutes of data, and the last sites came back on 18 April. Which number had the organisation actually been buying, and what would you change first?

    3 min answer atlassianrportodisaster-recovery
  9. Trade-off Analysis advanced

    Dropbox moved file storage off Amazon S3 onto its own infrastructure - Magic Pocket - completing the migration in 2016. What did that decision buy and what did it pay, and which facts about your own situation would have to be true before the same move is defensible?

    3 min answer dropboxbuild-vs-buyunit-economicsstorage
  10. Trade-off Analysis advanced

    Product requires strong consistency and p99 latency under 100 ms for a globally distributed user base. Both are non-negotiable. What do you do?

    2 min answer tradeoffsconsistencyrequirements
  11. Trade-off Analysis advanced

    When Notion sharded its Postgres database in 2021 it created 480 logical shards spread across 32 physical instances, 15 per host. In 2023 it moved to 96 hosts with 5 logical shards each, keeping the same 480. What did the original choice buy, what did it cost, and what would you have had to believe to choose 32 logical shards instead?

    3 min answer notionshardingreversibilitycapacity