1. Observability advanced

    Your observability bill is now 40% of your compute bill. Leadership wants it cut in half without going blind. What do you cut?

    2 min answer observabilitycostcardinalitysampling
  2. Observability advanced

    p99 latency on checkout tripled overnight. Dashboards look normal, no deployment went out, and every service reports healthy. How do you find it?

    3 min answer debuggingtracinglatencydiagnosis
  3. Peak Event Readiness advanced

    A retailer expects 20× normal traffic for a two-hour sale launch. Plan the readiness.

    2 min answer peakcapacitydegradationwar-room
  4. Polyglot Persistence advanced

    Four services want four different databases: Postgres, MongoDB, Cassandra and Neo4j. What do you say?

    2 min answer polyglotoperationsdecision-makingstandards
  5. Polyglot Persistence advanced

    Your search index and your database disagree — some products appear in search that were deleted, and some new ones never appear. How do you make this reliable?

    2 min answer consistencyindexingoutboxreconciliation
  6. Privacy Engineering advanced

    Product wants to add "customers who bought this also bought" using purchase history. What does privacy by design require here?

    2 min answer privacypurpose-limitationminimisationgdpr
  7. Quality Attributes advanced

    You are starting a greenfield platform. How do you establish what the architecture must satisfy before designing anything?

    2 min answer driversquality-attributesconstraintsmethod
  8. Query Optimisation advanced

    A query that ran in 50ms for two years now takes 90 seconds. Nothing was deployed and the data volume grew normally. What happened?

    2 min answer plannerstatisticsplan-flipdiagnosis
  9. Capacity Planning advanced

    Your platform runs across three availability zones. How would you determine whether it actually survives losing one?

    2 min answer capacityzone-failureheadroomtesting
  10. Reliability & Resilience advanced

    In July 2019 a single regex deployed globally took Cloudflare's network to 100% CPU within seconds. What does this say about how configuration should be released?

    2 min answer deploymentblast-radiusconfigcase-study
  11. Reliability & Resilience advanced

    Maersk rebuilt roughly 4,000 servers and 45,000 PCs in about ten days after NotPetya in 2017, and recovered its directory only because one data centre had been offline during the attack. What does this say about DR design?

    2 min answer disaster-recoveryransomwarebackupcase-study
  12. Replication advanced

    Users report seeing stale data intermittently. Replication lag is normally under a second but spikes to minutes twice a day. How do you handle it?

    2 min answer replicationlagroutingconsistency