1. Telemetry Cost intermediate Multiple choice

    A design platform of the kind Canva runs exports 40 metrics per service. An engineer adds a `pod_name` label so a noisy pod can be identified. The service runs 600 pods and deploys twice a day, so pod names turn over completely every 12 hours. Retention is 30 days. Roughly how many distinct series does that one label create over the retention window?

    2 min answer canvacardinalitytelemetry costmetrics
  2. Telemetry Cost intermediate

    A platform's log volume has grown until log storage is one of its largest infrastructure costs. What should change, and what should not?

    2 min answer elasticlogscostindexing
  3. Sampling intermediate Multiple choice

    A 400-service mesh runs Envoy sidecars, the data plane Lyft built and open-sourced in 2016. Tracing is on at 100% in staging and 0.1% uniform in production, and every incident review ends with "we did not have a trace for the failing request". The tracing backend can afford about 2% of production spans. Which sampling design do you adopt?

    4 min answer lyftenvoytracingsampling
  4. Test Architecture Strategy intermediate

    What should a test strategy specify, and what makes strategies fail in practice?

    2 min answer test-strategyownershiplayersrisk
  5. Test Data Management intermediate

    What are the test-data options, what must be true regardless, and which combination actually finds the bugs that reach production?

    2 min answer test-datamaskingsyntheticprivacy
  6. Test Pyramid Shapes intermediate

    A team is told its test suite should be a pyramid. When is that shape wrong?

    2 min answer growwpyramidriskintegration
  7. Test Pyramid Shapes intermediate

    A team's test distribution is an inverted pyramid — many end-to-end tests, few unit tests. What does that cost, and what causes it?

    2 min answer test-pyramidcostflakinessarchitecture
  8. Testing Strategies intermediate

    A team's test suite is slow, flaky and gives little confidence. What shape should the strategy have, and what determines it?

    2 min answer lineartestingpyramidflaky
  9. Testing Strategies intermediate Multiple choice

    An end-to-end suite has 200 tests. Each fails spuriously about 1% of the time, independently of the others. Roughly what fraction of runs on correct code go red anyway?

    2 min answer booking.comtestingflaky testsci
  10. Third-Party Risk intermediate

    Your KYC verification vendor is down for six hours. Customer onboarding stops. The board asks why a vendor outage became your outage.

    2 min answer vendorsresiliencedegradation
  11. Communicating Threat Models intermediate

    A security review produces findings that engineering teams dispute or ignore. How should the findings be communicated to be acted on?

    2 min answer freshworkssecurityfindingscommunication
  12. Communicating Threat Models intermediate

    A threat model is technically thorough and the engineering team does not act on it. What is wrong with how it is communicated?

    2 min answer threat-modelcommunicationactionabilityshopify