1. Platform SLOs advanced

    A product team commits to 99.95% availability. Their service depends on your platform's ingress, config service and secret manager. What do you tell them?

    2 min answer slodependenciesavailability
  2. Platform SLOs advanced

    Application teams say the platform is unreliable. The platform team's dashboard shows 99.95% on every component. How do you resolve this?

    2 min answer platformslomeasurementtrust
  3. Platform SLOs intermediate

    Should an internal platform have SLOs, and what should they cover?

    2 min answer dream11platform-sloreliabilitycontract
  4. Platform SLOs intermediate

    Why should a platform publish SLOs to internal consumers, and what should they cover?

    2 min answer platform-slosdependenciesexpectationsavailability
  5. Platform Team Topologies advanced

    A team of six owns nineteen services and is missing every deadline. They ask for more engineers. Is that the right answer?

    2 min answer team-topologiesboundariesorganisation
  6. Platform Team Topologies advanced

    How should platform work be organised as an organisation grows, and what structure prevents the platform becoming a bottleneck?

    2 min answer team-topologiesplatformenabling-teamsbottleneck
  7. Platform Team Topologies beginner Multiple choice

    Six platform engineers support forty product teams. Each team raises roughly 1.5 requests a week and each request costs about 90 minutes of an engineer's time once the context switch is counted. Roughly how much of the team's capacity is interrupt work and what does that number rule in or out?

    2 min answer platform teaminterruptscapacityteam topologies
  8. Platform Telemetry advanced

    A platform injects standard labels into every metric it scrapes: team, service, environment, pod and commit SHA. Between deploys the metrics backend is healthy. Within minutes of each fleet deploy, query latency triples, ingester memory climbs and dashboards covering the last six hours time out. Samples per second have not changed. Where is the time going?

    3 min answer prometheuscardinalityseries churnlabels
  9. Platform Telemetry intermediate

    A platform team cannot tell which of its capabilities are used, by whom, or whether they are working. What telemetry does a platform need about itself?

    2 min answer druvaplatform-telemetryusageadoption
  10. Platform Telemetry advanced

    On 11 December 2024 OpenAI rolled a new telemetry service across its Kubernetes fleet. Staging was clean; the change reached every cluster in under thirty minutes; all services degraded for roughly four hours. What is the failure chain, and which of its links is the one to design against?

    3 min answer openaikubernetescontrol planedns
  11. Platform Telemetry intermediate

    What should a platform team measure about its own platform, and what is most often missing?

    1 min answer platform-telemetryadoptionexperienceslos
  12. Platform Tenancy advanced

    A platform serves internal teams with very different scale and criticality. Should they share infrastructure?

    1 min answer platform-tenancyisolationnoisy-neighbourtiering