1. Platform SLOs advanced

    A product team commits to 99.95% availability. Their service depends on your platform's ingress, config service and secret manager. What do you tell them?

    2 min answer slodependenciesavailability
  2. Platform SLOs advanced

    Application teams say the platform is unreliable. The platform team's dashboard shows 99.95% on every component. How do you resolve this?

    2 min answer platformslomeasurementtrust
  3. Platform Team Topologies advanced

    A team of six owns nineteen services and is missing every deadline. They ask for more engineers. Is that the right answer?

    2 min answer team-topologiesboundariesorganisation
  4. Platform Team Topologies advanced

    How should platform work be organised as an organisation grows, and what structure prevents the platform becoming a bottleneck?

    2 min answer team-topologiesplatformenabling-teamsbottleneck
  5. Platform Telemetry advanced

    A platform injects standard labels into every metric it scrapes: team, service, environment, pod and commit SHA. Between deploys the metrics backend is healthy. Within minutes of each fleet deploy, query latency triples, ingester memory climbs and dashboards covering the last six hours time out. Samples per second have not changed. Where is the time going?

    3 min answer prometheuscardinalityseries churnlabels
  6. Platform Telemetry advanced

    On 11 December 2024 OpenAI rolled a new telemetry service across its Kubernetes fleet. Staging was clean; the change reached every cluster in under thirty minutes; all services degraded for roughly four hours. What is the failure chain, and which of its links is the one to design against?

    3 min answer openaikubernetescontrol planedns
  7. Platform Tenancy advanced

    A platform serves internal teams with very different scale and criticality. Should they share infrastructure?

    1 min answer platform-tenancyisolationnoisy-neighbourtiering
  8. Platform Tenancy advanced

    An internal platform hosts workloads from many teams. What isolation model should it use, and how are noisy neighbours handled?

    2 min answer multi-tenancyisolationkubernetesquotas
  9. Platform Tenancy advanced

    An internal platform serves many product teams with very different scale and criticality. How should isolation between them be designed?

    2 min answer razorpayplatform-tenancyisolationquotas
  10. Platform Tenancy advanced

    On 11 December 2024 OpenAI deployed a new telemetry service into its Kubernetes clusters. Its configuration caused every node to perform Kubernetes API operations whose cost scaled with the size of the cluster; API servers saturated and DNS-based service discovery failed across most large clusters, with services unavailable from 15:16 to 19:38 PST. The change had been tested in staging. What structurally failed, what could staging not test, and where would copying their response be a mistake?

    3 min answer openaikubernetescontrol-planemulti-tenancy
  11. Platform Tenancy advanced

    On a shared cluster serving twelve teams, one team's new operator creates a custom resource per user session, reaching about 50,000 objects in a day. Other teams start seeing slow deploys and intermittent API timeouts. What is happening inside the cluster?

    3 min answer kubernetesetcdmulti-tenancynoisy neighbour
  12. Platform Tenancy advanced Multiple choice

    Twelve teams, forty services. One shared Kubernetes cluster or a cluster per team? Justify.

    1 min answer kubernetestenancyisolation