1. Platform Tenancy advanced

    An internal platform hosts workloads from many teams. What isolation model should it use, and how are noisy neighbours handled?

    2 min answer multi-tenancyisolationkubernetesquotas
  2. Platform Tenancy advanced

    An internal platform serves many product teams with very different scale and criticality. How should isolation between them be designed?

    2 min answer razorpayplatform-tenancyisolationquotas
  3. Platform Tenancy advanced

    On 11 December 2024 OpenAI deployed a new telemetry service into its Kubernetes clusters. Its configuration caused every node to perform Kubernetes API operations whose cost scaled with the size of the cluster; API servers saturated and DNS-based service discovery failed across most large clusters, with services unavailable from 15:16 to 19:38 PST. The change had been tested in staging. What structurally failed, what could staging not test, and where would copying their response be a mistake?

    3 min answer openaikubernetescontrol-planemulti-tenancy
  4. Platform Tenancy advanced

    On a shared cluster serving twelve teams, one team's new operator creates a custom resource per user session, reaching about 50,000 objects in a day. Other teams start seeing slow deploys and intermittent API timeouts. What is happening inside the cluster?

    3 min answer kubernetesetcdmulti-tenancynoisy neighbour
  5. Platform Tenancy advanced Multiple choice

    Twelve teams, forty services. One shared Kubernetes cluster or a cluster per team? Justify.

    1 min answer kubernetestenancyisolation
  6. Self-Service Provisioning intermediate

    A platform offers self-service provisioning and teams still file tickets. Why, and what would change it?

    2 min answer delhiveryself-serviceprovisioningfriction
  7. Self-Service Provisioning advanced

    A self-service database provisioning workflow takes 40 minutes and fails about 15% of the time. The platform team says the cloud provider's API is unreliable. Retrying the whole workflow usually works. What is actually going on, and what would you change?

    3 min answer provisioningidempotencyeventual consistencyorphaned resources
  8. Self-Service Provisioning intermediate

    Design the self-service interface for provisioning a database, so that developers do not file tickets and security does not object.

    2 min answer platformguardrailssecurity
  9. Self-Service Provisioning beginner Multiple choice

    Provisioning a database on this platform takes about 4 engineer-hours of real work. The median request takes 6 working days end to end. The platform team of 5 is fully occupied. Why is the lead time roughly twelve times the work time and which change reduces it most?

    3 min answer self-servicequeueinglead timeticket ops
  10. Self-Service Provisioning intermediate Multiple choice

    Which platform operations should be self-service, and which should require review?

    1 min answer self-serviceguardrailsreviewboundaries
  11. Self-Service Provisioning advanced

    Your platform team is now a ticket queue — every environment, every permission, every new service goes through them. How do you get out of it?

    1 min answer platformself-servicebottleneckautomation
  12. Service Mesh Operations advanced

    A platform runs a service mesh. What are the operational realities teams underestimate?

    2 min answer service-meshoperationscontrol-planeupgrades