1. Rightsizing intermediate

    A platform's rightsizing programme reduces instance sizes based on average utilisation, and incidents follow. What was wrong with the method?

    2 min answer rightsizingheadroompercentilescapacity
  2. Rightsizing intermediate

    A rightsizing proposal moves 180 steady services from 4-vCPU general-purpose instances to 2-vCPU burstable instances because average CPU is 9%, projecting a 55% saving. Several of those services run a two-hour nightly batch. Review the proposal - what would you change, what would you keep, and what would you leave alone?

    2 min answer rightsizingburstablecpu-creditsheadroom
  3. Rightsizing intermediate

    An automated tool recommends reducing your instances by 60% based on average CPU utilisation. What do you check before accepting?

    2 min answer rightsizingheadroomfailure-domainsutilisation
  4. Showback & Chargeback intermediate

    A platform team wants engineering teams to reduce their infrastructure spend. Does showback or chargeback work better, and what determines the answer?

    2 min answer posthogshowbackchargebackincentives
  5. Showback & Chargeback advanced

    A team reduces its cloud bill 30% by removing a standby replica, then has an outage. What was wrong with the incentive structure?

    2 min answer chargebackincentivesreliabilityslo
  6. Showback & Chargeback advanced

    A year into chargeback every product team's reported infrastructure cost is down between 15% and 25%. Company cloud spend is up 8%. Tag coverage is 100% and no commitment expired. Where is the money going, and what would you look at first?

    3 min answer chargebackshared servicesincentivescost allocation
  7. Showback & Chargeback intermediate

    An enterprise is deciding between showback and chargeback for cloud costs. What does each achieve, and what goes wrong with chargeback?

    2 min answer showbackchargebackincentivesgovernance
  8. Showback & Chargeback intermediate

    Cloud spend is growing faster than revenue and no individual team feels responsible. How should cost allocation be designed to actually change behaviour?

    2 min answer finopsshowbackchargebackallocation
  9. Spot & Interruptible Capacity intermediate

    A data platform wants to use interruptible capacity to reduce cost. Which workloads are suitable, and what must be true of them?

    2 min answer databricksspotpreemptioncheckpointing
  10. Spot & Interruptible Capacity intermediate

    A nightly batch job takes six hours on on-demand instances. Someone proposes spot to save 70%. What must be true?

    2 min answer finopsspotresilience
  11. Spot & Interruptible Capacity intermediate

    A recommendation platform runs large offline training and feature computation jobs. What must be true architecturally before spot capacity is usable, and what does it cost?

    2 min answer spotpreemptioncheckpointingbatch
  12. Spot & Interruptible Capacity advanced

    A training pipeline runs on 400 interruptible GPU instances checkpointing every 30 minutes. Regional demand spikes and reclamation goes from a handful of instances an hour to most of the fleet inside ten minutes. What happens minute by minute, and what stops it?

    2 min answer spotgpucheckpointingreclamation