1. Compute Optimisation advanced

    A platform's accelerator fleet reports high utilisation, yet throughput per accelerator is far below the hardware's capability. Where does the waste usually hide?

    2 min answer compute-optimisationutilisationbatchingdata-pipeline
  2. Compute Optimisation advanced

    A team migrates its stateless serving fleet to an ARM instance family for roughly 20% better price-performance on paper. What has it given up, and when does that bill arrive?

    2 min answer arminstance-familybuild-matrixcommitments
  3. Compute Optimisation advanced

    Which compute cost lever has the largest payoff, and why is it the one teams attempt last?

    2 min answer computespotarchitecturecost
  4. Cost Allocation advanced

    A multi-tenant platform needs to know what each customer costs to serve. What makes this hard, and how is it approached?

    2 min answer cost-allocationmulti-tenancyshared-resourcesattribution
  5. Cost Allocation advanced

    Tag coverage is at 98% and teams still argue about their bill every month. Roughly a fifth of the spend is shared - the service mesh, the log pipeline, NAT gateways, the Kafka cluster, and the commitment discounts. How would you allocate it?

    3 min answer cost-allocationshared-costsshowbackcommitments
  6. Cost & FinOps advanced

    A design review presents a new event-driven platform. What cost questions do you ask before approving it?

    2 min answer finopscost-modelreviewegress
  7. Cost Governance advanced

    A commerce platform must control cost while guaranteeing capacity for peak retail events. How should governance handle the tension?

    2 min answer cost-governancepeakexceptionsbudgets
  8. Cost Governance advanced

    At 02:14 on a Saturday a storage-triggered function starts writing its output back into the prefix that triggers it. By Monday 09:00 the account has logged 290 million invocations against a monthly forecast of about $700. The budget alert fired on Sunday at 18:40. Which design decisions made this possible?

    2 min answer cost-governanceserverlessguardrailsquotas
  9. Architecture Cost Modelling advanced

    A GPU platform must decide between holding warm capacity and accepting cold starts. How should the two be compared economically?

    2 min answer together-aimodalcold-startwarm-pool
  10. Architecture Cost Modelling advanced

    A video analysis pipeline splits streams into frames, analyses each, and aggregates. Built as serverless functions passing frames through object storage, it costs far more than expected. Diagnose.

    2 min answer amazonserverlesscostboundaries
  11. Architecture Cost Modelling advanced

    On 8 August 2024 Hugging Face announced it had acquired XetHub, a 14-person storage startup, describing it as its largest acquisition to that date. Continuing on Git LFS, or building content-defined chunking in-house, were the alternatives. What cost model separates the three, and where would copying this be a mistake?

    3 min answer build vs buydeduplicationstorage economicsacquisition
  12. Cost per Request advanced

    An inference platform's cost per request varies by two orders of magnitude between requests. What does that imply for pricing, capacity and engineering priorities?

    2 min answer cost-per-requestvariancepricingbatching