advanced 3 min answer

You own the 200 ms p95 budget for a search API and measure 188 ms today. Product wants a personalisation call that costs 60 ms at p95. Walk me through the conversation and what you would commit to.

performance-budgetslatencyprioritisationdegradationinterview
Show the full answer Hide the answer

What the interviewer is testing

Whether you treat a budget as a balance to be spent rather than a target to be missed, and whether you can price a product request in milliseconds without becoming the engineer who says no. The weak version of this answer argues about whether 60 ms is a lot. The strong version establishes who pays and how the payment is verified, because a budget with no payment mechanism is a preference.

The clarifying questions that change the answer

  • Is the call on the critical path, or can it run concurrently with the existing work? If personalisation does not need the candidate set, it runs alongside retrieval and the composed cost is max(), not sum(). A 60 ms concurrent call under a 188 ms path can cost nothing at all.
  • Where did 200 ms come from? A measured conversion or abandonment relationship is a different object from a number someone typed in a design document. One can be renegotiated with evidence; the other should never have been a budget.
  • What does the feature move, per millisecond spent? If search conversion falls by a known amount per 100 ms, both sides of the trade are in the same currency and the decision stops being a taste argument.
  • Is the budget contractual? A partner SLA removes options that an internal target leaves open.

A strong answer's arc

There are four ways to pay, in order of preference:

  1. Parallelise. Issue personalisation concurrently and charge only the amount by which it extends the critical path. Verify with a shadow run, not with arithmetic.
  2. Precompute. Move the work off the request path into a per-user vector refreshed every few minutes. The bill becomes staleness and a pipeline to operate, and p95 pays a lookup of a few milliseconds.
  3. Buy headroom. Fund a profiling and optimisation slice to free 60 ms elsewhere. This is real work with real risk, and it must appear on the roadmap alongside the feature rather than be assumed.
  4. Degrade. Give the call a 40 ms deadline with a documented fallback to unpersonalised ranking. The feature's measured lift must then be computed including the fallback rate, or you will be shipping a feature that is silently absent for the slowest 5% of requests.

Prefer 1, then 2. Choose 4 only when the feature degrades gracefully, and treat 3 as a funded project rather than an assumption. Verify whichever you pick against shadow traffic in production, since a composed percentile measured in staging is a different number.

The subtlety that separates a strong answer: percentiles do not add. A 60 ms p95 call bolted onto a 188 ms p95 path does not yield 248 ms p95. The composed percentile depends on whether the two slow tails happen to the same requests, and the only honest way to find out is to run the call in shadow mode and measure the composed distribution before committing.

Common weak answers

  • "Raise the budget to 260 ms." If the number came from user behaviour, moving it without re-measuring that behaviour deletes the control and keeps the ceremony.
  • "Grant a temporary exemption." Exemptions without an expiry date and a named owner are how budgets become dashboards. Every waiver gets a date, and the build fails again on that date.
  • "It is only 60 ms." This is true of every individual request, which is why systems get slow without any single change being responsible.

What a strong answer adds

An explicit statement of who decides. The architect supplies the exchange rate and the options; the product owner spends the balance. A budget whose owner can also approve overruns is not a budget.