Your gateway samples 1% of platform API requests into traces and that trace data is the only usage record you have. One caller is a monthly batch job that makes 3 calls a month. Roughly what is the chance it appears at all in six months of this data and what does that number rule out?
Show the full answer Hide the answer
The arithmetic, stated
Assumptions: sampling is independent per request at 1%, and the caller's rate is steady at 3 calls a month.
- Calls in the window: 3 × 6 = 18 requests.
- Chance of never being sampled: 0.99^18 = 0.835.
- Chance of appearing at least once: about 16.5%, call it 1 in 6.
Push it the other way to see the shape. For a 99% chance of catching a caller you need roughly 460 sampled
opportunities at this rate - ln(0.01)/ln(0.99) ≈ 459 - which for a caller making 3 calls a month is
about 13 years of observation. The dominant error term is not the sampling maths, it is the assumption
of a steady rate: a job that runs on the first of the month only, or quarterly, or at year end, makes the
number worse rather than better.
What the number rules out
It rules out sampled traces as the inventory for any deprecation. The callers that sampling hides are exactly the ones that break a removal: low-frequency, unattended, usually unowned batch jobs. High-volume callers appear instantly and were never the risk.
It also rules out the common compromise of raising the sample rate. Going to 10% still leaves this caller with a roughly 84% chance of being seen over six months, and it multiplies trace storage tenfold to buy information about callers you already knew about.
What to measure instead
Counting and tracing are different jobs. Keep sampled traces for latency and causal debugging, and prefer an unsampled per-caller counter at the gateway for inventory - the arrangement platform teams actually run in production, because the counter costs a few hundred metric series and the trace pipeline costs storage per request:
- Key the counter on workload identity, not on the request. Cardinality is then bounded by the number of calling services - hundreds to low thousands - which is a trivial metric series count, while per-request data is not.
- Retain it longer than the longest caller period you believe exists, plus a margin. For an estate with annual batch jobs that means 13 to 15 months, not 30 days.
- Resolve each identity to an owner in the catalogue at query time, so the deprecation list is "service plus team" rather than "a service account nobody recognises".
- Use a scheduled brownout for the residue. Counting finds callers that call; a short planned failure finds callers whose health depends on the call in ways no counter shows.
Why the other options fail
- "Roughly 1 in 2" is the intuition that a long window compensates for a thin sample. It does not, because what matters is the product of rate and sample rate: 18 opportunities at 1% is 0.18 expected hits, and no amount of calendar changes that for a caller this quiet.
- "Roughly 99 in 100" is the law of large numbers applied to the wrong population. It holds for the gateway's aggregate traffic and says nothing about one caller's chance of appearing.
- "Roughly 1 in 100" reads the sample rate as the answer, which undercounts by treating 18 chances as one. The conclusion it leads to is also the expensive mistake: full-fidelity tracing of all traffic to solve an inventory problem that a bounded counter solves for almost nothing.
When this is the wrong answer
If the API is young, has a handful of callers, and every one of them is in a monorepo you can search exhaustively, then static discovery is sufficient and cheaper. The arithmetic matters once callers are numerous, unattended, or outside the code you can read - which is the normal state of an internal platform after two years.