metric

Caller Attribution Coverage

also called Owner Resolution Rate, Consumer Identity Coverage

The share of calls to a platform interface that resolve to a named owning service and team, which is the number that decides whether a removal date can be held or will slip.

platform telemetrydeprecationownershipidentityservice catalogue

A platform announces that v1 of its config API will be removed in six months, with telemetry showing 412 calling services. Five months later nine callers remain and four cannot be traced to any human: the requests arrive under a shared service account, the catalogue names a dissolved team, and nobody will authorise breaking something that might be a payroll job.

The deprecation did not fail on engineering effort; it failed on attribution - which this metric would have shown on day one, when it was cheap to fix.

Why it matters

Every platform change affecting consumers needs the same input: who is affected, and whose week the work lands in. Without attribution a platform can only ask the estate to volunteer, and volunteering has a long tail that no amount of communication shortens.

Two measures behave differently. Coverage by call volume flatters you: a few chatty services account for most traffic, so 99% of calls can be attributed while a third of distinct callers are anonymous. Coverage by distinct caller predicts whether the date holds.

Implementation patterns

  • Attribute at the authentication boundary, the only attribution a forgetful client cannot omit.
  • Key the counter on a stable catalogue service id, never a pod name or IP. Pod names churn daily, so a 3000-pod estate gives a huge key space useless for ownership, while 500 stable ids are a trivial series count that survives redeployment.
  • Count unsampled. At 1% a caller making 3 calls a month has roughly a 1 in 6 chance of appearing in six months of data, which is how a trace store becomes a false inventory.
  • Retain beyond the longest caller period you believe exists - 13 to 15 months where annual jobs exist - and resolve identity to owner at query time, so ownership changes are picked up.
  • Publish both numbers with a threshold - commonly above 99% of calls and 95% of distinct callers before a removal date is announced - and alert on any new unattributed caller.

Industry example

The mechanism is standard in platforms with a service catalogue. Backstage, the CNCF developer-portal project, exists partly so a service identifier resolves to an owner, its docs and a support channel. A catalogue alone does not produce coverage: that needs the join between runtime identity and the catalogue entry, and platforms running this in production usually find the first measurement far below expectation.

Failure scenarios

  • Shared service accounts, where 40 callers appear as one identity, or human tokens in automation, so a nightly job is attributed to someone who has left.
  • An owner field never revalidated, giving nominal coverage and an owner who cannot be paged.
  • Counters retained for 30 days, which structurally cannot see monthly or quarterly callers, or attribution set by a client header, so the callers who forget are the ones who will break.

Trade-offs

Choose Gains Pays
Unsampled counters keyed on service id A defensible caller list and holdable dates A series per caller plus 13 months of retention
Sampled traces as the only record Cheap and good for latency work Quiet callers invisible and every deprecation slips
Reject unidentified calls Coverage near 100% by construction A hard cutover that breaks legacy callers

The surprising cost is organisational: coverage is raised by chasing credentials and catalogue entries across teams, with no visible output until the first deprecation runs on time.

When not to use it

An API with a handful of callers in one repository does not need this. A code search is complete, cheap and current, and an attribution pipeline there is ceremony.

It is also wrong where the interface will never change incompatibly. Attribution earns its cost in proportion to how often the platform intends to change, so prefer spending on it when a deprecation or breaking release is coming within the year.

Interview question

Q: You are about to announce a platform API removal in six months. Your only usage data is 1% sampled traces for the last 90 days showing 180 distinct callers. What do you do before announcing anything?

What a strong answer covers: that both the sample and the window are wrong for an inventory - 3 calls a month at 1% gives roughly a 1 in 6 chance over six months · the replacement, an unsampled per-caller counter keyed on workload identity with long retention · that the announcement date is set by coverage crossing a threshold rather than by the calendar · funded migration for owned callers and brownouts for the rest · and that delaying the announcement by a month beats slipping the removal by a quarter.

Quick check

Quiz: Attribution is 99% by call volume. Why might the next deprecation still slip? Volume coverage is dominated by a few chatty services, while the callers that block a removal are quiet and unowned - so the number that matters is coverage by distinct caller.

Flashcard: Why can sampled traces not serve as a deprecation inventory? — A caller making 3 calls a month has roughly a 1 in 6 chance of being sampled at all in six months at 1%, and quiet unattended callers are the ones that block a removal date.