An interviewer gives you this - "Every service has to propagate W3C trace context. You have 200 services across 6 languages and no authority to mandate anything. Get me to full coverage, and tell me how you would know you got there." Walk through it.
Show the full answer Hide the answer
What the interviewer is testing
Whether you reach for enforcement or for defaults. The weak arc is checks, dashboards and escalation. The strong arc is that coverage is a property of the artifacts teams start from, not of the rules they are measured against, and that a platform with no mandate authority has exactly one lever: making the compliant path the path of least effort.
The clarifying questions that change the answer
- What is the shared surface? If all 200 services already use a platform-provided HTTP client and server middleware, this is 6 pull requests to 6 libraries plus a version bump. If each team wrote its own, it is 200.
- What fraction of hops traverse something the platform owns (an ingress, a gateway, a sidecar)? Those hops can be instrumented without touching a service.
- How is coverage defined? "Services with the library installed" is a procurement metric. "Share of requests whose trace has no broken parent link" is the metric that matches what a user of tracing needs, and it is traffic-weighted, so 20 services may be 70% of the answer.
- Which services cannot change - vendor images, an end-of-life runtime, something in a regulated freeze?
A strong answer's arc
- Put propagation in the default artifact. Header propagation goes into the shared client and server middleware for each of the 6 languages, enabled by default with no configuration. 6 units of platform work replace 200 units of team work, because the compliant behaviour now arrives with a dependency bump rather than with a task on someone else's backlog. The cost is that the platform owns six libraries in production forever, including the one written in the language nobody on the team knows well.
- Harvest the free hops. The ingress and any mesh sidecars inject and forward context for services that have not been touched yet, which lifts traffic-weighted coverage fast and makes the remaining gaps visible as broken links rather than as missing services.
- Order by traffic, not by alphabet. Instrument the services carrying most of the requests first; coverage measured by request share moves roughly ten times faster than coverage measured by service count.
- Ratchet, do not gate retroactively. A blocking check applies to new services and to services being changed, so it never produces a backlog of 140 failing repositories and never blocks a team that did not touch the code. Existing violations get a dated, enumerated exemption list that shrinks.
- Fund the long tail. For the last 10 to 20 services, the platform opens the pull request. At that point a platform engineer's day is cheaper than another quarter of chasing.
Common weak answers
- "Make it a CI gate." A gate on 200 non-compliant repositories blocks delivery on day one and gets switched off by Friday. It also checks source, not runtime, so a service can pass and still drop the header.
- "Take it to the architecture review board." Review produces agreement, not propagation. Nothing about a decision record changes a header.
- "Publish a dashboard and escalate the laggards." Measurement without a default path converts a technical problem into a political one, and the platform always loses that one.
What a strong answer adds
Two things. First, the regression story: coverage is not a project, it is a number that decays with every new service and every library fork, so it needs an owner and an alert at a threshold, not a launch. Second, when not to do any of this: with 12 services in one language, 12 pull requests beat building middleware, a ratchet and a coverage metric. The machinery is justified by the count of teams you cannot talk to individually, not by the elegance of the approach.