An interviewer gives you this. "Our platform takes a short YAML file per service and a controller reconciles it into real infrastructure. 300 services use it. Product wants to change the default CPU request for the services that never set one. Walk me through how you would ship that."
Show the full answer Hide the answer
What the interviewer is testing
Whether you know that in a reconciled declarative interface the default is part of the API, not an implementation detail. In a request-response API a default affects the next call a caller makes. In a reconciler, running state is recomputed from the declaration, so changing a default rewrites live infrastructure for every service that omitted the field, with no deploy and no action by the owning team. A candidate who treats this as a configuration tweak has missed the whole mechanism.
The clarifying questions that change the answer
- Is the field absent, or explicitly set to the old default? Only the absent ones move, and that count is the blast radius. It is computable before shipping anything.
- What triggers reconciliation - a periodic resync or an edit? With a resync every few minutes the change lands fleet-wide within the hour. If reconciliation is edit-triggered, it lands over weeks as unrelated changes arrive, and the team that breaks on Thursday will not connect it to a platform merge from eleven days earlier. Smeared blast radius is worse than a fast one, because attribution is lost.
- Does applying it restart anything? A CPU request change recreates pods. If 210 of 300 services omit the field, the merge triggers 210 rolling restarts.
- Which way does the default move? Raising a request can make pods unschedulable on full nodes; lowering it exposes services to CPU starvation that their owners never chose.
A strong answer's arc
- Query before deciding. Count the resources that omit the field, grouped by criticality. This is a one-line query against the declarations and it converts a debate into a number.
- Materialise the current default. Write the existing effective value into every resource that omits the field. This is a semantic no-op: nothing changes behaviour, and afterwards no live service depends on the default any more.
- Change the default for new resources only. Existing services keep what they had, new ones get the better value, and nobody is surprised.
- If existing services genuinely must move, that is a migration with a cohort plan, not a default change: ten services, then a hundred, with a per-service diff published in advance, an owner notified, and a documented revert.
- Build the pre-flight diff into the platform. Any change to the renderer or its defaults should be rendered against all 300 declarations in CI and the diff posted on the pull request. A platform whose own changes cannot be diffed against its consumers is shipping blind on every merge.
When this is the wrong answer
Materialising a default is the wrong move when the default's whole purpose is to keep moving - a pinned base image version, a security policy floor, a certificate lifetime. Freezing those into 300 files converts a central lever into 300 pull requests. For values the platform must be able to change centrally, keep the default dynamic and pay for the cohort rollout machinery instead.
What a strong answer adds
Versioning the schema so a resource pins the generation of defaults it was created against, which makes the default change opt-in at a per-service pace. Noting that the same reasoning applies to anything computed rather than declared - injected sidecars, labels, probe shapes, retry policy. And the operational tell to watch for after any such change: a rise in platform-attributed restarts in the week after a platform merge, which is the signal that a default moved under someone.