metric

Supported Path Maintenance Load

also called Golden Path Upkeep Cost, Path Matrix Load

The recurring engineering cost of keeping every supported path working against every change in the platform beneath it, which is what actually caps how many paths a platform can offer.

golden pathpaved roadplatform engineeringtemplatescapacity

A platform team of five offers four golden paths: stateless services, stateful services, batch jobs and inference workloads. The platform underneath changes six times a year - a cluster version, a mesh version, a base image series, a pipeline runner, a secrets backend. Every one has to be re-proven against every path, and nobody budgeted for it, because the paths were funded as projects.

The quantity is a product, not a sum: paths × environments × platform changes per year × days per revalidation. Four paths, six changes and 1 to 3 days each is 24 revalidations, roughly 24 to 72 engineer-days a year, about 0.1 to 0.3 of an engineer per path, before anyone adds a feature. If each path must also work across three environments, the matrix triples.

Why it matters

Platform teams decide how many paths to offer by asking which workloads deserve support. That is the wrong question, because the cost is not in creating a path. A path that is not revalidated becomes a trap: it still appears in the portal, teams still start from it, and it produces a service that fails on its first upgrade. The team that followed the supported route is punished for it, which is the fastest way to lose a paved road.

Naming the load also makes refusal defensible. "We cannot support a fifth path with five engineers" is arithmetic, and it survives a conversation that "we would rather not" does not.

Implementation patterns

  • Regenerate from every path in CI, on a schedule. A monthly job creates a fresh service from each path, deploys it to a production-like environment, runs its tests and tears it down. It is the only honest test of whether a path still works.
  • Track days since last successful regeneration per path, and treat anything over one platform version as broken, not stale.
  • Count the matrix when a new path or environment is proposed, and present the recurring cost rather than the build cost.
  • Prefer one parameterised path over two when the difference is configuration rather than operational shape. Two templates sharing 80% of their content double the load and halve the attention each gets.
  • Retire paths, and keep them thin: a path with fewer than three services should be a documented example, not a supported road.

Industry example

Microsoft's Radius, open-sourced by its Azure Incubations team in October 2023, makes this unit visible: operators author Recipes saying how each resource type is provisioned, defined per environment, while developers reference the resource abstractly. That is an explicit statement of the maintenance unit - the thing someone must keep current is the recipe-environment pair, so upkeep grows with the product of resource types and environments rather than with the number of applications. Any platform offering paths has this matrix whether or not it is written down.

Failure scenarios

  • The zombie path. Offered in the portal, last regenerated two platform versions ago, broken for everyone who starts from it.
  • Silent drift. Generated repositories diverge from the template, so a template fix reaches no existing service and "we fixed it" is not true in production.
  • Upgrade paralysis. The platform stops upgrading its own components because revalidating every path costs more than the quarter allows, so the fleet falls behind on patches.

Trade-offs

Choose Gains Pays
Few paths kept current Every supported route actually works More teams off-road, writing their own and asking for help
Many paths Broad coverage and happier specialist teams Thin attention on each, and the newest path silently rots
One parameterised path One matrix row, one test Conditional complexity inside the template

A workable rule: five platform engineers keep three to five paths current. Beyond that, add paths only by adding people or retiring a path.

When not to use it

Below roughly ten services there is one path and no matrix to measure, so tracking this is overhead for a number that is always 1. It does not apply to reference repositories that are explicitly unsupported, as long as the portal says so. The metric earns its keep when a second workload class appears, or when the platform changes more often than the team can revalidate by hand.

Interview question

Q: A team asks for a fifth golden path for their event-driven workload. You have five platform engineers. How do you decide, and what do you say if the answer is no?

What a strong answer covers: the recurring cost as paths times platform changes times environments rather than the build cost; the regeneration job as the test of whether existing paths work; parameterising an existing path when the operational shape is the same; retiring a low-use path to make room; an enabling engagement instead of a permanent road; and framing refusal as capacity arithmetic.

Quick check

Quiz: What is the maintenance unit of a golden path? The path multiplied by each environment it must work in, revalidated on every platform change - not the services that used it.

Flashcard: How do you know a supported path still works? A scheduled job creates a service from it, deploys it and runs its tests; if the last successful regeneration predates the current platform version, the path is broken.