concept

Diamond Dependency

also called Version Skew Diamond, Transitive Pin Conflict

The situation where one component depends on two others that each require a different version of a shared dependency, so neither side can be upgraded until both agree.

shared-librarybuild-time-couplingsecurity-patchingdependency-managementversion-skew

An internal library holds the authentication client, logging and the money type. It is on version 14, and 60 services depend on it — 22 of them transitively, through two other internal libraries. A security fix lands in version 15 and must reach production in a week.

The 38 direct consumers are a scheduling problem. The 22 transitive ones are a resolution problem. A service that depends on library A and library B, each pinning a different version of the shared library, cannot satisfy both. It cannot upgrade until A and B agree, and A and B are owned by other teams with their own release cycles.

Why it matters

A diamond converts an urgent change into an organisational negotiation. The engineering work is a version bump; the schedule is set by the slowest intermediate owner. Teams discover this during their first security incident, at which point the honest answer to "when will the estate be patched" is "we do not know", because nobody has the graph.

The deeper issue is that a shared library is build-time coupling, so the fix is only live when every consumer has been rebuilt and redeployed. The estate's patch latency is therefore the longest redeploy cycle among all consumers, not the time to merge the fix, and the long tail is services with no active owner.

Implementation patterns

  • Backport to every version line in use rather than releasing only on the newest. Publishing 11.x, 12.x, 13.x and 14.x with the same patch costs about a day and decouples the security fix from the upgrade work, which otherwise takes a quarter.
  • Widen ranges in intermediate libraries instead of pinning exact versions, so leaf services are not forced to move in lockstep. Pin only when a known incompatibility justifies it, and record why.
  • Upgrade in dependency order: shared library, then intermediates, then leaves, with the graph driving the sequence.
  • Automate the leaf change. One bot-raised pull request per repository with the bump and the build result attached; 60 hand-raised pull requests are the schedule.
  • Report resolved versions at runtime. Each service publishes its actual dependency versions at startup to an inventory. A lockfile states intent; the running process states fact, and the gap between them is where the silent failure lives.
  • Split the library along its reasons to change. Authentication changes on security timescales, a money type almost never, logging for developer-experience reasons. Bundling them makes every consumer inherit the union of three cadences.

Industry example

The problem is general enough to have shaped tooling across ecosystems: Go's minimal version selection, Maven's nearest-wins resolution, npm's nested installs and Python's flat single-version constraint are four different answers to the same diamond, each with its own pathology. The widely cited treatment is in Software Engineering at Google (2020), which argues for living at head with continuous integration across the estate precisely because version pinning defers the conflict until it is urgent. The practical lesson for a smaller organisation is cheaper: keep internal libraries narrow, and keep a runtime inventory of resolved versions.

Failure scenarios

  • A green deployment with the old version, because a transitive pin resolved below the patched release. The pull request merged, the build passed, and the vulnerability is still running.
  • Unintended changes riding along. A consumer taking an urgent authentication fix also receives a changed logging format, because both live in the same artefact.
  • A blocked leaf, where two intermediates disagree and the leaf service simply cannot upgrade until one of them releases.
  • Duplicate versions loaded at runtime in ecosystems that permit it, so two copies of a class or module exist and type identity checks fail in confusing ways.
  • An unowned tail, where the last 10% of consumers have no team and the patch programme stalls there for months.

Trade-offs

Narrow libraries mean more artefacts to publish, more release pipelines, and consumers managing more dependencies. A single fat internal library is genuinely more convenient on the day it is created. What it costs is paid later, at the worst moment: every consumer inherits every change, and urgent fixes cannot be isolated. The alternative delivery mechanisms (a sidecar, a service) remove the rebuild entirely and add a network hop, a control plane and a new failure domain on every request.

When not to use it

A shared library is the right mechanism for stable, logic-only concerns — a money type, a date utility, a domain primitive — where there is no runtime failure mode and a service would be absurd. Do not move a capability out of a library because of this risk alone; move it when the capability changes often and its changes must propagate quickly. For a single team with a handful of services deployed together, the diamond is theoretical and the inventory would be ceremony.

Interview question

Q: You are told a critical vulnerability exists in an internal library used across 60 services, and leadership wants a date. Walk me through what you would need to know before giving one, and what you would build so that next time you can answer in an hour.

What a strong answer covers: demanding the dependency graph with transitive paths before estimating · the backport strategy as the thing that makes a short deadline feasible · widened ranges versus pins in intermediate libraries · automation for the leaf bumps · the runtime version inventory as the only reliable completion signal, and the explicit statement that merged pull requests are not evidence · and the structural follow-up of splitting the library by reason to change, with the trade-off against a sidecar stated honestly.

Quick check

Quiz: Why can a service be fully patched in its lockfile and still run the vulnerable library? — Because a transitive constraint can resolve the shared dependency to an older version at build time, so the intent recorded in the manifest is not what the process loads.

Flashcard: Which step makes a one-week estate-wide security patch feasible? — Backporting the fix onto every version line still in use, so consumers take a patch rather than an upgrade.