beginner 3 min answer

Review this. An internal library called platform-common holds an HTTP client, a feature-flag reader, a money type, a retry helper and tenant-context propagation. 140 services depend on it. A one-line bugfix to the money type produces version 4.12.0, and until each service upgrades it is running a version missing the fix. The last upgrade wave took eleven weeks. What would you split, and what would you leave alone?

cohesionshared-librariesversioningrelease-cadencecoupling
Show the full answer Hide the answer

What is actually wrong

The five things in this library are related by topic — they are all "platform stuff" — and unrelated by the reason they change. The money type changes when tax or rounding rules change. The HTTP client changes when a transport defect is found. The feature-flag reader changes when the flag vendor's SDK changes. A versioned artifact has exactly one release cadence, so bundling them forces every consumer to accept every other component's change schedule.

That is why an eleven-week upgrade wave is the real defect rather than an operational annoyance. The cost of a one-line fix is not the fix; it is 140 upgrade decisions, and each one is a decision because the version also carries four unrelated diffs a team has to assess.

The test I would apply, and what it shows

For each pair of components, ask: has a change to one ever required a change to the other? Across a year of history the answer here is almost certainly no for every pair. A boundary that has never carried a coordinated change is not a boundary, it is a folder. The evidence is in the commit log and it takes an afternoon to extract: count, per component, how many releases touched it, and cross-tabulate.

What I would split

  • The money type into its own artifact. It is the one with correctness consequences and the one whose changes must propagate fastest. On its own, a tax fix is a version bump with a one-line diff nobody needs to read.
  • The feature-flag reader out, because its change driver is an external vendor and therefore outside the team's control. Anything whose release cadence is set by a third party should not set the cadence for anything else.

What I would leave alone, even though it looks odd

The HTTP client, the retry helper and tenant-context propagation stay together. They genuinely do change for the same reason: they are one policy about what an outbound call means. Splitting them would produce three artifacts whose versions must be compatible with each other, which converts an internal function call into a version-matrix problem. That is the failure mode of over-splitting, and it is worse than the one being fixed.

The change that matters more than the split

Publish independent versions and stop expecting synchronous upgrades. With four artifacts instead of one, a money fix reaches services that consume money and nobody else, and the upgrade wave for it is small enough to automate with a dependency-bump bot. Add one measurement: the age distribution of the version each service is running. If the p90 age exceeds a quarter, the artifact is too coupled to its consumers regardless of how it is split.

When this is the wrong answer

Below roughly ten consumers, one shared library is correct and splitting it is pure overhead: you can upgrade ten services in an afternoon, and the version matrix costs more than the coupling. The number of consumers is what flips this, not any property of the code. Monorepo estates with a single version of every internal dependency also sidestep the whole problem, at the price of making every upgrade an all-or-nothing change.