An internal library defines your authentication client, logging and money type. It is on version 14 and 60 services depend on it, 22 of them transitively through two other internal libraries. A security fix in version 15 needs to reach production in a week. Sequence it.
Show the full answer Hide the answer
The sequence
The constraint is the diamond: a service depending on library A and library B, each of which depends on the shared library at a different version, cannot resolve both. Until A and B agree, the leaf service cannot upgrade, which is why the 22 transitive consumers dominate the schedule.
- Map the graph before anything else. Produce the list of 60 consumers with their current version and their path to it. The build tooling can do this; a spreadsheet maintained by hand will be wrong within a fortnight.
- Publish version 15 as a patch of the oldest supported line, not as the newest. If consumers sit on 11, 12, 13 and 14, release 11.x, 12.x, 13.x and 14.x with the same fix. This is the step that makes the week possible, because it decouples the security fix from the upgrade work. Backporting four releases is a day; getting 60 teams to a new major is a quarter.
- Upgrade the intermediate libraries first, in dependency order, and release them with a widened version range on the shared library rather than a pin, so leaf services are not forced to move in lockstep.
- Automate the leaf change. One bot-raised pull request per repository with the version bump and the build result attached. For 60 repositories, hand-raised pull requests are the schedule.
- Rebuild and redeploy. This is the part teams forget: a library fix is only live when every consumer has been rebuilt and redeployed. The patch latency of the estate is the longest redeploy cycle among 60 services, not the time to merge.
- Verify at runtime, not in the manifest. Have each service report its resolved dependency versions at startup to an inventory. A lockfile says what was intended; the running process says what shipped.
Where it can diverge, and how you would know
The silent divergence is a service whose lockfile says 15 while a transitive pin resolves 14 at build time. The deployment is green and the fix is absent. A startup-reported version inventory is the only reliable detection, and the query to run is "which running instances report a shared-library version below 15", not "which pull requests merged".
The second divergence is behavioural: a shared library that bundles an authentication client, logging and a money type forces a consumer taking the security fix to also take whatever else changed. If version 15 altered logging format, you have shipped an unintended change to 60 services.
The structural fix, after the incident
Split the library along its reasons to change. The authentication client changes for security reasons on security timescales; the money type changes almost never; logging changes for developer-experience reasons. Bundling them means every consumer inherits the union of three change cadences. Three small libraries with narrow surfaces let a consumer take the urgent fix alone.
Then change the coupling mode where it is worth it. A capability consumed over the network, through a sidecar or a service, is patched by redeploying the sidecar rather than rebuilding 60 services. That is the real argument for moving authentication out of a library, and the real cost is a network hop and a new failure domain on every request.
How long it really takes
For a well-instrumented estate with automated dependency updates, a backported patch can reach most consumers in a week and the long tail in three to four. Without the version inventory, assume you will not know when you are done, which is the condition most organisations discover during their first urgent fix. In an estate of internal tools in the mould of a developer-tools vendor's own platform, the tail is typically services with no active owner, and the answer there is organisational rather than technical.
When not to move off a library
For genuinely stable, logic-only concerns — a money type, a date utility, a domain primitive — a library is simpler, faster and has no runtime failure mode, and moving it to a service would be absurd. The decision rule is change frequency: the faster a capability changes and the more urgently its changes must propagate, the worse a library is as its delivery mechanism.