practice

Managed versus Self-Managed

Whether to run a component yourself — a decision that should default to managed and require a specific articulated reason to deviate.

managed-servicesbuild-vs-buytcooperationsdecision-making

Definition

The choice between paying a provider to operate a component and operating it yourself. The premium is visible; the cost of self-managing is not, which is why the decision is systematically biased.

The default

Managed, unless you can name the specific reason not to. The premium is usually less than a fraction of an engineer's time, and operational capacity is the scarcest resource in most organisations.

Self-managing to save money almost always omits on-call, upgrades, incident response and the opportunity cost of the engineers doing it — the term that dominates the calculation and appears on no invoice.

The legitimate reasons to self-manage

  • A required capability the managed version does not expose — an extension, a version, a replication topology, a kernel parameter.
  • Cost at a scale where the premium becomes material. A calculation, and the threshold is much higher than teams assume.
  • Regulatory or residency constraints the provider cannot satisfy.
  • Existing deep operational expertise and a team whose job it is.
  • Unacceptable failure modes — a documented failover time exceeding your RTO, for instance.

The strongest public example of a correct repatriation — a very large storage workload moved off public cloud — met several of these simultaneously: one component dominating the entire cost base, at enormous scale, with a uniform workload and a team built to operate it. Those conditions are narrow, and citing that case as general support for self-hosting proves considerably less than it appears to.

What managed does not remove

Your responsibility for your own resilience. Failover is documented behaviour; the application must tolerate a connection reset at an arbitrary moment. Backups may exist and are almost never restore-tested by you. Maintenance happens on the provider's schedule.

"Managed" means somebody else does the operations, not that the operations do not happen — and reading the failure behaviour before adopting is the step most often skipped.

Managing the lock-in

Prefer managed versions of open interfaces over proprietary ones at similar cost, so the migration path is a restore rather than a rewrite. That preserves the option without paying for a portability abstraction that would probably never be exercised.

Interview question

"Under what circumstances would you self-manage a message broker, and how would you justify it financially?"