Managed versus Self-Managed
Whether to run a component yourself — a decision that should default to managed and require a specific articulated reason to deviate.
Definition
The choice between paying a provider to operate a component and operating it yourself. The premium is visible; the cost of self-managing is not, which is why the decision is systematically biased.
The default
Managed, unless you can name the specific reason not to. The premium is usually less than a fraction of an engineer's time, and operational capacity is the scarcest resource in most organisations.
Self-managing to save money almost always omits on-call, upgrades, incident response and the opportunity cost of the engineers doing it — the term that dominates the calculation and appears on no invoice.
The legitimate reasons to self-manage
- A required capability the managed version does not expose — an extension, a version, a replication topology, a kernel parameter.
- Cost at a scale where the premium becomes material. A calculation, and the threshold is much higher than teams assume.
- Regulatory or residency constraints the provider cannot satisfy.
- Existing deep operational expertise and a team whose job it is.
- Unacceptable failure modes — a documented failover time exceeding your RTO, for instance.
The strongest public example of a correct repatriation — a very large storage workload moved off public cloud — met several of these simultaneously: one component dominating the entire cost base, at enormous scale, with a uniform workload and a team built to operate it. Those conditions are narrow, and citing that case as general support for self-hosting proves considerably less than it appears to.
What managed does not remove
Your responsibility for your own resilience. Failover is documented behaviour; the application must tolerate a connection reset at an arbitrary moment. Backups may exist and are almost never restore-tested by you. Maintenance happens on the provider's schedule.
"Managed" means somebody else does the operations, not that the operations do not happen — and reading the failure behaviour before adopting is the step most often skipped.
Managing the lock-in
Prefer managed versions of open interfaces over proprietary ones at similar cost, so the migration path is a restore rather than a rewrite. That preserves the option without paying for a portability abstraction that would probably never be exercised.
Interview question
"Under what circumstances would you self-manage a message broker, and how would you justify it financially?"