advanced 3 min answer

Review this. A 40-engineer company runs on one provider but has banned managed services for portability - self-hosted Postgres and Kafka on Kubernetes, an S3-compatible object store on block volumes, and an in-house wrapper over two infrastructure-as-code providers. The claim is that it could move in a quarter. It has never run anything in a second provider. What would you remove, what would you change, and what would you leave alone?

multi-cloudlock-inmanaged-servicesarchitecture-reviewcost
Show the full answer Hide the answer

What is actually required

Three different risks get bundled into the word portability, and each has a different and much cheaper control. Pricing power at renewal is a commercial problem solved by a costed exit document and a credible second quote, not by code. A contractual or regulatory exit obligation is satisfied by an assessed, evidenced plan. Surviving a provider-wide failure is the only one of the three that needs anything running in a second provider, and this company has no evidence it has ever suffered one.

What I would remove

The infrastructure-as-code wrapper. It abstracts over the three things that differ most between providers — identity, networking and the managed-service surface — so it converges on the lowest common denominator and then has to be maintained by the people who could be shipping product. A portability layer does not reduce migration cost; it relocates it from a one-off project into a permanent tax.

Put a number on that tax in the review. The self-hosted estate here is realistically three engineers of standing operational load — Postgres major-version upgrades, Kafka rebalances and partition arithmetic, the object store's durability story — which at fully loaded cost is on the order of $600k a year, spent every year, against a migration the company has never been asked to perform. Stated plainly: the portability spend has already exceeded the lock-in it was bought to avoid.

I would also remove the self-hosted object store. The S3 API is the portable thing; running your own durability is not a portability requirement, it is a liability, because the team now owns availability and durability claims it cannot evidence.

The one change that matters

Write the exit plan and rehearse one piece of it. Restore the production database into a second provider's managed Postgres once a year and time it. That converts "a quarter" from an assertion into a measured number, produces the artefact auditors and boards actually ask for, and costs about a week. Every company that has done this discovers the real blocker is not the database but identity, secrets and DNS.

What I would leave alone

Kubernetes and container packaging, Postgres as a wire protocol, the S3 API as the storage interface, and open formats on disk. Portability that comes from a standard interface is free; portability that comes from an abstraction you maintain is not. Leave Kafka too if the streaming requirement is real — the argument against self-hosting it is operational cost, not portability, and that is a separate review.

How I would argue this in the review

Not on principle, on evidence. Ask for the last three incidents and the last renewal. If no incident was provider-wide and no renewal turned on price, the programme has no supporting evidence and the counter-proposal is concrete: adopt the managed database, delete the wrapper, fund the annual exit rehearsal, report the headcount released.

When this is the right architecture

A company whose contracts require workloads to run in a second provider or a sovereign environment is not optimising cost, it is meeting a condition of sale. Then the lowest common denominator is the specification. Say so explicitly in the architecture decision record, so nobody later mistakes a contractual constraint for an engineering preference and tries to optimise it away.