practice

Replatform

Moving a workload with targeted changes that capture most of the platform's benefit without redesigning the application.

migrationreplatformmanaged-servicesoptimisationpragmatism

Definition

Replatforming makes deliberate, bounded changes during or after a move: adopting a managed database, containerising, right-sizing, replacing a self-managed queue with a service, moving files to object storage. The application's architecture is unchanged.

Why it is usually the best value

It captures the majority of the benefit for a fraction of the cost of refactoring:

  • Managed services remove the largest operational burden. Backups that exist, patching that happens, failover that has been tested by someone else.
  • Right-sizing against real utilisation typically removes 30–50%.
  • Containerisation improves deployment consistency and density without redesign.
  • Object storage for files removes a whole class of capacity management.

Each is a bounded change with a known payoff, which makes it fundable in a way that a redesign frequently is not.

The judgement about which changes to make

Prioritise by operational burden removed per unit of effort:

  1. The component consuming the most operational time. Usually a database or a message broker.
  2. The component with the worst reliability record.
  3. The largest cost line, where right-sizing or a tier change applies.
  4. The dependency approaching end of support, since the deadline is external.

Explicitly do not replatform things that work, are supported and rarely change. Age is not a defect, and modernising a stable low-change component is frequently the least valuable work available.

The traps

  • Scope creep into refactoring. "While we are here" turns a six-week replatform into a nine-month redesign, which is how these programmes lose their funding.
  • Managed service failure modes not read. Failover times, maintenance windows and connection limits are documented behaviour; the application must tolerate them.
  • Connection limits, particularly when moving to serverless or high-replica-count topologies. A pooler is usually mandatory and is discovered late.
  • Behaviour differences between a self-managed engine and its managed equivalent — extensions, versions, configuration flags.

Failure scenarios

  • Replatforming everything rather than the components with the largest burden.
  • Managed database adopted without retry and reconnect handling, so routine failovers become incidents.
  • Cost assumed to fall without measuring; a managed service can be more expensive at high utilisation.

Interview question

"You have twelve weeks and a large estate. Which replatforming changes do you make and why those?"