intermediate 3 min answer

Review this. A 12-engineer startup with 6 services has built an internal developer platform - a React portal, a YAML DSL that generates Terraform, a bespoke secrets service backed by cloud KMS, a service catalogue and a maturity scorecard with 23 checks. Two of the twelve engineers work on it full time. Deploys take 14 minutes and the DSL covers about 60% of what teams need, with the rest hand-written Terraform merged by those two engineers. What would you remove, what would you change, and what would you leave alone?

internal developer platformpragmatismabstractionbuild-vs-buydeveloper experience
Show the full answer Hide the answer

What is actually required

Six services and twelve engineers need four things: a way to ship a change safely, a way to know who owns what, a way to get secrets into a process, and a way to see what is running. None of the four requires a product. The current shape spends 17% of engineering capacity, about \(500k a year loaded, on internal infrastructure for 6 services. That is **\)83k per service per year**, and the comparison is a managed application platform at a few thousand dollars a month.

What I would remove, and why it is safe to

The YAML DSL. An abstraction that covers 60% of cases is not a 60% win; it is a layer that must be learned, debugged and maintained, and 40% of the time the engineer has to learn the thing underneath anyway. Worse, every failure now has two possible locations, and the team's only two experts are the generator's authors. Replace it with a small set of reviewed Terraform modules and three worked examples. The test for an abstraction is whether a reader can predict what it produces; a generator with a 40% escape hatch fails it.

The bespoke secrets service. A homegrown component in the credential path, maintained by one person, is a security liability that gets audited eventually. The cloud provider's secret manager plus workload identity is strictly better and costs nothing to run.

The portal, and 19 of the 23 scorecard checks. Discovery across 6 services is a README. A scorecard with 23 checks and no consequence is a dashboard nobody opens; keep the four that matter (owner, on-call route, runbook link, resource limits set) and make them visible where the work happens.

The one change that matters

Buy the platform. A managed application platform or the cloud provider's own app runtime absorbs the build, the deploy, the rollback and the base-image refresh. Then spend the recovered 2 engineers on the 14-minute deploy, which is the number every one of those 12 engineers feels several times a day.

What I would keep, even though it looks odd

The service catalogue, as a file in each repository rather than a central registry, because the ownership and on-call route is the one piece of metadata that has to be right during an incident and the one that rots fastest when it lives somewhere else. And keep the single paved pipeline, which is the part of this platform that is actually load-bearing.

The decision rule

Build a platform when at least one of these is true: three or more teams have independently solved the same infrastructure problem, the infrastructure request queue exceeds roughly one full-time engineer, or compliance requires an enforced path rather than a documented one. Typically that lands somewhere past 25 to 30 engineers and 10 to 15 services. Below it, the correct platform is one paved pipeline, a README and a managed runtime.

How I would argue this in the review

Not as "you over-engineered". Present the per-service cost, the two-expert bus factor on the DSL and the secrets service, and the 14-minute deploy as the only user-visible number. Then offer the trade: the team loses its own abstraction and inherits a vendor's conventions, which is a real loss of control and the reason this is a decision rather than an obvious correction.

Common weak answers

  • "Adopt Backstage." Replacing a homegrown portal with a larger portal adds a system to operate for a discovery problem 6 services do not have.
  • "Keep the DSL and raise coverage to 95%." That is the trap: coverage chases workload variety, and the last 35% is the hardest and least reusable.
  • "Remove everything and go back to manual deploys." The pipeline is the piece that is paying for itself. Critique is not demolition.