Review this design. 12 services and 8 engineers in one cloud region. The proposal adds a service mesh with mutual TLS between every pair of services plus per-service microsegmentation policy plus an external policy decision point consulted on every internal call plus envelope encryption with one data key per record. What would you remove?
Show the full answer Hide the answer
What is actually required
Three things, for a system of 12 services. Workload identity so no service holds a static credential. Encrypted wires. And correct object-level authorisation on the read path, which is the control that maps to the risk that actually materialises for a twelve-service product.
What I would remove, and why it is safe
The external policy decision point on every call. A synchronous dependency in front of every internal request inherits the availability of the thing it calls, so the platform now fails whenever the policy service does, and 8 engineers cannot run that service to a higher availability than the 12 services depending on it. Keep central policy authorship and move evaluation into a local library or sidecar fed by a pushed policy bundle. The decision then costs microseconds and survives the policy service being down, at the price of a staleness window equal to the bundle push interval.
One data key per record. At that granularity every read is a decrypt plus, without caching, a key-service call, so key-service traffic scales with read volume rather than with data volume. Per-record keys are justified when destroying one key is how you delete one subject's data, or when records have different legal custodians. Neither is stated here. Per-tenant keys with a bounded data-key cache give the same at-rest guarantee for roughly 1000x fewer key-service requests, because the cache amortises one wrapped key over many records.
Hand-authored microsegmentation policy. With twelve services the call graph fits on a whiteboard, and security groups per service express the same control with no mesh to operate. Microsegmentation starts paying when the rule set is too large for anyone to safely delete from, which is a problem at 100 services and not at 12.
The one change that matters
Authorise the object, not the endpoint, in one shared data-access path. None of the four proposals touches the vulnerability class that will actually be found: an endpoint that returns a record the caller is not entitled to. That is a five-figure engineering change and it removes the risk that mTLS, segmentation and envelope encryption all leave untouched.
What I would keep even though it looks odd
Workload identity, even without the mesh. Cloud instance identity or a SPIFFE-style issuer removes the long-lived secret from every service, and every other control on the list assumes it. It is the one item here whose absence makes the rest cosmetic.
When not to remove any of it
Change three facts and the proposal becomes correct. If you run other people's code on shared compute, mutual TLS and microsegmentation are the boundary. If a regulated scope boundary must be demonstrable to an assessor, the mesh produces the evidence. If the estate is past roughly 100 services, hand-written security groups have already become a rule set nobody will prune and a policy engine is the cheaper path. The design is not wrong in the abstract. It is wrong for 12 services and 8 engineers, and saying so is the job. The reference point worth raising in the review is that Envoy was open-sourced in 2016 by a company already operating a large microservice fleet, and the running cost a mesh assumes was sized for an estate of that shape.