You have joined a 55-engineer product organisation with a six-year-old monolith. The VP of Engineering wants a microservices roadmap by Friday. Deploys queue for 40 minutes, the test suite takes 55 minutes, and two candidates recently turned down offers citing the monolith. Walk me through the conversation.
Show the full answer Hide the answer
What the interviewer is testing
Whether you can take a topology request and convert it into the measurements that would settle it, without either capitulating or lecturing the VP about distributed systems. The trap is that both refusing and agreeing are wrong answers here: the pain is real, and microservices are a plausible but unevidenced response to it.
The clarifying questions that change the answer
- How is lead time actually spent? A 40-minute deploy queue and a 55-minute suite total about 1.6 hours. If a typical change takes three days to reach production, those 1.6 hours are not the constraint and services will not touch the rest.
- What fraction of changes touch more than one area? This is the one number that says whether boundaries are wrong. If 70% of changes are confined to a single area, the codebase is already modular and extraction buys deployment independence it does not need. If 70% span three areas, extraction would produce services that must be released together.
- Where does the merge queue time go? A 40-minute queue is usually serialisation plus flaky retries, and both are fixable in weeks.
- What did the candidates actually object to? "The monolith" from a candidate is often a proxy for a 55-minute feedback loop and unclear ownership, both of which survive decomposition intact.
A strong answer's arc
Give the VP the roadmap they asked for, and make its first milestone the part that is valuable on either path. Enforced module boundaries checked in CI, named ownership per module, test selection so a change runs the tests it can affect, and a merge queue with parallel workers. That is roughly 2 to 4 engineer-weeks and it cuts the feedback loop, which is the hiring complaint. It is also the prerequisite for any extraction, because you cannot cleanly extract a module whose boundary is not enforced.
Then run one extraction as an experiment, chosen for a named constraint rather than for size: a module with a different scaling shape, a compliance boundary, or a team that genuinely owns it end to end. Measure lead time and incident count for that module over a quarter and bring the numbers back.
Common weak answers
- "Microservices would not help, keep the monolith." Correct conclusion, no evidence, and it loses the room. The VP has three real symptoms and you have dismissed them.
- "Start with the strangler pattern and extract the edges." A method with no argument about whether the destination is right.
- "Fix the CI pipeline and the request goes away." It does not. The ownership and hiring concerns are real, and CI work alone does not address who decides what inside a shared codebase.
What a strong answer adds
The cost of being wrong in each direction, stated in the currency the VP uses. Decomposing a six-year-old monolith with 55 engineers is a 12 to 18 month programme that consumes most of the platform capacity and, done without enforced boundaries first, yields services that deploy together and page each other. Doing the CI and boundary work and being wrong costs a month and leaves the code better for the decomposition if it is still needed. The asymmetry decides the sequence, and it is the sequence rather than the destination that you are being asked for on Friday.