Walk me through how you would decide whether a new analytics platform needs a batch processing path at all. The team already runs a stream processor, and nobody has asked for historical reprocessing yet.
Show the full answer Hide the answer
What the interviewer is testing
Whether you understand that the second path is not about freshness. Both paths produce the same numbers; the batch path exists as a correction mechanism and an audit artifact. A candidate who argues from "streaming is strictly more general" has not had to fix a month of wrong output.
The clarifying questions that change the answer
- How is a logic bug corrected today, and how far back? This is the whole question in one sentence. If the answer is "we reprocess 90 days", the single path must be able to replay 90 days.
- What is the source retention? Replay is bounded by the shortest retention anywhere on the path, not by intent. A 7-day topic cannot serve a 90-day correction no matter what the architecture diagram says.
- Are the sinks keyed? Replay into an append-only sink duplicates everything. Replay into an upsert keyed by (entity, window) is safe and cheap.
- What else reads the output? If the processor's output triggers notifications, payments or webhooks, replay re-fires them unless those effects are separated from the computation.
- Does anyone sign the number? An auditor or controller needs a figure that does not change after publication. A provisional aggregate that corrects itself is correct behaviour and unsignable.
A strong answer's arc
One path is right when replay works, and replay works only when retention, keyed sinks and effect isolation all hold. Establish those three, and the batch path is duplicated logic you should delete. If any one of them is missing, "we have a single path" means "we have a single path that cannot be re-run", which is strictly worse than two paths, because the correction capability has quietly disappeared while the architecture looks modern.
Then propose the middle option, because it is almost always the right call: one codebase, two runners. The transformation is a library with no I/O. A streaming runner feeds it from the log; a batch runner feeds it from object storage over a date range. The two numbers agree because the code is the same, and you get the re-runnable correction path without paying 90 days of hot log retention.
Common weak answers
- "Kappa is the modern approach, so one path." Names a conclusion. The interviewer wants the three preconditions.
- "Keep both for safety." Two implementations of the same logic drift, and the drift is discovered during a reconciliation argument rather than by a test. Every metric change becomes two changes plus a disagreement.
- "Use the batch path for history and streaming for recent data." That is Lambda, which is a defensible answer, but only if you also say where the boundary between the two views sits and how the serving layer handles records near it.
What a strong answer adds
The cost of duplication measured as change coupling: a team that ships 30 metric changes a quarter pays 30 double-implementations plus reconciliation. Argue for the version that tests cheapest, which is the shared library, and name the artifact that would prove the single path works: a quarterly replay drill that rebuilds one derived table from the log into a shadow destination and compares it. If that drill has never run, the replay capability does not exist.