advanced 3 min answer

You are the architect for a platform where 30 teams both publish and consume events, and leadership asks you to set the company standard for exposing real-time data between teams. What do you standardise, what do you deliberately leave to teams, and how would you know the standard is working?

contractsownershipslosplatformgovernance
Show the full answer Hide the answer

What the interviewer is testing

Whether you can separate the parts of a streaming platform where variation causes real harm from the parts where it only causes discomfort. Weak answers standardise the fun things (frameworks, languages, topologies) and leave the harmful variation untouched.

The clarifying questions that change the answer

  • What actually broke in the last six months? Schema breakage, cost, or nobody knowing which topic to read? Each points at a different standard.
  • How many consumers does a typical topic have? With one consumer, a stream is a private channel between two teams and needs no contract. With nine, it is an API.
  • Who gets paged when a stream is late? If the answer is the platform team, no standard will help until ownership moves.
  • Is there a regulated consumer downstream? That converts a freshness guideline into a control with evidence requirements.

The arc of a strong answer

Standardise the things that make a stream readable by someone who did not write it:

  • A required envelope on every event: event id, event time, ingestion time, producer, schema id and partition key. This is what lets any consumer deduplicate, order and trace without asking the producer anything.
  • A registry with a declared compatibility mode per topic, so "who upgrades first" is a property of the topic rather than a negotiation each time.
  • A published freshness and completeness objective with a named owner for every topic with more than one consumer. A stream without a stated staleness bound is not a product; it is a side effect someone is reading.
  • A retention floor that makes replay possible — long enough to cover the longest consumer outage you intend to survive without a full rebuild, with tiered storage so that floor is a cost decision rather than a broker-sizing one.
  • Dead-letter conventions: where diverted records go, what they carry, and a budgeted threshold above which the pipeline is failing rather than coping.

Leave to teams: processing framework, language, internal state stores, topology, and whether they use a stream processor at all. None of those are visible across the boundary, and mandating them buys nothing while making the platform team the bottleneck for every change.

Common weak answers

  • "Mandate one framework." It standardises the invisible and leaves event contracts free, which is the exact inversion of the value.
  • "A central review board for every new topic." It works at 30 teams for about a quarter, then becomes a queue, and teams route around it with private Kafka clusters.
  • "One giant company-wide schema." Canonical data models fail because they require agreement that does not exist; per-topic contracts with compatibility rules do not.

What a strong answer adds

Measure adoption rather than compliance. Share of topics with a named owner and a published freshness objective; number of incidents attributed to a schema change; median time for a new consumer to go from discovery to first record. Make the paved road cheaper than the alternative — a generated producer library that emits the envelope for free beats a policy document — and add a sunset path, because topics with no consumer for 90 days are cost and confusion and deleting them is part of keeping the standard credible.

When a standard is the wrong move

With three teams and six topics, a page in the wiki and a code review is the correct amount of governance. A standard is worth its cost once no single person can hold the topic list in their head, and it should be written down only for the failures that have actually happened.