1. Backfill & Reprocessing beginner Multiple choice

    A logic bug means the last 10 days of a derived table are wrong. An engineer proposes resetting the running consumer group's offsets back 10 days and letting the job reprocess. What does that actually do?

    2 min answer offsetsconsumer-groupsreplaybackfill
  2. Dead Letter Handling beginner Multiple choice

    A platform team sets one default for every stream consumer: retry a failing record three times then route it to a dead-letter topic and carry on. The ledger team's consumer derives account balances from ordered per-account events. What trade-off does that default make on their behalf and what should they run instead?

    3 min answer dead letterorderingavailabilitycorrectness
  3. Stream Processing Frameworks beginner

    A team runs 400 events per second through a pipeline with Kafka Streams for routing, Flink for windowed aggregation and a Spark job for a nightly correction pass, maintained by four engineers. Two of the three have had production incidents this quarter. What would you remove, what would you keep, and what would you leave alone even though it looks odd?

    3 min answer simplificationoperational costframework choiceright-sizing
  4. Stream-Table Duality beginner

    A service keeps an in-memory map of product prices built by consuming a compacted topic, instead of calling the pricing service on each request. A colleague asks why the map is not just a cache. What is the honest answer, and what does the service now own?

    2 min answer materialized viewcompactioncacheconsistency
  5. Stream-Table Duality beginner

    A team is told that a table and a stream are two views of the same thing. They have a Postgres table holding 12 million account balances and are asked to produce the stream of changes that produced it. Why does the equivalence only run one way in practice, and what has to exist before it runs the other way?

    3 min answer changelogcompactioncdcoutbox
  6. Streaming vs Batch beginner Multiple choice

    An hourly batch job writes each hour's aggregate into a dt/hour partition and is safe to rerun after any crash. The team reimplements the same logic as a streaming job writing to the same table. The streaming job is at-least-once and now produces duplicate rows after every restart. What made the batch version's retries safe?

    2 min answer streamingbatchidempotencysinks