beginner 2 min answer Multiple choice

A logic bug means the last 10 days of a derived table are wrong. An engineer proposes resetting the running consumer group's offsets back 10 days and letting the job reprocess. What does that actually do?

offsetsconsumer-groupsreplaybackfillidempotency
Pick one
Show the full answer Hide the answer

The mechanism

An offset is a cursor stored per (group, topic, partition), written to the broker's internal offsets topic. It records how far that group has read. It records nothing about what the job did with those records, which sink it wrote, or whether that write can be repeated. Moving the cursor backwards means the same job, with the same configuration and the same sink, reads and writes the same records again.

Three outcomes follow from the sink's shape, not from the reset:

  • Upsert sink keyed by the same business key. The rewrite overwrites itself, and this costs only pipeline time. Ten days of records at 5,000 per second is 4.3 billion records, so at a catch-up rate of 20,000 records per second the replay occupies the job for roughly 60 hours, and live traffic waits behind it.
  • Append-only sink. You have just created a second copy of 10 days of rows, and the table is now wrong in a new and harder way.
  • A job with side effects. Emails, webhooks, push notifications, ledger postings. Ten days of them go out again, which is the failure mode that turns a data repair into a customer incident.

There is also an operational catch that surprises people: you cannot move a live group's offsets while it has members. The group must be stopped for the reset, so the live path is down for the whole catch-up.

Why the other options fail

  • "A reset starts a new group generation." A generation increments on every rebalance and exists to fence off stale members. It has no relationship to the sink. Nothing anywhere creates a fresh copy of a destination on your behalf.
  • "Only records whose keys changed." Brokers track offsets, not per-key delivery. Compaction is per key, but it decides what the log retains, not what a consumer has seen.
  • "Offsets cannot move backwards." They can be set to any valid position, which is precisely why this mistake is available.

The reversible shape

Run a second consumer group with a different group id over the same partitions, writing to a shadow destination, with every side effect switched off by configuration rather than by a code branch. Live traffic never stops. Compare the shadow output against the live table, and swap by pointer or by a view definition when the comparison passes. Rollback is pointing the view back.

When this is the wrong repair

If the wrong rows can be identified by a predicate, fix them with a targeted write instead of a replay. A bug that mis-set one field for one merchant is an update statement over a few million rows, which is hours of work rather than days of pipeline time. Choose replay only if the derivation itself was wrong and cannot be inverted; prefer the targeted write whenever the correction is expressible as a query over the output, because it skips the side-effect risk entirely. This is the ordinary case in production and it is routinely skipped in favour of the more impressive option.