advanced 3 min answer

At Kafka Summit in May 2017, The New York Times described a publishing pipeline in which every piece of published content is appended to a Kafka log that is the source of truth, and the systems behind each reading experience are built by consuming that log. That buys a rendering strategy most sites cannot afford. What does it buy exactly, and when does the bill arrive?

new-york-timeslog-as-source-of-truthstatic-renderingcdnreplay
Show the full answer Hide the answer

What is gained

If every published item is replayable in order, every rendered artefact is derivable. That single property is what makes the expensive rendering strategies affordable.

  • A new front end, a new search index or a new mobile surface can be built by replaying the log, without asking the producing systems for a backfill. The talk's account of the previous API-based model is that consumers had to know about every producer and there was no efficient way to read old content in bulk, which is what made replacing a service's datastore hard.
  • Pages can be pre-rendered and served from a CDN, because regenerating them is a derived operation rather than a migration. For an election night, that converts a traffic spike into a cache problem: the origin serves only what changed.
  • A bug in a rendered view is fixable by rebuilding the view, not by writing a correction script against production state.

What is paid

  • Schema compatibility for the life of the log. A consumer written in 2026 must still read items written in 2017. You are committed to forward and backward compatible published payloads, and a genuinely breaking change costs a rewrite of every consumer plus a full replay.
  • Replay time, which is real and must be rehearsed. Rebuilding a view over millions of items is hours of work. A team that has never timed it will discover the number during an incident.
  • A single global order couples unrelated content. A consumer that falls behind is behind on everything, not only on the items it cares about.
  • N copies of the content, N freshness problems. Every consumer owns a store, so "is the article live?" has as many answers as there are reading experiences.

When the cost becomes visible

Three moments, reliably. The first breaking content-model change, when the compatibility promise turns into a multi-team programme. The first consumer rebuild during a live event, when replay duration becomes the incident. And the first time two surfaces disagree in public and somebody has to say which one is behind — a question that needs per-experience consumer lag as a published signal, not a cluster-level dashboard.

How to keep the option to reverse

  • Version the published payload and keep it separate from internal models, so the log's contract is a deliberate artefact rather than a leaked database schema.
  • Keep a compacted current-state topic beside the full history, so a new consumer bootstraps from the current set and only the rare consumer needs the whole replay.
  • Make consumer lag a service level indicator per reading experience, with a threshold expressed in what a reader would notice, such as "an article is live on the site within 30 s of publication".
  • Keep one consumer deliberately thin — a plain renderer with no derived state — so there is always a path back to serving directly from the log.

When the log is the wrong answer

A site with one reading experience and a content database is better served by incremental static regeneration plus a CDN, and nothing else. The log pays for itself when independent consumers of published content exceed roughly a handful and are built, replaced and operated on different schedules. If one team owns every consumer, a log adds a hop, a compatibility obligation and an operational surface to solve a problem that a database and a cache invalidation already solved. The honest test: name the consumer you would add next year, and the backfill you would otherwise have to run for it. If you cannot, you do not need the log yet.