Every source behind this page, graded. GitHub rows were fetched and read in
full; the rest were retrieved through a server-side search tool, which the ledger records
per claim. Filter by kind.
Source code
LinkedIn2026-09
linkedin/kafka, pull request 542: the 3.9 side of the 3.0-li protocol bridge
States that 3.0-li changed LeaderAndIsr v3, UpdateMetadata v6 and StopReplica v2 by
inserting MaxBrokerEpoch, that Apache later gave those same version numbers different
meanings, and that the remedy is a dynamic cluster config which makes the controller send
the newest versions shared with the old brokers. Closed unmerged and split into a
reviewable stack.
Carry forwardA private extension to someone else's version
number is a dated liability, and the interest is paid at upgrade time.
https://github.com/linkedin/kafka/pull/542
Source code
LinkedIn2023-10
linkedin/kafka, pull request 479: RPCs for xinfra topic creation and deletion
Adds create, delete and list operations that write a znode under
/xinfraTopics/topicName carrying the namespace as its value, reusing the
existing topic change listener in the Kafka server. Dated twenty months before Northguard
and Xinfra were publicly announced.
Carry forwardThe abstraction that lets you leave gets installed
inside the incumbent, early, as metadata.
https://github.com/linkedin/kafka/pull/479
Source code
LinkedIn2023-11
linkedin/kafka, pull request 494: change federated topic znode structure
Moves the namespace from a node's value into its path, for space and for control-plane
simplicity, and states the exit criterion in the description: the change can be deprecated
once all clients move to xinfra clients and all topic lifecycle operations go through
xmd.
Carry forwardWrite the deletion condition into the
compatibility change itself, or the change outlives the reason for it.
https://github.com/linkedin/kafka/pull/494
Source code
LinkedIn2026-09
linkedin/kafka branches and open pull requests
The only branch ending in -li is 3.0-li; the nineteen most
recently updated branches are all 3.9-li-bridge/... or
3.0-li-bridge/.... Open titles include closing default-off gaps in operational
compatibility, failing closed on unresolved upgrade admission decisions, and restoring
LinkedIn ZooKeeper operational behaviour.
Carry forwardA fork's real cost shows up as a branch programme
five years later, not as a merge conflict today.
https://github.com/linkedin/kafka/branches/all
Source code
LinkedIncurrent
linkedin/li-apache-kafka-clients
A wrapper library over the vanilla Kafka clients, fully compatible with their interfaces,
adding large-message segmentation and pluggable auditing. Its README tells readers who do
not need those functions to prefer the vanilla clients.
Carry forwardHouse behaviour belongs in a removable layer that
implements the standard interface, and the README should say when to remove it.
https://github.com/linkedin/li-apache-kafka-clients
Source code
LinkedIncurrent
linkedin/rest.li and its closed-unmerged pull requests
No active development on new features, with gRPC named as the direction and explicitly
"not a drop-in replacement". The frozen repository nonetheless carries 2025 and 2026 pull
requests about gRPC name resolution and xDS client metrics, which is what a shared
discovery plane looks like from the inside of the old framework.
Carry forwardA framework you are leaving still needs
maintenance on the seam, long after you stop adding features.
https://github.com/linkedin/rest.li
Source code
LinkedIncurrent
linkedin/venice and linkedin/coral
Venice offers three clients against one API at three cost points, ending with Da Vinci:
stateful local cache, zero network hops, under one millisecond. Coral defines a SQL
intermediate representation independent of any dialect, translating HiveQL and Spark SQL
in, and Hive, Spark and Trino out.
Carry forwardTwo seam shapes: one on the data path that trades
latency, one on the authoring path that costs nothing at runtime.
https://github.com/linkedin/venice
Source code
Voldemort projectcurrent
voldemort/voldemort README
"Voldemort is no longer under development." LinkedIn was its primary maintainer and user
and "stopped all production usage in 2018"; most read-only use cases and some read-write
ones migrated to Venice.
Carry forwardThe predecessor's repository is where you find the
honest end date, and it is usually years before the archive notice.
https://github.com/voldemort/voldemort
Decision record
Apache Kafka2019-2021
KIP-500: replace ZooKeeper with a self-managed metadata quorum
The compatibility plan calls for a "bridge release" in which the ZooKeeper dependency is
well isolated: the release does not remove ZooKeeper but eliminates most of the touch
points, confining access to the controller rather than brokers, clients and tools.
Carry forwardThe canonical name for the artefact, from the
project that had to remove its own foundation. Isolate first, remove second.
https://cwiki.apache.org/confluence/display/KAFKA/KIP-500
Decision record
gRPC2020-03
gRFC A27: xDS-based global load balancing
gRPC moves off its own grpclb protocol to converge with the industry trend, on the stated
grounds that the xDS API is evolving into a standard that other data plane software will
use. The rationale concedes the design would have been much simpler had load balancing
been the only thing configured through it.
Carry forwardThe written form of the decision LinkedIn's D2
format eventually forced: adopt the standard before your format becomes the obstacle.
https://github.com/grpc/proposal/blob/master/A27-xds-global-load-balancing.md
Postmortem
Reddit2023-03
You broke Reddit: the Pi-Day outage
About 314 minutes of downtime after a Kubernetes 1.23 to 1.24 upgrade on a cluster built
with kubeadm rather than a standard template. There was no supported way to downgrade, so
recovery meant restoring from backup with an outdated procedure, and the actual cause was
found hours later.
Carry forwardRehearse the reverse direction. An upgrade without
a tested downgrade converts a config mistake into a restore.
https://postmortem.io/incidents/reddit--2023-03-21--pi-day-outage/
Postmortem
Roblox2022-01
Roblox return to service, 28 to 31 October 2021
73 hours, attributed to a streaming feature taken on in a Consul 1.9 to 1.10 upgrade,
under heavy read and write load, plus a BoltDB performance problem. Recovery came through
a configuration change disabling the feature. Monitoring that would have helped depended
on the system that was down.
Carry forwardThe runtime switch is the recovery path. Also:
never let the observability of a migration depend on the thing being migrated.
https://blog.roblox.com/2022/01/roblox-return-to-service-10-28-10-31-2021/
Postmortem
LinkedIn2016-05
Kafkaesque days at LinkedIn
Joel Koshy's account of the 2015 incidents: offset rewinds where resetting to the earliest
offset duplicates and resetting to the latest loses, with fixes landing in offset
management, log compaction and monitoring rather than in applications.
Carry forwardConsumer position is state that a migration must
own explicitly, which is why Xinfra's dual writes are reported as epoch-ordered.
https://engineering.linkedin.com/blog/2016/05/kafkaesque-days-at-linkedin--part-1
Case study
InfoQ2025-06
LinkedIn announces Northguard and Xinfra: scaling beyond Kafka
The reported scale at which the incumbent stopped being operable: 32 trillion records and
17 PB a day, 400,000 topics, 150 clusters. Xinfra abstracts the underlying log systems so
applications use one pub/sub interface whether data sits in Kafka or Northguard.
Carry forwardThe replacement shipped with the abstraction, not
after it. Coexistence was a launch requirement.
https://www.infoq.com/news/2025/06/linkedin-northguard-xinfra/
Case study
InfoQ2026-02
How LinkedIn rebuilt service discovery
More than ten years on ZooKeeper with the house D2 format, whose custom schemas are
reported as incompatible with modern data planes such as gRPC and Envoy. The rebuild moves
writes through Kafka and reads through an xDS-pushing observer, migrated as a dual read
with ZooKeeper as source of truth and silent comparison in the background.
Carry forwardShadow mode is how you find reconnection storms
and inconsistencies while they are still free.
https://www.infoq.com/news/2026/02/linkedin-service-discovery/
Case study
CNBC2023-12
LinkedIn shelved its plan to migrate to Microsoft Azure
A June 2022 memo from LinkedIn's CTO said the company would keep using some Azure services
and focus its efforts on scaling and innovating its on-premises infrastructure. Executives
framed it as on hold rather than cancelled; reporting attributes the difficulty to lifting
and shifting existing tools rather than refactoring them. Tech Monitor reports a
spokesperson saying the move will not go ahead.
Carry forwardThe only migration here with no compatibility layer
is the only one that stopped, and sources still disagree on whether it is paused or
dead.
https://www.cnbc.com/2023/12/14/linkedin-shelved-plan-to-migrate-to-microsoft-azure-cloud.html
Eng blog
LinkedIn2019
How LinkedIn customizes Apache Kafka for 7 trillion messages per day
LinkedIn's 2019 position: this is "not a fork", each release branch comes off the matching
Apache branch with an -li suffix, and the aim is to stay as close as possible
to upstream. Over 100 clusters, more than 4,000 brokers, 100,000 topics, 7 million
partitions.
Carry forwardRead this next to the 2026 branch list. Intending
to stay close to upstream is not a mechanism for staying close to upstream.
https://www.linkedin.com/blog/engineering/open-source/apache-kafka-trillion-messages
Eng blog
LinkedIn2023-04
Unified streaming and batch pipelines, reducing processing time by 94%
One pipeline definition in Apache Beam running on two engines, Samza for streaming and
Spark for backfill, cutting a backfill from 7.5 hours to 25 minutes on roughly half the
memory and CPU.
Carry forwardA seam on the authoring path buys you engine
choice for free at runtime. This is the cheapest version of the pattern.
https://engineering.linkedin.com/blog/2023/unified-streaming-and-batch-pipelines-at-linkedin--reducing-proc
Eng blog
LinkedIn2022-09
Open sourcing Venice, LinkedIn's derived data platform
Roughly 500 Voldemort read-only use cases were fully migrated to Venice by 2018 using its
Full Push capability as a drop-in replacement, and in 2020 a second client library, Da
Vinci, served the same APIs from local state instead of remote queries.
Carry forward"Drop-in replacement" is a design requirement with
a number attached, not a marketing phrase: 500 use cases that nobody had to rewrite.
https://www.linkedin.com/blog/engineering/open-source/open-sourcing-venice-linkedin-s-derived-data-platform
Eng blog
SoftwareMill2025
Northguard and Xinfra: what LinkedIn changed when Kafka was no longer enough
Reports that successive segments in the same range can hold different replica sets, which
LinkedIn calls log striping, that a segment seals at 1 GB or one hour or on losing a
replica, and that Xinfra supports dual writes with epoch-based ordering during
migration.
Carry forwardChanging the unit of placement from partition to
sealed segment is what makes the new system balanceable at the old one's topic count.
https://softwaremill.com/linkedin-northguard-xinfra-kafka/
Talk
QCon London2024-04
gRPC migration automation at LinkedIn (Karthik Ramgopal, Min Chen)
50,000 production endpoints moved from Rest.li to gRPC across roughly 8,000 services and
about 100M lines of code, with a planned two-to-three-year manual migration reported as
delivered in two to three quarters once the edits were automated, using an intermediate
bridged mode so both protocols ran side by side.
Carry forwardTwo artefacts decide whether a framework migration
is feasible: the bridged mode, and the tool that rewrites the call sites.
https://qconlondon.com/presentation/apr2024/grpc-migration-automation-linkedin
Talk
LinkedIn, Kafka Summit2016
Kafkaesque days at LinkedIn in 2015 (slides)
The deck's own taxonomy of that year's incidents: offset rewinds, data loss, cluster
unavailability, "(in)compatibility", and blackout.
Carry forwardCompatibility was already a named incident class
at LinkedIn in 2015, a decade before the protocol collision showed up in a pull
request.
https://www.slideshare.net/jjkoshy/kafkaesque-days-at-linked-in-in-2015
Paper
Google2013
Online, asynchronous schema change in F1 (VLDB)
Replaces corruption-causing schema changes with a sequence of changes "guaranteed to
avoid corrupting the database so long as all servers are no more than one schema version
behind at any time".
Carry forwardThe formal statement of the bridge constraint. It
is also the test for your own plan: can any two participants be two steps apart?
http://www.vldb.org/pvldb/vol6/p1045-rae.pdf
Paper
LinkedIn2013
On brewing fresh Espresso: LinkedIn's distributed data serving platform (SIGMOD)
Describes the store as offering a hierarchical document model, transactional modification
of related documents, real-time secondary indexing, on-the-fly schema evolution and "a
timeline-consistent change capture stream".
Carry forwardThe one LinkedIn store nobody had to replace is
the one that shipped a change stream as a first-class feature, which let everything
downstream of it be replaced instead.
https://dl.acm.org/doi/10.1145/2463676.2465298
Paper
LinkedIn2011
Kafka: a distributed messaging system for log processing (NetDB)
Kreps, Narkhede and Rao describe a system "developed for collecting and delivering high
volumes of log data with low latency", suitable for offline and online consumption.
Carry forwardThe scope in the original paper is log delivery.
Fourteen years of scope growth is a large part of why the replacement was needed.
https://www.odbms.org/2011/01/kafka-a-distributed-messaging-system-for-log-processing/
Vendor
Confluent2020
Kafka needs no keeper: removing the ZooKeeper dependency
Explains the bridge release as able to coexist with versions on both sides of the change,
enabling a zero-downtime upgrade, with all brokers except the controller treating
ZooKeeper as read-only.
Carry forwardCoexistence, not conversion, is the property that
makes the upgrade schedulable.
https://www.confluent.io/blog/removing-zookeeper-dependency-in-kafka/
Vendor
AWS2025
In-place ZooKeeper-to-KRaft upgrades for Amazon MSK
A managed in-place metadata migration, with the caveat that attempting it on a cluster
with altered advertised.listeners fails the upgrade.
Carry forwardEven a vendor-run bridge has preconditions on your
configuration. Enumerate yours before you schedule the window.
https://aws.amazon.com/blogs/big-data/announcing-in-place-zookeeper-to-kraft-cluster-upgrades-for-amazon-msk/
Vendor
Apache Kafka release coverage2025-03
Kafka 4.0: the ZooKeeper path closes
4.0 is the first Apache Kafka release that runs entirely without ZooKeeper, and brokers
upgrade directly to it only from KRaft clusters on 3.3.x through 3.9.x, so a ZooKeeper
cluster must migrate first through the 3.9 bridge release.
Carry forwardThis is the clock on LinkedIn's bridge programme:
3.9 is the last door, and the fork is on 3.0.
https://softwaremill.com/apache-kafka-4-0-0-released-kraft-queues-better-rebalance-performance/