Leaving your own infrastructure  / field guide
Practitioner field guide · 10 October 2026

Swapping the system behind the name: ten years of LinkedIn

LinkedIn has replaced its log store, its RPC framework, its service discovery control plane, its derived-data store and its stream processor inside one decade, and abandoned the one migration it attempted without a compatibility layer. Reconstructed from the branches and pull requests it still maintains in public, this guide gives you the shape of the artefact that makes an exit possible, and the conditions under which building it is the wrong call.

41 sourced claims 12 organisations 4 incidents Evidence through September 2026 Read: 20 min
01

The territory

The problem, stated without naming a technology: how do you get out of a system that thousands of callers address directly, when you built it yourself and nobody else in the industry runs it?

32T
records a day through Kafka at LinkedIn before the replacement was announced
3.0-li
newest release branch of LinkedIn's public Kafka fork, while Apache shipped 4.0
50,000
production endpoints moved off the RPC framework LinkedIn wrote and published
~500
Voldemort use cases moved to Venice on a write path that was API-compatible

LinkedIn is an unusually clean specimen because it is on both sides of the problem. It wrote Kafka, published it, and watched the industry standardise on it; Jay Kreps, Neha Narkhede and Jun Rao described it in 2011 as "a distributed messaging system that we developed for collecting and delivering high volumes of log data with low latency". Fourteen years later, in June 2025, LinkedIn announced Northguard, a new log store, and Xinfra, a virtualised pub/sub layer whose job is to let Northguard and Kafka serve the same applications at the same time. In between it did the same thing to four other load-bearing systems. The question this guide answers is not why a company outgrows its own software, which is uninteresting, but what it has to have built in advance to be able to leave.

The answer, across every case in the record, is a layer between the caller and the system that owns the name: the topic, the service, the dataset, the pipeline. Sources call it a wrapper library, a virtualised pub/sub layer, a bridge mode, a drop-in replacement, an intermediate representation and a shadow read. I am going to call it a seam, and say plainly that I am naming it, because the absence of a shared name is part of why the pattern keeps being rebuilt from scratch. A seam has four properties that matter and that this guide traces: it implements the incumbent's interface exactly, it is installed inside the incumbent rather than beside the replacement, it carries a switch that defaults to the old behaviour, and somebody has written down the condition under which it gets deleted.

The surprise is where the seams were built. Xinfra reads, from the outside, like a layer that arrived with Northguard in 2025. It did not. On 20 October 2023, a pull request against LinkedIn's public Kafka fork added three remote procedure calls for "xinfra topic operations", the first of which "adds a znode under /xinfraTopics/topicName with the namespace name as its value". The abstraction that would later let topics move off Kafka was first built as metadata inside Kafka, twenty months before the system it was abstracting away was publicly named. That ordering is the whole finding: the seam goes into the thing you are leaving, not into the thing you are going to.

What this guide covers, and what it does not

It covers the mechanics of leaving infrastructure you built and operate yourself: where the indirection lives, what it costs, and how published incidents say it fails. It is built mainly from one company's public record, cross-checked against Apache Kafka's own removal plan for ZooKeeper, gRPC's decision record for adopting xDS, and three unrelated postmortems. It does not cover Northguard's internals beyond what the announcement describes, does not compare Northguard to Kafka as products, does not touch the economics of running your own data centres, and says nothing about migrating off a vendor you do not control, which is a different problem with a different remedy.

Figure 1 · One decade, five exits and one abandonment

Attempted without one

Replaced through a seam

Voldemort RO
to Venice
seam: Full Push API

Rest.li to gRPC
seam: bridged mode

Samza only to
Samza + Spark
seam: Beam pipeline

Kafka to Northguard
seam: Xinfra namespace

ZooKeeper D2 to xDS
seam: dual read

All workloads
to Azure

Completed

Paused 2022,
reported dead 2023

Attempted without one

Replaced through a seam

Voldemort RO
to Venice
seam: Full Push API

Rest.li to gRPC
seam: bridged mode

Samza only to
Samza + Spark
seam: Beam pipeline

Kafka to Northguard
seam: Xinfra namespace

ZooKeeper D2 to xDS
seam: dual read

All workloads
to Azure

Completed

Paused 2022,
reported dead 2023

Every replacement that completed went through a named indirection layer; the one attempt with no layer, reported as a lift and shift to Azure, is the one that stopped. Compiled from linkedin/kafka PR 479, InfoQ on LinkedIn's service discovery rebuild and CNBC, December 2023.
Diagram source
02

How it is actually built

Five independent seams in one company's record, plus Apache's own and gRPC's, reduce to the same five parts. Where they differ is which part carries traffic.

Figure 2 · The reference shape of a seam

resolve name

traffic while gate is off

traffic once gate is on

Application code
unchanged, compiled
against the old interface

Seam client library
implements the old interface
exactly

Control plane
name to system mapping
+ capability gate

Incumbent system

Replacement system

Comparator
reads both, reports diffs

resolve name

traffic while gate is off

traffic once gate is on

Application code
unchanged, compiled
against the old interface

Seam client library
implements the old interface
exactly

Control plane
name to system mapping
+ capability gate

Incumbent system

Replacement system

Comparator
reads both, reports diffs

The caller keeps the interface it already compiled against; everything that knows there are two systems lives in the resolver and the control plane. Reconstructed from li-apache-kafka-clients, the xinfra topic RPCs, Venice's client tiers and the dual-read discovery migration. Dashed boxes are the parts that exist only during the migration.
Diagram source

Start with the oldest artefact in the record, because it is the clearest. LinkedIn's li-apache-kafka-clients is, in its own words, "a wrapper Kafka clients library built on top of vanilla Apache Kafka clients", "fully compatible with Apache Kafka vanilla clients", whose producer splits oversized messages into segments and whose consumer reassembles them "without needing an external storage dependency". The README then does something most internal libraries never do: it tells you not to use it. "If you don't need these extra functions, the vanilla Kafka Java clients should be preferred." That sentence is the design. LinkedIn-only behaviour lives in a layer that implements the standard interface, so a caller that stops needing the behaviour drops the layer rather than rewriting against a different API. Xinfra is the same move generalised: one pub/sub client interface that resolves a topic name to whichever log system currently holds it, so an application cannot tell whether it is talking to Kafka or to Northguard.

The second part is the control plane that holds the mapping, and here the record is unusually specific about an implementation detail that is easy to get wrong. The 2023 pull request that restructured LinkedIn's federated topic znodes moved them from /kafka-tracking/federatedTopics/PageViewEvent(data = tracking) to /kafka-tracking/federatedTopics/tracking/PageViewEvent, putting the namespace in the path rather than in the node's payload. The author's stated reasons are that it "saves ZooKeeper space and simplifies logic on the Mario side", Mario being LinkedIn's Kafka control plane. The deeper consequence is that a name becomes a hierarchical, listable, watchable thing rather than an opaque blob, which is what lets a resolver enumerate what it owns. If your seam's mapping is a string inside a value, you will rewrite it.

The third part is the gate, and the public record gives us its exact form. When LinkedIn prepared to move its Kafka fork from the 3.0-li line to Apache 3.9, it added li.protocol.bridge.mode.enable, described in the pull request as a cluster-level dynamic configuration, off by default, and ignored entirely under KRaft. Three properties, each load-bearing. Cluster-level, so the decision is made once in the place that can see all participants. Dynamic, so turning it off does not require the rolling restart that turning it on did. Off by default, so a broker that comes up without having been told anything behaves like the old fleet. The accompanying open pull requests are blunt about why this matters: one is titled "close default-off gaps in operational compatibility", another "Fail closed on unresolved upgrade admission decisions".

The fourth part is the comparator, and LinkedIn's most recent published migration is built around it. Its service discovery control plane ran on ZooKeeper with a house registration format called D2 for more than ten years. The rebuild moves writes through Kafka and reads through an observer that pushes updates over xDS, and the migration ran as a dual read: clients read from both ZooKeeper, which stayed the source of truth, and the new system in shadow mode, with metrics comparing the two silently in the background. InfoQ's February 2026 account reports that this surfaced reconnection storms, data inconsistencies, subscription bugs and TLS misconfigurations before any of them could route real traffic. One observer instance is reported to hold 40,000 client streams and process 10,000 updates per second.

The fifth part is the one almost nobody builds, and it is a sentence rather than a component. The federated-topics pull request says the change "can be deprecated once all Kafka clients move to xinfra clients and all topic create/update/delete operations go through xmd". That is an exit criterion, written into the artefact, naming the observable condition under which the compatibility code is deleted. Without it, a seam becomes permanent, and a permanent seam is a second system to operate rather than a path out of the first.

Figure 5 · The states a seam passes through, and the one most never reach

mismatch classes named and closed

gate flipped per cluster

config change, no restart

no caller on the old path

nobody wrote the criterion

Seam installed in the incumbent

Shadow read, old system authoritative

Dual write, ordering defined

Gate on, new system authoritative

Gate off again

Seam deleted, exit criterion met

Seam becomes a second system to operate

mismatch classes named and closed

gate flipped per cluster

config change, no restart

no caller on the old path

nobody wrote the criterion

Seam installed in the incumbent

Shadow read, old system authoritative

Dual write, ordering defined

Gate on, new system authoritative

Gate off again

Seam deleted, exit criterion met

Seam becomes a second system to operate

Every state in this lifecycle is attested somewhere in the record; only the last one is attested as a sentence of intent rather than as shipped code, which is the honest limit of the evidence. Compiled from PR 494's exit criterion, the bridge-mode gate and the dual-read discovery migration.
Diagram source

Where the seam diverges: traffic or metadata

Xinfra and the Kafka client wrapper sit on the data path and pay latency for it. Coral, LinkedIn's SQL intermediate representation, and the Beam pipeline definition sit on the authoring path and cost nothing at runtime because translation happens before execution. Prefer the authoring path when you can: a seam that only compiles is a seam that cannot page you.

Runs this way at: Coral, Beam at LinkedIn

The interface is the unit of migration, not the system

Venice ships three clients against one API: a thin client under 10ms, a fast client under 2ms, and Da Vinci, "Stateful local cache, 0 network hops, < 1ms latency". A caller changes its cost and failure profile by changing a dependency, not a call site. The same logic retired Voldemort: roughly 500 production use cases moved on Venice's Full Push, described as a drop-in replacement.

Runs this way at: Venice, Voldemort, retired 2018

Upstream builds the same thing and calls it a bridge

Apache Kafka's KIP-500 plans ZooKeeper's removal through a "bridge release" in which the dependency is isolated rather than removed, with access confined to the controller instead of brokers, clients and tools. Confluent's account adds the point: the bridge release can coexist with versions on both sides of the change, which is what makes a zero-downtime upgrade possible at all.

Runs this way at: KIP-500, Confluent

03

The decisions that matter

Four forks where LinkedIn's record and the open-source record disagree usefully, each with the condition that flips it.

Decision: do you extend the incumbent's wire protocol, or refuse to?

Chosen
  • LinkedIn extended it. 3.0-li changed the wire formats of LeaderAndIsr v3, UpdateMetadata v6 and StopReplica v2 by inserting a MaxBrokerEpoch field.
  • It shipped years earlier than an upstream proposal would have.
Rejected
  • Waiting for Apache, or carrying the extension only in private request numbers.
  • The bill arrived in 2026: "Apache later assigned different meanings to those same version numbers", so a 3.9 controller must avoid the colliding versions while any 3.0-li broker is alive.
Flips when
  • Never bump a version number in a space somebody else allocates. Private namespaces survive; private numbers collide. LinkedIn's own private API identifiers in the 1000 range came through the same upgrade intact.

Decision: publish your own framework, or adopt the emerging standard?

Chosen
  • LinkedIn published Rest.li and D2, and in 2023 started leaving both.
  • The README is now explicit: "There is no active development at LinkedIn on new features for Rest.li", and "gRPC is not a drop-in replacement".
Rejected
  • Converging early. gRPC's own decision record for adopting xDS gives the reason LinkedIn eventually hit: the API "is evolving into a standard" that other data plane software will use.
  • LinkedIn's stated cost: Rest.li was open-sourced "but this led to only limited adoption outside LinkedIn", and its non-Java support was "patchy or non-existent".
Flips when
  • Publish when no standard exists yet, which was true of Kafka in 2011 and of nothing in this decade. Once a standard has a second independent data plane implementing it, a house format stops being an asset and becomes the reason you cannot adopt the tooling everyone else tests.

Decision: migrate by lift and shift, or by seam?

Chosen
  • In 2019 LinkedIn chose lift and shift to Azure: "a multi-year migration of all LinkedIn workloads to the public cloud".
  • In June 2022 its CTO told R&D staff the company would instead "focus our efforts on scaling and innovating our on-prem infrastructure".
Rejected
  • Refactoring first, then moving. Reporting attributes the trouble to moving existing tools without the adjustments needed to run them well.
  • No engineer-written account of the failure exists, so the mechanism is reported, not documented.
Flips when
  • Lift and shift is defensible only where the unit you are moving already has an interface somebody else implements. LinkedIn's stack was built on Kafka it forked, a discovery format it invented and an RPC framework it published, which means almost nothing in it had a second implementation to land on.

Figure 3 · Which seam, and where it lives

yes

no

yes

no

yes

no

Can you redeploy
every caller?

Does the old interface
have a second
implementation?

Is the change visible
at authoring time?

Seam in the client library
e.g. Xinfra, Da Vinci

Seam as a wire-compatible
endpoint or proxy
e.g. Kafka bridge mode

Seam in the translation layer
e.g. Coral IR, Beam pipeline

Stop: you are rewriting,
not migrating.
Budget for it honestly

yes

no

yes

no

yes

no

Can you redeploy
every caller?

Does the old interface
have a second
implementation?

Is the change visible
at authoring time?

Seam in the client library
e.g. Xinfra, Da Vinci

Seam as a wire-compatible
endpoint or proxy
e.g. Kafka bridge mode

Seam in the translation layer
e.g. Coral IR, Beam pipeline

Stop: you are rewriting,
not migrating.
Budget for it honestly

The first question is not which system wins, it is whether you can redeploy the callers. Terminal nodes name the artefact to build, drawn from the five LinkedIn migrations in this guide.
Diagram source
DecisionChosenRejectedBecauseEvidence
Where the indirection livesClient library implementing the old interfaceBroker-side or proxy shimThe client is the only participant that can hold both mappings and be redeployed independentlyli-apache-kafka-clients
How the namespace is storedNamespace in the pathNamespace in the node valueSaves coordination-store space and lets the control plane list and watch what it ownsPR 494
Default of the compatibility gateOff, dynamic, cluster-levelOn by default, or staticA broker that was told nothing must behave like the old fleet, and reversing must not need a restartPR 542
Unit of placement in the replacementSegment, striped across replica setsPartition, as in KafkaPartition-level placement is what stopped balancing at 400,000 topics; a sealed 1 GB segment can be placed anywhereSoftwareMill, 2025
How the migration is validatedDual read with silent comparisonCut over and watch error ratesReconnection storms and data inconsistencies showed up in shadow mode before they could route trafficInfoQ, Feb 2026
How the mechanical edits get doneAutomated code change at scaleTeam-by-team manual migrationA migration planned at two to three years was reported delivered in two to three quartersQCon, 2024
04

What broke in production

Four published incidents, three failure classes. None of them is a failure of the replacement system; all of them are failures of the transition between two versions.

Class A: no reverse path

Postmortem

Reddit, Pi Day 2023: the upgrade had no downgrade

AssumptionThat a Kubernetes minor-version upgrade which had worked on other clusters was reversible if it did not.
What happenedUpgrading one cluster from 1.23 to 1.24 changed a control-plane node label that a Calico route-reflector selector still pointed at, and routing collapsed.
Blast radiusRoughly 314 minutes of site downtime on 14 March 2023. The actual cause was not found until hours after service was restored.
FixRestoring from backup, because "there was no supported way to downgrade Kubernetes", with a restore procedure the postmortem describes as outdated.
Design ruleA seam you can only cross in one direction is not a seam, it is a cliff. Test the reverse before the forward, and date-stamp the restore runbook the same day you upgrade.
Postmortem

LinkedIn, 2015: both recovery choices lose

AssumptionThat a consumer whose offset is unusable can be recovered by resetting it.
What happenedOffsets rewound. Resetting to the earliest offset reprocesses everything; resetting to the latest loses whatever arrived between the reset and the next fetch. Neither option is safe, and the decision is made under pressure.
Blast radiusNot quantified in the post. LinkedIn's own 2015 taxonomy lists offset rewinds, data loss, cluster unavailability, "(in)compatibility" and blackout as the year's classes.
FixChanges in offset management, log compaction and monitoring rather than in the application layer.
Design rulePosition is state, and a seam that proxies a position has to own it. Xinfra's reported epoch-based ordering across dual writes is the same lesson, applied ten years later by the same company.

Class B: the new default was the hazard

Postmortem

Roblox, October 2021: 73 hours from a minor-version feature

AssumptionThat a new Consul feature, enabled deliberately on one traffic-routing service, had a blast radius limited to that service.
What happenedThe streaming feature introduced in the 1.9 to 1.10 upgrade, under heavy read and write load, combined with a BoltDB performance problem to degrade the cluster the whole platform depended on.
Blast radius73 hours. Diagnosis was slowed because "critical monitoring systems that would have provided better visibility relied on affected systems, such as Consul".
FixDisabling streaming by configuration change, then switching BoltDB's freelist implementation and recovering in stages.
Design ruleThe gate saved them. A new behaviour reachable only through a runtime switch is recoverable; the same behaviour compiled in is a redeploy under load. This is why LinkedIn's bridge work carries pull requests titled "close default-off gaps".
Source code

LinkedIn, 2026: the fork's version numbers collided with upstream's

AssumptionThat adding a field to an internal request and bumping its version was a local change, because the fork was described in 2019 as "not a fork" and kept "as close as possible to upstream".
What happenedApache assigned different meanings to LeaderAndIsr v3, UpdateMetadata v6 and StopReplica v2. A 3.9 controller and a 3.0-li broker in one cluster can now agree on a version number and disagree on its contents.
Blast radiusNo published incident. The visible cost is a programme of work: 51 open pull requests in September 2026, more than twenty of them bridge compatibility gates on two parallel base branches.
FixA dynamic, off-by-default cluster config that makes the new controller send the newest versions shared with the old brokers, plus the same change prepared independently on both base branches.
Design ruleTreat an upstream version number as somebody else's primary key. Extensions go in a namespace you allocate, or they go upstream, or you accept that the upgrade five years out is a project rather than a release.

Figure 4 · How a version number becomes an outage

3.9-li broker3.0-li broker3.9 controller3.9-li broker3.0-li broker3.9 controllerMixed-version rollout, bridge mode OFFParses v3 as the LinkedInformatwith MaxBrokerEpochBridge mode ONLeaderAndIsr v3 (Apache meaning)1ok2LeaderAndIsr v3 (Apachemeaning)3accepted, wrong leadershipstate4LeaderAndIsr v2 (newestshared version)5ok6
3.9-li broker3.0-li broker3.9 controller3.9-li broker3.0-li broker3.9 controllerMixed-version rollout, bridge mode OFFParses v3 as the LinkedInformatwith MaxBrokerEpochBridge mode ONLeaderAndIsr v3 (Apache meaning)1ok2LeaderAndIsr v3 (Apachemeaning)3accepted, wrong leadershipstate4LeaderAndIsr v2 (newestshared version)5ok6
The failure needs no bug: both participants are correct about the version they speak. Reconstructed from the description of linkedin/kafka PR 542.
Diagram source
What the record does not contain

No public postmortem in this corpus describes a seam that worked and was then deleted on schedule. The evidence that these layers come out is a sentence in a pull request description, not an incident review or a deprecation notice. That asymmetry should make you suspicious in a specific direction: the risk with a seam is not that it fails, it is that it becomes permanent and nobody writes a postmortem about a layer that simply stayed.

05

Numbers you can plan against

Growth on one side, migration effort on the other. The gap between the two is the argument for building the seam early.

MetricValueAtContextAs ofSource
Messages per day through Kafka7 trillionLinkedInOver 100 clusters, 4,000+ brokers, 100,000 topics, 7M partitions2019LinkedIn blog
Records per day through Kafka32 trillionLinkedIn17 PB/day, 400,000 topics, 150 clusters2025InfoQ
Nodes behind those clusters10,000+LinkedInReported alongside the 150-cluster figure2025BigDATAwire
Growth in daily volume, 2019 to 20254.6x (derived)LinkedIn32 trillion divided by 7 trillion. Units differ, messages against records, so treat as a magnitude2026arithmetic on the two rows above
Segment seal threshold in Northguard1 GBLinkedInOr one hour active, or on losing a replica2025SoftwareMill
Endpoints migrated off Rest.li50,000LinkedInAcross roughly 8,000 services and about 100M lines of code2024QCon London
Migration elapsed time, planned against actual2-3 years to 2-3 quartersLinkedInManual plan against automated delivery, as presented by the engineers who ran it2024InfoQ Q&A
Latency reduction from Protobufup to 60%LinkedInLargest gains on services with very large and complex payloads; validated by synthetic benchmarks and side-by-side production ramps2023InfoQ Q&A
Backfill processing time, before and after Beam7.5 hours to 25 minutesLinkedInOne pipeline definition, Samza for streaming and Spark for backfill; roughly half the memory and CPU2023LinkedIn blog
Streams per discovery observer instance40,000LinkedInPlus 10,000 updates per second, after the xDS rebuild2026InfoQ
Use cases moved off Voldemort read-only~500LinkedInFully migrated by 2018 on Venice's Full Push, described as a drop-in replacement2022LinkedIn blog
Open pull requests on the public Kafka fork51 of 595LinkedInSeptember 2026, overwhelmingly bridge compatibility work on two parallel base branches2026linkedin/kafka
Closed-unmerged pull requests on the same fork93LinkedInIncludes an entire kernel-TLS workstream abandoned in 20232026linkedin/kafka
Years the house discovery format ran10+LinkedInZooKeeper plus the D2 registration format, until the xDS rebuild2026InfoQ
Read these carefully

Measured and primary: the 2019 Kafka fleet figures, the Beam backfill times and the Voldemort migration count, all published by LinkedIn about its own systems. Reported second-hand: every 2025 and 2026 figure in this table, because LinkedIn's engineering blog could not be fetched directly from the container this guide was built in; the primary URLs are cited and the numbers come from InfoQ, BigDATAwire, SoftwareMill and SiliconANGLE reporting on those posts. Vendor: Apache Beam's case-study claim of a 2x cost-to-serve improvement, which no engineer-written source corroborates. Derived: the 4.6x growth row, shown with its arithmetic, and weak because the two figures count different things. Unknown, and worth knowing that it is unknown: what any of this costs. There is no public figure for what LinkedIn's own data centres cost per unit of work, for what the Azure attempt cost, or for the engineering cost of the bridge programme. Anyone who quotes you a return-on-investment number for repatriation or for a seam is modelling, not citing.

06

The evidence wall

Every source behind this page, graded. GitHub rows were fetched and read in full; the rest were retrieved through a server-side search tool, which the ledger records per claim. Filter by kind.

Source code LinkedIn2026-09

linkedin/kafka, pull request 542: the 3.9 side of the 3.0-li protocol bridge

States that 3.0-li changed LeaderAndIsr v3, UpdateMetadata v6 and StopReplica v2 by inserting MaxBrokerEpoch, that Apache later gave those same version numbers different meanings, and that the remedy is a dynamic cluster config which makes the controller send the newest versions shared with the old brokers. Closed unmerged and split into a reviewable stack.

Carry forwardA private extension to someone else's version number is a dated liability, and the interest is paid at upgrade time.
https://github.com/linkedin/kafka/pull/542
Source code LinkedIn2023-10

linkedin/kafka, pull request 479: RPCs for xinfra topic creation and deletion

Adds create, delete and list operations that write a znode under /xinfraTopics/topicName carrying the namespace as its value, reusing the existing topic change listener in the Kafka server. Dated twenty months before Northguard and Xinfra were publicly announced.

Carry forwardThe abstraction that lets you leave gets installed inside the incumbent, early, as metadata.
https://github.com/linkedin/kafka/pull/479
Source code LinkedIn2023-11

linkedin/kafka, pull request 494: change federated topic znode structure

Moves the namespace from a node's value into its path, for space and for control-plane simplicity, and states the exit criterion in the description: the change can be deprecated once all clients move to xinfra clients and all topic lifecycle operations go through xmd.

Carry forwardWrite the deletion condition into the compatibility change itself, or the change outlives the reason for it.
https://github.com/linkedin/kafka/pull/494
Source code LinkedIn2026-09

linkedin/kafka branches and open pull requests

The only branch ending in -li is 3.0-li; the nineteen most recently updated branches are all 3.9-li-bridge/... or 3.0-li-bridge/.... Open titles include closing default-off gaps in operational compatibility, failing closed on unresolved upgrade admission decisions, and restoring LinkedIn ZooKeeper operational behaviour.

Carry forwardA fork's real cost shows up as a branch programme five years later, not as a merge conflict today.
https://github.com/linkedin/kafka/branches/all
Source code LinkedIncurrent

linkedin/li-apache-kafka-clients

A wrapper library over the vanilla Kafka clients, fully compatible with their interfaces, adding large-message segmentation and pluggable auditing. Its README tells readers who do not need those functions to prefer the vanilla clients.

Carry forwardHouse behaviour belongs in a removable layer that implements the standard interface, and the README should say when to remove it.
https://github.com/linkedin/li-apache-kafka-clients
Source code LinkedIncurrent

linkedin/rest.li and its closed-unmerged pull requests

No active development on new features, with gRPC named as the direction and explicitly "not a drop-in replacement". The frozen repository nonetheless carries 2025 and 2026 pull requests about gRPC name resolution and xDS client metrics, which is what a shared discovery plane looks like from the inside of the old framework.

Carry forwardA framework you are leaving still needs maintenance on the seam, long after you stop adding features.
https://github.com/linkedin/rest.li
Source code LinkedIncurrent

linkedin/venice and linkedin/coral

Venice offers three clients against one API at three cost points, ending with Da Vinci: stateful local cache, zero network hops, under one millisecond. Coral defines a SQL intermediate representation independent of any dialect, translating HiveQL and Spark SQL in, and Hive, Spark and Trino out.

Carry forwardTwo seam shapes: one on the data path that trades latency, one on the authoring path that costs nothing at runtime.
https://github.com/linkedin/venice
Source code Voldemort projectcurrent

voldemort/voldemort README

"Voldemort is no longer under development." LinkedIn was its primary maintainer and user and "stopped all production usage in 2018"; most read-only use cases and some read-write ones migrated to Venice.

Carry forwardThe predecessor's repository is where you find the honest end date, and it is usually years before the archive notice.
https://github.com/voldemort/voldemort
Decision record Apache Kafka2019-2021

KIP-500: replace ZooKeeper with a self-managed metadata quorum

The compatibility plan calls for a "bridge release" in which the ZooKeeper dependency is well isolated: the release does not remove ZooKeeper but eliminates most of the touch points, confining access to the controller rather than brokers, clients and tools.

Carry forwardThe canonical name for the artefact, from the project that had to remove its own foundation. Isolate first, remove second.
https://cwiki.apache.org/confluence/display/KAFKA/KIP-500
Decision record gRPC2020-03

gRFC A27: xDS-based global load balancing

gRPC moves off its own grpclb protocol to converge with the industry trend, on the stated grounds that the xDS API is evolving into a standard that other data plane software will use. The rationale concedes the design would have been much simpler had load balancing been the only thing configured through it.

Carry forwardThe written form of the decision LinkedIn's D2 format eventually forced: adopt the standard before your format becomes the obstacle.
https://github.com/grpc/proposal/blob/master/A27-xds-global-load-balancing.md
Postmortem Reddit2023-03

You broke Reddit: the Pi-Day outage

About 314 minutes of downtime after a Kubernetes 1.23 to 1.24 upgrade on a cluster built with kubeadm rather than a standard template. There was no supported way to downgrade, so recovery meant restoring from backup with an outdated procedure, and the actual cause was found hours later.

Carry forwardRehearse the reverse direction. An upgrade without a tested downgrade converts a config mistake into a restore.
https://postmortem.io/incidents/reddit--2023-03-21--pi-day-outage/
Postmortem Roblox2022-01

Roblox return to service, 28 to 31 October 2021

73 hours, attributed to a streaming feature taken on in a Consul 1.9 to 1.10 upgrade, under heavy read and write load, plus a BoltDB performance problem. Recovery came through a configuration change disabling the feature. Monitoring that would have helped depended on the system that was down.

Carry forwardThe runtime switch is the recovery path. Also: never let the observability of a migration depend on the thing being migrated.
https://blog.roblox.com/2022/01/roblox-return-to-service-10-28-10-31-2021/
Postmortem LinkedIn2016-05

Kafkaesque days at LinkedIn

Joel Koshy's account of the 2015 incidents: offset rewinds where resetting to the earliest offset duplicates and resetting to the latest loses, with fixes landing in offset management, log compaction and monitoring rather than in applications.

Carry forwardConsumer position is state that a migration must own explicitly, which is why Xinfra's dual writes are reported as epoch-ordered.
https://engineering.linkedin.com/blog/2016/05/kafkaesque-days-at-linkedin--part-1
Case study InfoQ2025-06

LinkedIn announces Northguard and Xinfra: scaling beyond Kafka

The reported scale at which the incumbent stopped being operable: 32 trillion records and 17 PB a day, 400,000 topics, 150 clusters. Xinfra abstracts the underlying log systems so applications use one pub/sub interface whether data sits in Kafka or Northguard.

Carry forwardThe replacement shipped with the abstraction, not after it. Coexistence was a launch requirement.
https://www.infoq.com/news/2025/06/linkedin-northguard-xinfra/
Case study InfoQ2026-02

How LinkedIn rebuilt service discovery

More than ten years on ZooKeeper with the house D2 format, whose custom schemas are reported as incompatible with modern data planes such as gRPC and Envoy. The rebuild moves writes through Kafka and reads through an xDS-pushing observer, migrated as a dual read with ZooKeeper as source of truth and silent comparison in the background.

Carry forwardShadow mode is how you find reconnection storms and inconsistencies while they are still free.
https://www.infoq.com/news/2026/02/linkedin-service-discovery/
Case study CNBC2023-12

LinkedIn shelved its plan to migrate to Microsoft Azure

A June 2022 memo from LinkedIn's CTO said the company would keep using some Azure services and focus its efforts on scaling and innovating its on-premises infrastructure. Executives framed it as on hold rather than cancelled; reporting attributes the difficulty to lifting and shifting existing tools rather than refactoring them. Tech Monitor reports a spokesperson saying the move will not go ahead.

Carry forwardThe only migration here with no compatibility layer is the only one that stopped, and sources still disagree on whether it is paused or dead.
https://www.cnbc.com/2023/12/14/linkedin-shelved-plan-to-migrate-to-microsoft-azure-cloud.html
Eng blog LinkedIn2019

How LinkedIn customizes Apache Kafka for 7 trillion messages per day

LinkedIn's 2019 position: this is "not a fork", each release branch comes off the matching Apache branch with an -li suffix, and the aim is to stay as close as possible to upstream. Over 100 clusters, more than 4,000 brokers, 100,000 topics, 7 million partitions.

Carry forwardRead this next to the 2026 branch list. Intending to stay close to upstream is not a mechanism for staying close to upstream.
https://www.linkedin.com/blog/engineering/open-source/apache-kafka-trillion-messages
Eng blog LinkedIn2023-04

Unified streaming and batch pipelines, reducing processing time by 94%

One pipeline definition in Apache Beam running on two engines, Samza for streaming and Spark for backfill, cutting a backfill from 7.5 hours to 25 minutes on roughly half the memory and CPU.

Carry forwardA seam on the authoring path buys you engine choice for free at runtime. This is the cheapest version of the pattern.
https://engineering.linkedin.com/blog/2023/unified-streaming-and-batch-pipelines-at-linkedin--reducing-proc
Eng blog LinkedIn2022-09

Open sourcing Venice, LinkedIn's derived data platform

Roughly 500 Voldemort read-only use cases were fully migrated to Venice by 2018 using its Full Push capability as a drop-in replacement, and in 2020 a second client library, Da Vinci, served the same APIs from local state instead of remote queries.

Carry forward"Drop-in replacement" is a design requirement with a number attached, not a marketing phrase: 500 use cases that nobody had to rewrite.
https://www.linkedin.com/blog/engineering/open-source/open-sourcing-venice-linkedin-s-derived-data-platform
Eng blog SoftwareMill2025

Northguard and Xinfra: what LinkedIn changed when Kafka was no longer enough

Reports that successive segments in the same range can hold different replica sets, which LinkedIn calls log striping, that a segment seals at 1 GB or one hour or on losing a replica, and that Xinfra supports dual writes with epoch-based ordering during migration.

Carry forwardChanging the unit of placement from partition to sealed segment is what makes the new system balanceable at the old one's topic count.
https://softwaremill.com/linkedin-northguard-xinfra-kafka/
Talk QCon London2024-04

gRPC migration automation at LinkedIn (Karthik Ramgopal, Min Chen)

50,000 production endpoints moved from Rest.li to gRPC across roughly 8,000 services and about 100M lines of code, with a planned two-to-three-year manual migration reported as delivered in two to three quarters once the edits were automated, using an intermediate bridged mode so both protocols ran side by side.

Carry forwardTwo artefacts decide whether a framework migration is feasible: the bridged mode, and the tool that rewrites the call sites.
https://qconlondon.com/presentation/apr2024/grpc-migration-automation-linkedin
Talk LinkedIn, Kafka Summit2016

Kafkaesque days at LinkedIn in 2015 (slides)

The deck's own taxonomy of that year's incidents: offset rewinds, data loss, cluster unavailability, "(in)compatibility", and blackout.

Carry forwardCompatibility was already a named incident class at LinkedIn in 2015, a decade before the protocol collision showed up in a pull request.
https://www.slideshare.net/jjkoshy/kafkaesque-days-at-linked-in-in-2015
Paper Google2013

Online, asynchronous schema change in F1 (VLDB)

Replaces corruption-causing schema changes with a sequence of changes "guaranteed to avoid corrupting the database so long as all servers are no more than one schema version behind at any time".

Carry forwardThe formal statement of the bridge constraint. It is also the test for your own plan: can any two participants be two steps apart?
http://www.vldb.org/pvldb/vol6/p1045-rae.pdf
Paper LinkedIn2013

On brewing fresh Espresso: LinkedIn's distributed data serving platform (SIGMOD)

Describes the store as offering a hierarchical document model, transactional modification of related documents, real-time secondary indexing, on-the-fly schema evolution and "a timeline-consistent change capture stream".

Carry forwardThe one LinkedIn store nobody had to replace is the one that shipped a change stream as a first-class feature, which let everything downstream of it be replaced instead.
https://dl.acm.org/doi/10.1145/2463676.2465298
Paper LinkedIn2011

Kafka: a distributed messaging system for log processing (NetDB)

Kreps, Narkhede and Rao describe a system "developed for collecting and delivering high volumes of log data with low latency", suitable for offline and online consumption.

Carry forwardThe scope in the original paper is log delivery. Fourteen years of scope growth is a large part of why the replacement was needed.
https://www.odbms.org/2011/01/kafka-a-distributed-messaging-system-for-log-processing/
Vendor Confluent2020

Kafka needs no keeper: removing the ZooKeeper dependency

Explains the bridge release as able to coexist with versions on both sides of the change, enabling a zero-downtime upgrade, with all brokers except the controller treating ZooKeeper as read-only.

Carry forwardCoexistence, not conversion, is the property that makes the upgrade schedulable.
https://www.confluent.io/blog/removing-zookeeper-dependency-in-kafka/
Vendor Apache Kafka release coverage2025-03

Kafka 4.0: the ZooKeeper path closes

4.0 is the first Apache Kafka release that runs entirely without ZooKeeper, and brokers upgrade directly to it only from KRaft clusters on 3.3.x through 3.9.x, so a ZooKeeper cluster must migrate first through the 3.9 bridge release.

Carry forwardThis is the clock on LinkedIn's bridge programme: 3.9 is the last door, and the fork is on 3.0.
https://softwaremill.com/apache-kafka-4-0-0-released-kraft-queues-better-rebalance-performance/
07

Build a miniature, then productionise it

Six rungs. The line from toy to real is between four and five, where the seam starts carrying traffic it can be blamed for.

Wrap one client, change nothing else

Take any client library your services use directly. Write a wrapper that implements its interface exactly and delegates every call. Ship it to one service.

Done when: the service's behaviour and metrics are indistinguishable from before, and the diff in the service is one dependency line.  Teaches: what "implements the old interface exactly" costs, which is usually more than you expect because of the parts of the interface nobody documented.

Put the name in a control plane

Move the mapping from name to backend out of configuration and into a small service or a coordination store, with the namespace in the path rather than in the value. Have the wrapper resolve through it.

Done when: you can list every name the seam owns, and watch for changes, without parsing a value.  Teaches: why LinkedIn moved its federated topic znodes from value to path, and why a resolver that cannot enumerate is a resolver you will replace.

Add the gate, default off

Introduce one dynamic, cluster-scoped switch that selects the new path. Default it off. Make the off state the behaviour of a process that was never told the switch exists.

Done when: a fresh instance with no configuration behaves exactly like the old fleet, and flipping the switch back does not need a restart.  Teaches: the difference between the Roblox recovery, a config change, and the Reddit recovery, a restore from backup.

Run a shadow read and compare

Read from both systems, serve only the old one, and emit a comparison metric per mismatch class. Leave it running for longer than feels necessary.

Done when: you have a named list of mismatch classes with counts, and at least one of them surprised you.  Teaches: that shadow mode's output is a taxonomy of your own wrong assumptions, which is what LinkedIn's discovery migration reported finding.

Cross the line: dual write with an ordering guarantee

Write to both systems for one low-value name. Decide and document what orders the two writes, and what happens to a consumer that reads the old system after the new one has advanced.

Done when: you can state, in one sentence, what a consumer observes during the switch, and you have tested the case where the second write fails.  Teaches: why position is the hard part, and why both of the 2015 offset recovery options at LinkedIn lose something.

Write the exit criterion and wire it to a metric

State the observable condition under which the seam is deleted, in the code, next to the seam. Then build the dashboard that measures it: how many callers still use the old path, by name.

Done when: somebody who did not build the seam can read the condition and tell you how far away it is.  Teaches: the one habit from this record that costs nothing and is almost always skipped, as in "can be deprecated once all Kafka clients move to xinfra clients".

08

Keep hunting

The queries that actually produced this page. The GitHub ones did most of the work, because a company's migration is visible in its branches long before it is visible in its blog.

Find the migration in the repository

  • github.com/<org>/<repo>/branches/all
  • repo:<org>/<repo> is:pr is:closed is:unmerged bridge OR compat OR migration
  • repo:<org>/<repo> is:pr "can be deprecated once"
  • repo:<org>/<repo> "default off" OR "opt-in" OR "fail closed" in:title

Find the thing they are leaving

  • "no active development" OR "no longer under development" <framework> site:github.com
  • "not a drop-in replacement" migration <framework>
  • <company> "stopped all production usage in"
  • <company> engineering "we outgrew" OR "no longer enough" <system>

Find the abandoned migration

  • <company> migration "put on hold" OR shelved OR paused cloud memo
  • <company> "lift and shift" problems refactor internal tools
  • <company> "multi-year migration" announcement after:2018

Find the upgrade failures

  • postmortem "no supported way to downgrade"
  • outage "enabled by default" minor version upgrade root cause
  • "bridge release" OR "bridge mode" protocol version collision
09

References

The full per-claim ledger, with the quote supporting each claim and a note on how each page was retrieved, is in sources.md beside this file.

  1. LinkedIn, linkedin/kafka README GitHub, branch current. Checked 2026-10-10.
  2. LinkedIn, linkedin/kafka branch list GitHub, updated September 2026. Checked 2026-10-10.
  3. LinkedIn, PR 542: add the 3.9 side of the 3.0-li protocol bridge GitHub, opened 2026-08-27, closed 2026-09-02. Checked 2026-10-10.
  4. LinkedIn, PR 479: add rpc for xinfra topic creation and deletion GitHub, 2023-10-20. Checked 2026-10-10.
  5. LinkedIn, PR 494: change federated topic znodes structure GitHub, 2023-11-08. Checked 2026-10-10.
  6. LinkedIn, linkedin/kafka open pull requests GitHub, September 2026. Checked 2026-10-10.
  7. LinkedIn, linkedin/kafka closed-unmerged pull requests GitHub, 2022 to 2026. Checked 2026-10-10.
  8. LinkedIn, li-apache-kafka-clients GitHub, current. Checked 2026-10-10.
  9. LinkedIn, rest.li GitHub, current. Checked 2026-10-10.
  10. LinkedIn, rest.li closed-unmerged pull requests matching grpc or protobuf GitHub, 2020 to 2026. Checked 2026-10-10.
  11. LinkedIn, venice GitHub, current. Checked 2026-10-10.
  12. LinkedIn, coral GitHub, current. Checked 2026-10-10.
  13. Voldemort project, voldemort README GitHub, current. Checked 2026-10-10.
  14. gRPC, gRFC A27: xDS-based global load balancing GitHub, 2020-03-18, status Implemented. Checked 2026-10-10.
  15. Apache Kafka, KIP-500 Apache Software Foundation wiki, 2019 onward. Checked 2026-10-10.
  16. Confluent, Kafka needs no keeper Confluent blog, 2020. Checked 2026-10-10.
  17. AWS, in-place ZooKeeper-to-KRaft upgrades for Amazon MSK AWS Big Data blog. Checked 2026-10-10.
  18. SoftwareMill, Apache Kafka 4.0.0 released SoftwareMill blog, March 2025. Checked 2026-10-10.
  19. LinkedIn, How LinkedIn customizes Apache Kafka for 7 trillion messages per day LinkedIn engineering blog, 2019. Checked 2026-10-10.
  20. LinkedIn, Introducing Northguard and Xinfra LinkedIn engineering blog, 2025-06-25. Checked 2026-10-10.
  21. InfoQ, LinkedIn announces Northguard and Xinfra InfoQ, June 2025. Checked 2026-10-10.
  22. BigDATAwire, LinkedIn introduces Northguard BigDATAwire, 2025-06-25. Checked 2026-10-10.
  23. SoftwareMill, Northguard and Xinfra SoftwareMill blog, 2025. Checked 2026-10-10.
  24. SiliconANGLE, LinkedIn introduces Northguard and Xinfra SiliconANGLE, 2025-06-25. Checked 2026-10-10.
  25. InfoQ, Why LinkedIn chose gRPC and Protobuf over REST and JSON InfoQ, December 2023. Checked 2026-10-10.
  26. QCon London, gRPC migration automation at LinkedIn QCon London, April 2024. Checked 2026-10-10.
  27. InfoQ, How LinkedIn rebuilt service discovery InfoQ, February 2026. Checked 2026-10-10.
  28. LinkedIn, Unified streaming and batch pipelines LinkedIn engineering blog, April 2023. Checked 2026-10-10.
  29. Apache Beam, case study: 4 trillion events daily at LinkedIn Apache Beam site, 2023. Checked 2026-10-10.
  30. LinkedIn, Open sourcing Venice LinkedIn engineering blog, September 2022. Checked 2026-10-10.
  31. Joel Koshy, Kafkaesque days at LinkedIn, part 1 LinkedIn engineering blog, May 2016. Checked 2026-10-10.
  32. Joel Koshy, Kafkaesque days at LinkedIn in 2015 (slides) SlideShare, 2016. Checked 2026-10-10.
  33. Reddit, You broke Reddit: the Pi-Day outage postmortem.io, March 2023. Checked 2026-10-10.
  34. Roblox, Return to service 10/28 to 10/31 2021 Roblox blog, January 2022. Checked 2026-10-10.
  35. CNBC, LinkedIn shelved plan to migrate to Microsoft Azure cloud CNBC, 2023-12-14. Checked 2026-10-10.
  36. The Register, Shift to Azure abandoned The Register, 2023-12-14. Checked 2026-10-10.
  37. Tech Monitor, LinkedIn cans Azure cloud migration Tech Monitor, December 2023. Checked 2026-10-10.
  38. Kreps, Narkhede, Rao, Kafka: a distributed messaging system for log processing NetDB workshop, 2011. Checked 2026-10-10.
  39. Qiao et al., On brewing fresh Espresso SIGMOD 2013. Checked 2026-10-10.
  40. Rae et al., Online, asynchronous schema change in F1 PVLDB 6(11), 2013. Checked 2026-10-10.