Evidence ledger
One row per claim in Swapping the system behind the name: ten years of LinkedIn: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.
One row per claim. Tiers follow the skill's grading: postmortem, source, adr, casestudy,
blog, paper, talk, vendor. All links checked 2026-10-10.
Network note, and it matters for how you read this ledger. This session ran in a container
whose egress policy allowed direct fetches only to GitHub hosts. Every github.com row below was
fetched and read in full this session: READMEs, branch listings, pull request descriptions and
comments. Every other row was retrieved and quoted through the session's search tool, which
fetches pages server-side and returns the text; where that tool returned a paraphrase rather than
a verbatim sentence, the quote column says so and the claim is written as reported rather than
quoted. LinkedIn's own engineering blog is the primary source for roughly a third of this guide
and could not be fetched directly from this container, so those rows carry the secondary account
that reported the figure alongside the primary URL. No claim in the guide is cited from memory.
| # | Org | Title | Tier | Published | Checked | URL | Claim I take from it | Supporting quote or figure |
|---|---|---|---|---|---|---|---|---|
| 1 | linkedin/kafka README | source | branch current 2026 | 2026-10-10 | https://github.com/linkedin/kafka | The public fork is the running version, and it deliberately carries changes nobody upstream wants | "This is the version of Kafka running at LinkedIn"; "We run a slightly modified version of Apache Kafka trunk"; branch contents include "Patches that are on their way upstream but we have deployed internally in the meantime" and "Patches that are of no interest to upstream"; "At this moment we are not accepting external contributions directly" | |
| 2 | linkedin/kafka branch list | source | updated 2026-09 | 2026-10-10 | https://github.com/linkedin/kafka/branches/all | The newest release branch is still 3.0-li while the live work is a named bridge from 3.0-li to 3.9 |
Only branch ending in -li is 3.0-li, updated Sep 27 2026; the next 19 branches are all 3.9-li-bridge/... or 3.0-li-bridge/..., updated Sep 11-12 2026 |
|
| 3 | PR #542 "add the 3.9 side of the 3.0-li protocol bridge" (closed unmerged) | source | opened 2026-08-27, closed 2026-09-02 | 2026-10-10 | https://github.com/linkedin/kafka/pull/542 | The fork's private protocol extensions collided with Apache's later use of the same version numbers, and the remedy is a mode that speaks the newest shared version | Description: 3.0-li changed the wire formats of LeaderAndIsr v3, UpdateMetadata v6 and StopReplica v2 by inserting MaxBrokerEpoch, and "Apache later assigned different meanings to those same version numbers"; adds li.protocol.bridge.mode.enable, cluster-level, dynamic, off by default, ignored in KRaft; when on, the controller sends LeaderAndIsr v2, UpdateMetadata v5, StopReplica v1 |
|
| 4 | PR #542 closing comment | source | 2026-09-02 | 2026-10-10 | https://github.com/linkedin/kafka/pull/542 | The single bridge change was split into a reviewable stack rather than landed whole | Author's comment says the draft was split into a GitHub stack so each change can be reviewed and merged on its own, "That stack starts at #543 and ends at #555"; "PR #541 holds the matching 3.0 bridge" | |
| 5 | PR #479 "add rpc for xinfra topic creation and deletion" (closed unmerged) | source | opened and closed 2023-10-20 | 2026-10-10 | https://github.com/linkedin/kafka/pull/479 | Xinfra's namespace indirection was being plumbed into the incumbent's metadata store 20 months before Northguard was announced | "adds three RPCs for xinfra topic operations": create "adds a znode under /xinfraTopics/topicName with the namespace name as its value", delete removes it, list "returns all xinfra topics as strings combining topic and namespace names", "helping reuse the TopicChangeListener in kafka server" | |
| 6 | PR #494 "[LI-HOTFIX] Change federated topic znodes structure" (closed unmerged) | source | 2023-11-08 | 2026-10-10 | https://github.com/linkedin/kafka/pull/494 | The compatibility change names its own deletion condition in the description | Znodes move from "/kafka-tracking/federatedTopics/PageViewEvent(data = tracking)" to "/kafka-tracking/federatedTopics/tracking/PageViewEvent"; the author writes the PR "can be deprecated once all Kafka clients move to xinfra clients and all topic create/update/delete operations go through xmd" | |
| 7 | linkedin/kafka open pull requests | source | 2026-09 | 2026-10-10 | https://github.com/linkedin/kafka/pulls | The bridge work is overwhelmingly about keeping new defaults off and restoring fork behaviour, not about new features | 51 open / 544 closed. Open titles include "kafka: close default-off gaps in operational compatibility" (#594), "Fail closed on unresolved upgrade admission decisions" (#600), "kafka: restore LinkedIn ZooKeeper operational behavior" (#549), "kafka: restore LinkedIn operational metrics and watchdog" (#551), "kafka: gate 3.9 bridge-state MBeans and require the diagnostics opt-in" (#589) | |
| 8 | linkedin/kafka closed-unmerged pull requests | source | 2022-2026 | 2026-10-10 | https://github.com/linkedin/kafka/pulls?q=is%3Apr+is%3Aclosed+is%3Aunmerged | Whole workstreams in the fork were abandoned unmerged, including kernel TLS and federated-topic znode changes | 93 closed-unmerged. Includes "Use Kernel TLS in Kafka" (#420, closed 2023-06-20), "Changes for KTLS integration" (#460), "3.0 li ktls release" (#467), "[LI-HOTFIX] Fix zk listener for federated topics znode" (#496), "Add max request count limit for LiCombinedControlRequest" (#519) | |
| 9 | linkedin/li-apache-kafka-clients README | source | current | 2026-10-10 | https://github.com/linkedin/li-apache-kafka-clients | The earliest seam in the record is a client wrapper that implements the vanilla interfaces, so the LinkedIn-only behaviour can be removed without touching callers | "a wrapper Kafka clients library built on top of vanilla Apache Kafka clients", "fully compatible with Apache Kafka vanilla clients", producer splits large messages into segments and the consumer reassembles them "without needing an external storage dependency"; "if you don't need these extra functions, the vanilla Kafka Java clients should be preferred" | |
| 10 | linkedin/rest.li README | source | current | 2026-10-10 | https://github.com/linkedin/rest.li | The framework LinkedIn published and the industry did not adopt is now frozen, and its replacement is explicitly not drop-in | "There is no active development at LinkedIn on new features for Rest.li"; moving to gRPC for "better performance, support for more programming languages, streaming, and a robust open source community"; "gRPC is not a drop-in replacement" | |
| 11 | linkedin/rest.li closed-unmerged PRs matching grpc or protobuf | source | 2020-2026 | 2026-10-10 | https://github.com/linkedin/rest.li/pulls?q=is%3Apr+is%3Aclosed+is%3Aunmerged+grpc+OR+protobuf | The frozen framework is still being patched in 2026 for gRPC and xDS behaviour, which is what a shared discovery plane looks like from inside | 33 results, all closed and unmerged; "Fix gRPC hanging forever on name resolution errors in XdsClientImpl" (#1154, closed 2026-03-09), "Add support for XdsClient metrics for grpc" (#1051, closed 2025-09-05) | |
| 12 | linkedin/venice README | source | current | 2026-10-10 | https://github.com/linkedin/venice | Venice ships three client shapes at different cost-performance points behind one API, which is what lets a caller move without a rewrite | Da Vinci Client: "Stateful local cache, 0 network hops, < 1ms latency"; Thin Client under 10ms; Fast Client under 2ms; "active-active replication between regions using CRDT-based conflict resolution"; BSD-2-Clause, 613 stars | |
| 13 | linkedin/coral README | source | current | 2026-10-10 | https://github.com/linkedin/coral | LinkedIn's answer to engine churn in the data plane is a dialect-independent IR, not a chosen engine | "a translation, analysis, and query rewrite engine for SQL and other relational languages"; Coral IR "is independent of any SQL dialect"; HiveQL and Spark SQL in, HiveQL, Spark SQL and Trino SQL out; Trino-to-IR "in progress"; "The project is under active development" | |
| 14 | Voldemort project | voldemort/voldemort README | source | current | 2026-10-10 | https://github.com/voldemort/voldemort | The predecessor store was retired from production eight years before its repository stopped being the reference | "Voldemort is no longer under development."; "LinkedIn was the primary maintainer and user of Voldemort, and stopped all production usage in 2018."; "Most of the Voldemort Read-Only use cases and some of the Read-Write use cases have migrated to" Venice |
| 15 | Apache Kafka | KIP-500: Replace ZooKeeper with a Self-Managed Metadata Quorum | adr | 2019-2021 | 2026-10-10 | https://cwiki.apache.org/confluence/display/KAFKA/KIP-500:+Replace+ZooKeeper+with+a+Self-Managed+Metadata+Quorum | The canonical name for the pattern comes from Apache's own removal plan: a release whose only job is to isolate the dependency being removed | Retrieved via the session's search tool, which returned the compatibility plan as calling for a "bridge release" in which the ZooKeeper dependency is well isolated; the release does not remove ZooKeeper but eliminates most touch points, with access happening in the controller rather than in brokers, clients or tools |
| 16 | Confluent | Kafka Needs No Keeper: Removing the ZooKeeper Dependency | vendor | 2020 | 2026-10-10 | https://www.confluent.io/blog/removing-zookeeper-dependency-in-kafka/ | The bridge release's purpose is coexistence, which is what makes a zero-downtime upgrade possible at all | Reported via search: the bridge release "can coexist with both pre- and post-KIP-500 versions of Kafka", enabling zero-downtime upgrades; in the bridge release all brokers except the controller must treat ZooKeeper as read-only, with very limited exceptions |
| 17 | gRPC | gRFC A27: xDS-Based Global Load Balancing | adr | 2020-03-18, status Implemented | 2026-10-10 | https://github.com/grpc/proposal/blob/master/A27-xds-global-load-balancing.md | The opposite decision to a house protocol, written down: adopt the emerging standard because other data planes will speak it | Abstract says gRPC is moving from grpclb to xDS to "converge with this industry trend", and that the API "is evolving into a standard" other data plane software will use; Rationale: "This design could have been much simpler if we never intended to configure anything other than load balancing via xDS." |
| 18 | AWS | Announcing in-place ZooKeeper-to-KRaft cluster upgrades for Amazon MSK | vendor | 2024-2025 | 2026-10-10 | https://aws.amazon.com/blogs/big-data/announcing-in-place-zookeeper-to-kraft-cluster-upgrades-for-amazon-msk/ | Even a managed in-place metadata migration has a configuration precondition that fails the upgrade if violated | Reported via search: if you attempt to upgrade an MSK cluster to KRaft with altered advertised.listeners, the upgrade operation fails |
| 19 | How LinkedIn customizes Apache Kafka for 7 trillion messages per day | blog | 2019 | 2026-10-10 | https://www.linkedin.com/blog/engineering/open-source/apache-kafka-trillion-messages | LinkedIn's stated position in 2019 was that this is not a fork and stays close to upstream, at 7 trillion messages a day | Reported via search from the post: its version "is not a fork of Apache Kafka", each release branch "is branched off of the corresponding release branch of Apache Kafka", suffixed -li, with the aim to "maintain our releases as close as possible to upstream"; over 100 clusters, more than 4,000 brokers, more than 100,000 topics, 7 million partitions, over 7 trillion messages per day |
|
| 20 | Introducing Northguard and Xinfra: scalable log storage at LinkedIn | blog | 2025-06-25 | 2026-10-10 | https://www.linkedin.com/blog/engineering/data-streaming-processing/introducing-northguard-and-xinfra | Kafka was replaced at LinkedIn by a new log store plus a virtualised pub/sub layer that lets both run at once | Reported via search (primary post not fetchable from this container): Northguard is a scalable log storage system, Xinfra a virtualised pub/sub layer that abstracts the underlying log systems so applications use one interface whether data sits in Kafka or Northguard; engineers said they had "successfully migrated thousands of topics from Kafka to Northguard, accounting for trillions of records" | |
| 21 | InfoQ | LinkedIn Announces Northguard and Xinfra: Scaling beyond Kafka | casestudy | 2025-06 | 2026-10-10 | https://www.infoq.com/news/2025/06/linkedin-northguard-xinfra/ | The scale at which the incumbent stopped being operable, as reported by a second party | 32T records/day, 17 PB/day, 400K topics, 150 clusters; segments described as immutable units of replication |
| 22 | BigDATAwire | LinkedIn Introduces Northguard, Its Replacement for Kafka | blog | 2025-06-25 | 2026-10-10 | https://www.hpcwire.com/bigdatawire/2025/06/25/linkedin-introduces-northguard-its-replacement-for-kafka/ | Node count behind the 150 clusters | "17 PB across 400,000 topics, which run on more than 150 clusters accounting for more than 10,000 individual nodes" |
| 23 | SoftwareMill | Northguard and Xinfra: What LinkedIn Changed When Kafka Was No Longer Enough | blog | 2025 | 2026-10-10 | https://softwaremill.com/linkedin-northguard-xinfra-kafka/ | Northguard's unit of placement is the segment, not the partition, and Xinfra carries dual writes during migration | Reported via search: successive segments in the same range can have different replica sets, which LinkedIn calls log striping; a segment is sealed when it reaches 1 GB, has been active for over an hour, or loses a replica; Xinfra supports dual writes during migration with epoch-based ordering |
| 24 | SiliconANGLE | LinkedIn introduces Northguard and Xinfra to replace Kafka | blog | 2025-06-25 | 2026-10-10 | https://siliconangle.com/2025/06/25/linkedin-introduces-northguard-xinfra-replace-kafka/ | Northguard's metadata is itself a sharded Raft system rather than an external coordinator | Reported via search: metadata is stored in a fault-tolerant, sharded system backed by the Raft consensus protocol |
| 25 | InfoQ | Why LinkedIn chose gRPC+Protobuf over REST+JSON: Q&A with Karthik Ramgopal and Min Chen | blog | 2023-12 | 2026-10-10 | https://www.infoq.com/news/2023/12/linkedin-grpc-protobuf-rest-json/ | The stated reason for leaving a framework LinkedIn wrote is that nobody else adopted it, and the efficiency case was proven by running both side by side in production | Reported via search: gRPC supports bidirectional streaming, flow control and deadlines "which Rest.li does not support"; Rest.li is "primarily implemented in Java, with patchy or non-existent support for other programming languages"; LinkedIn open-sourced Rest.li "but this led to only limited adoption outside LinkedIn"; efficiency "validated via synthetic benchmarks as well as production ramps of gRPC and Rest.li services running side by side" |
| 26 | QCon London / InfoQ | gRPC Migration Automation at LinkedIn (Ramgopal, Chen) | talk | 2024-04 | 2026-10-10 | https://qconlondon.com/presentation/apr2024/grpc-migration-automation-linkedin | The migration's cost was dominated by mechanical edits, which is what made automation the deciding factor | Reported via search: AI helped change the RPC protocol for 50,000 production endpoints from Rest.li to gRPC; a planned 2-3 year manual migration "turned into an AI-supported migration lasting 2-3 quarters"; the listing cites roughly 8k services and about 100M lines of code; an intermediate gRPC bridged mode ran Rest.li and gRPC side by side on clients and servers |
| 27 | InfoQ | How LinkedIn rebuilt service discovery (coverage of Bohan Yang's post) | casestudy | 2026-02 | 2026-10-10 | https://www.infoq.com/news/2026/02/linkedin-service-discovery/ | A decade-old house format for service registration became the thing blocking adoption of the standard data planes, and the migration ran as a silent dual read | Reported via search: for over ten years LinkedIn used Apache ZooKeeper as its service discovery control plane with a custom format called D2; read storms caused high latency for reads and writes; the custom schemas "were incompatible with modern data plane technologies like gRPC and Envoy"; writes now go through Kafka and reads through an observer pushing xDS; one observer instance maintains 40k client streams and 10k updates per second; clients read from both ZooKeeper (source of truth) and the new system in shadow mode with metrics compared in the background |
| 28 | Unified streaming and batch pipelines at LinkedIn: reducing processing time by 94% with Apache Beam | blog | 2023-04 | 2026-10-10 | https://engineering.linkedin.com/blog/2023/unified-streaming-and-batch-pipelines-at-linkedin--reducing-proc | One pipeline definition running on two different engines is the same seam applied to compute | Reported via search from the post: 94% of processing time saved; backfilling moved to a unified Beam batch job, roughly halving memory and CPU and cutting processing time "from 7.5 hours to 25 minutes"; one codebase runs on Samza for streaming and Spark for backfilling | |
| 29 | Apache Beam | Case study: 4 trillion events daily at LinkedIn | vendor | 2023 | 2026-10-10 | https://beam.apache.org/case-studies/linkedin/index.html | Vendor-side figure for the same work, useful only as a magnitude | Reported via search: 4 trillion events daily; stream and batch unification credited with optimising cost-to-serve by 2x |
| 30 | Open Sourcing Venice: LinkedIn's Derived Data Platform | blog | 2022-09 | 2026-10-10 | https://www.linkedin.com/blog/engineering/open-source/open-sourcing-venice-linkedin-s-derived-data-platform | The predecessor's 500 production use cases moved because the replacement offered an API-compatible write path, not because teams rewrote | Reported via search from the post: "By 2018, the roughly 500 Voldemort Read-Only use cases running in production had been fully migrated to Venice, leveraging its Full Push capability as a drop-in replacement"; "In 2020, we built an alternative client library called Da Vinci... by serving all calls from local state rather than remote queries" | |
| 31 | Kafkaesque days at LinkedIn, part 1 (Joel Koshy) | postmortem | 2016-05 | 2026-10-10 | https://engineering.linkedin.com/blog/2016/05/kafkaesque-days-at-linkedin--part-1 | The oldest published LinkedIn incidents in this corpus are offset-management failures where neither recovery choice is safe | Reported via search from the post: resetting to the earliest offset causes duplicate consumption, while resetting to the latest can lose messages arriving between the reset and the next fetch; the incidents led to fixes in offset management, log compaction and monitoring | |
| 32 | LinkedIn / Kafka Summit | Kafkaesque days at LinkedIn in 2015 (slides) | talk | 2016 | 2026-10-10 | https://www.slideshare.net/jjkoshy/kafkaesque-days-at-linked-in-in-2015 | LinkedIn's own 2015 incident taxonomy already lists compatibility as a failure class alongside data loss | Reported via search: the deck lists the 2015 incidents as offset rewinds, data loss, cluster unavailability, "(in)compatibility", and blackout |
| 33 | You Broke Reddit: The Pi-Day Outage | postmortem | 2023-03-21 | 2026-10-10 | https://postmortem.io/incidents/reddit--2023-03-21--pi-day-outage/ | The missing artefact was the reverse path: an in-place Kubernetes minor upgrade with no supported downgrade | Reported via search from the postmortem: the site was down about 314 minutes on 2023-03-14 after a 1.23 to 1.24 upgrade; "there was no supported way to downgrade Kubernetes", so rolling back meant restoring from a backup, the restore procedure was outdated, and the subtle cause (a control-plane label a Calico route-reflector selector still pointed at) was not found until hours after the restore | |
| 34 | Roblox | Roblox Return to Service 10/28-10/31 2021 | postmortem | 2022-01 | 2026-10-10 | https://blog.roblox.com/2022/01/roblox-return-to-service-10-28-10-31-2021/ | A 73-hour outage from a minor-version feature, where the only fast remedy was the configuration switch that turned the new behaviour off | Reported via search from the post: Roblox moved from Consul 1.9 to 1.10 for a new streaming feature in the months before; the outage is attributed to that feature under heavy read and write load plus a BoltDB performance issue; recovery came from disabling streaming by configuration change; diagnosis was slowed because "critical monitoring systems that would have provided better visibility relied on affected systems, such as Consul" |
| 35 | CNBC | LinkedIn shelved plan to migrate to Microsoft Azure cloud | casestudy | 2023-12-14 | 2026-10-10 | https://www.cnbc.com/2023/12/14/linkedin-shelved-plan-to-migrate-to-microsoft-azure-cloud.html | The migration with no seam is the one that stopped, and the reported mechanism is lift and shift | Reported via search: CTO Raghu Hiremagalur's June 2022 memo said LinkedIn would keep using some Azure services and would "focus our efforts on scaling and innovating our on-prem infrastructure"; executives "stressed that the project was being put on hold, rather than getting canceled altogether"; reporting says issues arose when LinkedIn attempted to lift and shift its existing tools to Azure rather than refactor them; LinkedIn is building another data centre |
| 36 | The Register | Not even LinkedIn is that keen on Microsoft's cloud: shift to Azure abandoned | blog | 2023-12-14 | 2026-10-10 | https://www.theregister.com/2023/12/14/linkedin_abandons_migration_to_microsoft/ | The 2019 commitment, in the words of the engineering leader who made it | Reported via search, quoting Mohak Shroff's 2019 post: "With the incredible member and business growth we're seeing, we've decided to begin a multi-year migration of all LinkedIn workloads to the public cloud" |
| 37 | Tech Monitor | LinkedIn snubs its owner Microsoft and cans Azure cloud migration | blog | 2023-12 | 2026-10-10 | https://www.techmonitor.ai/hardware/cloud/linkedin-cloud-migration-microsoft-azure | Sources disagree on whether the migration is paused or dead, and that disagreement is itself the state of the record | Reported via search: a LinkedIn spokesperson confirmed "the move will now not go ahead", against CNBC's "put on hold" |
| 38 | Kreps, Narkhede, Rao (LinkedIn) | Kafka: a Distributed Messaging System for Log Processing | paper | 2011, NetDB workshop | 2026-10-10 | https://www.odbms.org/2011/01/kafka-a-distributed-messaging-system-for-log-processing/ | The system LinkedIn is now replacing was published by LinkedIn engineers as a log collector, not a universal substrate | Reported via search from the abstract: "a distributed messaging system that we developed for collecting and delivering high volumes of log data with low latency", suitable for both offline and online consumption |
| 39 | Qiao et al. (LinkedIn) | On Brewing Fresh Espresso: LinkedIn's Distributed Data Serving Platform | paper | SIGMOD 2013 | 2026-10-10 | https://dl.acm.org/doi/10.1145/2463676.2465298 | The one store that was never replaced is the one that shipped a change-capture stream as a first-class feature | Reported via search from the abstract: Espresso offers a hierarchical document model, transactional modification of related documents, real-time secondary indexing, on-the-fly schema evolution, and "a timeline-consistent change capture stream" |
| 40 | Rae et al. (Google) | Online, Asynchronous Schema Change in F1 | paper | VLDB 2013 (PVLDB 6:11) | 2026-10-10 | http://www.vldb.org/pvldb/vol6/p1045-rae.pdf | The formal statement of the bridge constraint: correctness survives only while no participant is more than one step behind | Reported via search from the abstract: corruption-causing schema changes are replaced with "a sequence of schema changes that is guaranteed to avoid corrupting the database so long as all servers are no more than one schema version behind at any time" |
| 41 | Apache Kafka / SoftwareMill | Apache Kafka 4.0.0 released: KRaft, queues, rebalance performance | vendor | 2025-03 | 2026-10-10 | https://softwaremill.com/apache-kafka-4-0-0-released-kraft-queues-better-rebalance-performance/ | The upstream deadline that dates LinkedIn's bridge: after 4.0 there is no ZooKeeper path at all, and 3.9 is the last door | Reported via search: Kafka 4.0, released 2025-03-18, is the first Apache Kafka to run entirely without ZooKeeper; direct broker upgrades to 4.0 come only from KRaft clusters on 3.3.x to 3.9.x, so ZooKeeper clusters must migrate first via the 3.9 bridge release |
What is not in the record
- No public figure for what LinkedIn's own data centres cost per unit of work, nor for what the abandoned Azure migration cost to attempt. Every cost claim in this guide is about effort or throughput, never dollars, because the dollars are not public.
- No published LinkedIn postmortem for the 3.0-li protocol version collision. The collision is documented in a pull request description, not an incident review, so whether it caused an outage or was caught before it could is not in the record.
- No account from inside the Azure migration. Everything about Blueshift's technical failure mode comes from reporters paraphrasing sources, not from engineers writing about their own work.