Storage migrations  / field guide
Practitioner field guide · 30 September 2026

The shim outlives the migration

Ten years of Sentry's data platform, reconstructed from the commits that deleted things: five storage engines retired behind four unchanged interfaces, a task queue that took ten years and five months to remove, and three dual-backend shims whose measured lifetimes are the number nobody budgets for. A reader finishes able to cost the transition period of a storage migration rather than the migration.

31 primary sources 7 repositories 6 incidents Evidence through September 2026 Read: 18 min
01

The territory

Every team that outgrows its first database writes a plan for the new one. Almost nobody writes a plan for the period when both are running. This guide measures that period.

8→69
Containers in the reference deployment, June 2016 to July 2026
91 mo
Age of the Redis/ClickHouse dual-read shim, still in the tree
126 mo
Celery's life in the monolith, from first commit to deletion
3.8→14 GB
Minimum RAM the project will start on, 2021 to 2024

State the problem without naming a product. A system has an interface that the rest of the application calls to read and write some class of data. The engine behind that interface stops fitting, for reasons that are usually about the shape of the data rather than its volume. You write a second implementation of the same interface, run both for a while so you can compare and roll back, then delete the first. The question this guide answers is what the middle step costs, because that is the step that appears in no design document and has no owner.

Sentry is an unusually good subject for that question. The company has published its error-tracking server as source since 2008; the monolith repository carries 111,408 commits as of 30 September 2026, and the self-hosted deployment repository has a commit on 11 February 2016, which gives a clean ten-year window. More usefully, Sentry does its storage swaps as pluggable backends in a directory per interface, so the arrival and the deletion of each engine are both single commits with dates. The transition period is not inferred here. It is subtracted.

The finding

Across the four interfaces examined, the engines behind them lasted between 17 and 72 months. The dual-backend shims written to bridge them lasted 4, 66 and 91 months. Two of the three shims outlived the engine they were introduced to retire, and the longest one is still in the tree, being pruned by a commit dated 25 August 2026, seven and a half years after it was added.

What this guide covers: the data platform and the asynchronous work that feeds it, read from public repositories and package registries. What it deliberately does not cover: Sentry's SaaS topology, its capacity, its costs and its published incident narratives, all of which live on hosts this session's egress policy could not reach. Where an incident appears below, it is reconstructed from the commit that fixed it and the pull request that explains it, never from an incident report, and that limit is stated on each card.

Figure 1 · Ten years, read from arrivals and deletions

20168 containers, Postgresand Redis2019Kafka and ClickHousearriveRiak, Cassandra, tagstores deleted2022RFC namesRabbitMQ as the limit2025Taskbroker shipsCelery deleted202669 containers2019 shim still beingprunedArrivals and deletions in the public record
20168 containers, Postgresand Redis2019Kafka and ClickHousearriveRiak, Cassandra, tagstores deleted2022RFC namesRabbitMQ as the limit2025Taskbroker shipsCelery deleted202669 containers2019 shim still beingprunedArrivals and deletions in the public record
Notice that the deletions cluster years after the arrivals, and that 2019 carries both the largest arrival and the largest set of deletions. Dates from the git history of getsentry/self-hosted and getsentry/sentry.
Diagram source
02

The architecture is the interface, not the engine

Four named interfaces have survived the decade unchanged in purpose. Everything behind them has been replaced at least once, and in one case four times.

The shape below is what the repositories describe today. Read it as two separate things. The flow from Relay through Kafka to the consumers is the pipeline, and it is the part that gets rewritten in other languages. The links from the monolith down to the stores are the interfaces, and they are the part that does not move. Sentry's directory layout says this out loud: the monolith carries src/sentry/nodestore, src/sentry/tagstore, src/sentry/search and src/sentry/tsdb, each holding one subdirectory per engine plus a base class, which is why a swap is a directory and a deletion rather than a rewrite.

Figure 2 · Reference architecture, September 2026

Storage

Consumers and async work

Transport

Edge, Rust

envelopes to topics

gRPC GetTask

backend interfaces

SDKs

Relay

Kafka

Ingest consumers
(Arroyo strategies)

Taskbroker, Rust

SQLite
inflight tasks

Task workers
(Python)

Snuba
query service

ClickHouse

Node store
object storage

Product monolith, Django

Storage

Consumers and async work

Transport

Edge, Rust

envelopes to topics

gRPC GetTask

backend interfaces

SDKs

Relay

Kafka

Ingest consumers
(Arroyo strategies)

Taskbroker, Rust

SQLite
inflight tasks

Task workers
(Python)

Snuba
query service

ClickHouse

Node store
object storage

Product monolith, Django

The pipeline runs top to bottom and is polyglot at both ends; the product monolith reaches storage only through backend interfaces. Reconstructed from the service list in docker-compose.yml, the Snuba architecture overview and the Taskbroker README.
Diagram source

The edge is a proxy, and it is Rust

Relay's README describes it as a service that "pushes some functionality from the Sentry SDKs as well as the Sentry server into a proxy process". Its first commit is 17 January 2018, under the name smith, renamed to semaphore on 9 May 2018 and later to Relay. The Python binding that lets the monolith call it, sentry-relay, has been on PyPI since 27 January 2020.

Evidence: relay README, PyPI

The store is ClickHouse behind a query service

Snuba's README states it "was originally developed to replace a combination of Postgres and Redis to search and provide aggregated data on Sentry errors". The architecture document gives the reason for the engine: ClickHouse "provides a good balance of the real time performance Snuba needs, its distributed and replicated nature, its flexibility in terms of storage engines and consistency guarantees".

Evidence: snuba README, architecture overview

Async work is a broker with its own database

Taskbroker consumes task activations from Kafka and "stores them in a SQLite database to avoid head-of-line blocking, enable out-of-order execution and per-task acknowledgements". Workers pull over gRPC. The interesting move is putting a local durable store back in front of workers because Kafka's offset model cannot express a per-task acknowledgement.

Evidence: taskbroker README, first commit 2024-11-01

The consumer layer deserves its own note, because it is the part that was extracted into a library rather than a service. Arroyo, first committed on 11 June 2021 and published to PyPI as sentry-arroyo on 29 June 2021, supplies the Kafka consumer and a set of composable strategies, described in its README as pre-built steps such as RunTask, Filter, Reduce and CommitOffsets that are "chained together to form complex message processing pipelines". By July 2026 the reference deployment runs 27 containers whose names begin snuba- and a further 14 consumers on the Sentry side. That is the real reason a consumer framework became a library: the number of consumers grew faster than the number of services.

Figure 3 · The instrument: one interface, many engines

writes, legacy reads

range and aggregate reads

Product code
calls one interface

BaseTSDB
abstract backend

RedisSnubaTSDB
routes per method

Redis
added 2014-04-24

Snuba / ClickHouse
added 2018-03-30

writes, legacy reads

range and aggregate reads

Product code
calls one interface

BaseTSDB
abstract backend

RedisSnubaTSDB
routes per method

Redis
added 2014-04-24

Snuba / ClickHouse
added 2018-03-30

The pattern behind every swap in this guide. The multi or combined backend is the shim, and it is the box that turns out to be permanent. Reconstructed from the directory layout and git history of src/sentry/tsdb.
Diagram source

One more component belongs in the reference shape even though it is not a data store. Between September 2022 and September 2023 the monolith grew src/sentry/silo and then src/sentry/hybridcloud, splitting the single application into a control plane and regional planes. That change was not driven by load. It is the shape a product takes when customers require their data to stay in a named jurisdiction, and it is worth noticing that it arrived through the same mechanism as the storage swaps: a decorator layer over the existing ORM models rather than a new system.

03

The decisions that matter

Four forks with the reasons written down at the time, and the condition that would flip each one.

Decision: what carries a task whose duration varies by four orders of magnitude?

Chosen
  • Kafka for transport, plus a broker holding inflight tasks in SQLite, with gRPC pull and per-task acknowledgement
  • Head-of-line blocking and out-of-order completion are solved in the broker, not the queue
Rejected
  • Staying on RabbitMQ with Celery. The 2022 RFC states it plainly: "Our use of RabbitMQ is reaching the limits of what can be done with this system safely. If we hit disk our entire pipeline crawls to a grind"
  • Plain Kafka consumers, which the RFC's own proposal for "increasingly slower topics" was an attempt to work around
Flips when
  • Task durations are tightly clustered. The whole design exists because, in the RFC's words, "a JavaScript event can make it in the low milliseconds through the entire pipeline whereas a native event can spend up to 30 minutes or more"
  • If your p99 task is within about 10x of your median, a partitioned queue is enough and the broker is overhead

The chain between those two artefacts is the most instructive thing in this guide. The constraint was named in an RFC dated 21 July 2022. The repository that answers it was created on 1 November 2024. The containers appeared in the reference deployment on 11 June 2025, and the last Celery code left the monolith on 26 September 2025, in a pull request whose body reads "Continue from #100302 and #99677 and remove another chunk of celery related code. Only one consumer and some settings remain." Three years and two months from naming a constraint to deleting the thing it constrained, with a numbered internal ticket reference as the only visible planning artefact.

Decision: where does Rust go, and where does it stay out?

Chosen
  • Rust at the edge (Relay, 2018), for symbolication (Symbolicator, 2019), for the task broker (2024) and for the uptime checker (2025)
  • Python keeps the product surface and the task bodies
Rejected
  • Publishing a Rust binding per crate to PyPI. The 2023 RFC records the reason: "Maintaining Python Bindings is itself a huge burden and cumbersome to do", with "SemVer API maintenance obligations" and repeated publishing failures
  • Rewriting the monolith. It is still Django in 2026
Flips when
  • The hot path is not separable from the product logic. Every Rust component here sits at a boundary where the data has a schema and the work is uniform
  • The RFC's stated answer was a single internal bindings package rather than many public ones, which is the move when your Rust is for you and not for the ecosystem

Decision: keep the in-house deployment tool or buy general pipelining?

Chosen
  • GoCD, described in the RFC as "general-purpose, mature (2007) pipelining tech", specifically for pipelines that gate other pipelines
Rejected
  • Continuing to build Freight, their own tool, which "has been in maintenance mode for a long time" because "primary contributors are no longer here"
  • ArgoCD, rejected as "not general-purpose pipelining technology"; Tekton as "a complex system with moving parts in k8s"; Argo Workflows as analogous to Tekton with a worse UI
Flips when
  • Your deployment unit is one Kubernetes cluster and your rollout question is convergence rather than sequencing. Then a cluster controller is the right shape and the pipeline engine is the overhead

That RFC also contains the single most useful sentence about in-house infrastructure in the corpus, and it is about a database migration in a deployment tool: "no one knows what alembic is and this actually resulted in some downtime as a recent Freight migration had to be figured out adhoc". The tool that deploys everything else had become a system whose schema nobody could change safely. That is the same failure as the long-lived shim, in a different subsystem.

DecisionChosenRejectedBecauseEvidence
Issue search engine, 2013 to 2015Back to PostgresSolr, then ElasticsearchBoth were added and deleted inside 22 and 17 months respectively; no public statement of the reason survivesgit history, src/sentry/search
Search and aggregation engine, 2018ClickHouse behind SnubaPostgres plus RedisReal-time performance, replication, and storage-engine flexibility in one systemSnuba architecture overview
Exactly-once ingestionDeduplicating table enginesTransactional writes"we can achieve exactly once semantics if we accept eventual consistency"Snuba architecture overview
Task serialization formatMove off pickleKeeping pickle"it's not possible for code outside of the Sentry monolith to dispatch tasks"RFC 0002, 2022-07-21
Kafka coordinationKRaft, ZooKeeper removedKeeping ZooKeeperOne fewer stateful system in a deployment that had reached 56 containersself-hosted#3263, 2024-08-13
Change data capture from PostgresRemovedKeeping wal2json CDCDeleted alongside the ZooKeeper change; the compose file simply stops carrying itself-hosted#3260, 2024-08-12
Event body storageObject storage, 2025Bigtable-only or Postgres-only node storeBrings the self-hosted shape in line with an S3 API rather than a cloud-specific oneself-hosted#3498, 2025-09-13
Container base image, 2026DistrolessDebian slim images"hundreds of OS packages that the application itself never uses", each needing CVE triageRFC 0157, 2026-03-27

Two things about that table are worth saying directly. First, the 2013 to 2015 search experiments are the only decisions here with no recorded reason at all: two engines were added and removed, and the argument is not in the repository. Second, the decisions with the best-recorded reasoning are the ones about tooling and packaging, not about data stores. Sentry's public RFC repository opened on 21 July 2022, long after the storage moves that mattered most, and of its 64 merged records the overwhelming majority concern SDK behaviour and product semantics. The biggest architectural changes in this guide were decided somewhere that does not publish.

Figure 4 · Should you write the shim?

yes

no

no

yes

no

yes

Can the old engine
be read-only during
the cutover?

Copy, freeze, switch.
No shim, no dual path

Do you need to compare
old and new answers
in production?

Dual write, single read.
Delete the writer on cutover

Is there a named owner
and a dated deletion
ticket for the shim?

Do not start.
The shim becomes
permanent infrastructure

Route per method,
alarm on legacy reads,
delete on the date

yes

no

no

yes

no

yes

Can the old engine
be read-only during
the cutover?

Copy, freeze, switch.
No shim, no dual path

Do you need to compare
old and new answers
in production?

Dual write, single read.
Delete the writer on cutover

Is there a named owner
and a dated deletion
ticket for the shim?

Do not start.
The shim becomes
permanent infrastructure

Route per method,
alarm on legacy reads,
delete on the date

The terminal nodes are actions, and the left-hand one is the one teams skip. Derived from the observed lifetimes in section 05, not from any published recommendation.
Diagram source
04

What broke in production

Sentry numbers its incidents and the numbers leak into public commit messages. None of the narratives below is published; each card is reconstructed from the commit that closed the incident and the pull request that explains the change.

This is worth dwelling on as a method, because it generalises to any company that develops in the open. Commit subjects in the Snuba, Arroyo and Relay repositories carry references of the form inc-NNNN. The lowest observed is INC-181, fixed on 28 July 2022; the highest is INC-2179, fixed on 19 May 2026. If the identifiers are sequential and cover one incident stream, that is roughly two thousand numbered incidents in forty-six months, which works out near forty a month. That inference is mine and the numbering scheme is not published, so treat the rate as an order of magnitude and the individual identifiers as solid. What matters more is the distribution: the fixes cluster in the consumer framework and the query service, not in the storage engines that the migrations were about.

Class A · The error path fails harder than the path it protects

Incident fix

The dead-letter queue took down the consumer it was protecting

AssumptionA dead-letter queue is a safety valve, so its own failure is a minor event.
What happenedThe author's description is one sentence: "Right now, if the produce fails for any reason, the consumer backlogs, allow graceful degradation instead". A failed produce to the DLQ propagated as an error into the consumer loop, so bad messages stopped the pipeline instead of leaving it.
Blast radiusNot published. The mechanism is per-consumer, and Arroyo is the shared library under both Snuba and Sentry consumers, so the class of exposure is every topic.
FixFailures to produce to the DLQ are absorbed rather than raised, merged 24 January 2025.
Design ruleEvery error path needs an error path. Write down what happens when the DLQ, the retry topic or the quarantine table is the thing that is unavailable, and make that outcome weaker than stopping.
Incident fix

A timeout that did not time out under the load it existed for

AssumptionA join with a timeout returns within the timeout.
What happenedThe pull request states that "Pool.join will block for more than timeout seconds if there is a task that takes more than timeout seconds", and notes the problem is made worse when the next step rejects messages and batches are reprocessed. The shutdown and rebalance paths therefore hung exactly when the pool was overloaded.
Blast radiusNot published. Two commits on the same day, plus a follow-up to stop the redundant reprocessing.
FixTerminate the pool rather than wait, using the framework's own result-checking loop, merged 22 May 2023.
Design ruleA timeout inherited from a library is a claim, not a guarantee. Test the deadline under the condition that motivates it, which is the overloaded case your load test does not produce.

Figure 5 · How a failing dead-letter queue stops a healthy pipeline

OffsetsDLQ topicConsumerKafka topicOffsetsDLQ topicConsumerKafka topicoffsets stop advancing,lag grows on a healthy topicvalid messagecommitinvalid messageproduce to DLQproduce failsraise, stop processing
OffsetsDLQ topicConsumerKafka topicOffsetsDLQ topicConsumerKafka topicoffsets stop advancing,lag grows on a healthy topicvalid messagecommitinvalid messageproduce to DLQproduce failsraise, stop processing
The ordering is the point: the invalid message is handled, and the handling is what fails. Reconstructed from the change in arroyo#425; the incident narrative is not public.
Diagram source

Class B · Degrade the answer rather than the availability

Incident fix

Rate limiting was the only lever, so queries were refused instead of coarsened

AssumptionUnder storage overload the correct response is to admit fewer queries.
What happenedDuring high-load incidents affecting ClickHouse the available control was aggressive rate limiting. The change adds a switch that forces queries to a downsampled tier, tier 8 rather than tier 1, which the author says "will lead to a degraded performance for customers" while being better than the alternative.
Blast radiusNot published. The same commit turns off a cluster load-info lookup that was not being used elsewhere, which suggests the overload path was also doing unnecessary work.
FixA tier-forcing killswitch, merged 8 June 2026. Relay carries the same idea: a killswitch for trace-id partitioning added 7 May 2025.
Design ruleIf your data is sampled or tiered, resolution is a load-shedding dimension and it is cheaper than rejection. Build the coarse path before you need it, because you cannot introduce sampling during an incident.
Incident fix

The liveness probe failed on a cluster the request path did not need

AssumptionA health check should verify everything the service can reach.
What happenedThe liveness health check checked all ClickHouse clusters, so a non-essential cluster being unavailable marked otherwise serviceable instances unhealthy. The fix narrows the check to essential clusters only.
Blast radiusNot published. The failure converts a partial dependency outage into an instance-level one, which is the mechanism by which health checks amplify rather than contain.
FixMerged 14 May 2026, four weeks after a related change raising a distributed DDL timeout to 300 seconds for a separate incident.
Design ruleA liveness probe answers one question: should this process be killed. Anything it checks beyond that becomes a way for a degraded dependency to remove your capacity.

Class C · Input the schema permits and the code does not expect

Incident fix

A flag said an event existed, and the edge dropped the events where it did not

AssumptionAn envelope with creates_event set to true contains an event.
What happenedIn edge relays, metrics extraction encountered envelopes flagged as creating an event that carried no event, particularly Unreal Engine payloads and potentially minidumps, and dropped them. Customer data was discarded at the point of the system furthest from anyone who could notice.
Blast radiusNot published. Scope is the point-of-presence relay configuration and the affected payload types.
FixMerged 28 July 2022, referencing both an incident and an issue identifier.
Design ruleA boolean that describes another field is a second source of truth. At an edge that may drop data, derive from the payload and treat the flag as a hint.
Incident fix

Timestamps outside the retention window reached a time-partitioned store

AssumptionIngested timestamps fall inside the window the tables are partitioned for.
What happenedTwo commits five days apart drop messages with out-of-range timestamps on the EAP items and accepted-outcomes topics, and add a job to drop stored data older than thirty days. The sequence reads as containment first, cleanup second.
Blast radiusNot published. In a store partitioned by time, out-of-range rows create partitions that nothing queries and nothing reaps.
FixMerged 19 and 20 May 2026. A related change on 5 May 2026 put a hard limit on buffered messages for batched deletes.
Design ruleValidate against the partitioning key at ingestion, not at query time. A store whose physical layout encodes an assumption needs that assumption enforced at the boundary.
What the record cannot tell you

No incident in this corpus is attributed to a storage engine losing data, and none is attributed to a dual-backend shim returning inconsistent answers. Both absences are ambiguous. They may mean the interface-and-shim method works, or they may mean that inconsistency between two read paths is not the kind of thing a commit message records. The second reading is more likely, because a shim's whole purpose is to let one side be wrong quietly. If you run one, the reconciliation job that compares the two paths is the only instrument that would tell you.

05

Numbers you can plan against

Every figure below is counted out of a public repository or a package registry on 30 September 2026. None is a vendor claim, and none describes Sentry's own production capacity, which is not public.

MetricValueAtContextAs ofSource
Containers in reference deployment8 / 7 / 23 / 30 / 38 / 56 / 60 / 69self-hostedMid-year snapshots for 2016, 2019, 2020, 2021, 2023, 2024, 2025, 20262026-07compose file
Minimum RAM enforced by the installer3,800 MB → 15,900 MB → 14,000 MBself-hostedHard floor; a soft floor of 7,800 MB existed in 20212021-03 to 2024-08_min-requirements.sh
Minimum CPU cores2 → 4self-hostedRaised with the RAM floor in August 20242024-08_min-requirements.sh
Engine lifetime, node store (Riak)72 monthssentryAdded 2013-10-17, deleted 2019-10-142019-10git history
Engine lifetime, node store (Cassandra)66 monthssentryAdded 2013-11-05, deleted 2019-04-262019-04git history
Engine lifetime, issue search (Solr)22 monthssentryAdded 2013-10-20, deleted 2015-08-182015-08git history
Engine lifetime, issue search (Elasticsearch)17 monthssentryAdded 2014-05-20, deleted 2015-10-042015-10git history
Engine lifetime, tag store v222 monthssentryAdded 2017-11-27, deleted 2019-09-30; a full rewrite abandoned2019-09git history
Shim lifetime, tag store multi-backend4 monthssentryAdded 2017-12-19, deleted 2018-04-24; the only shim that was removed promptly2018-04git history
Shim lifetime, node store multi-backend66 monthssentryAdded 2013-10-17, deleted 2019-04-262019-04git history
Shim lifetime, Redis and ClickHouse time series91 months, opensentryAdded 2019-02-28, last pruned 2026-08-252026-09redissnuba.py
Celery in the monolith126 monthssentry2015-04-09 to 2025-09-26; replacement landed 2025-01-312025-09sentry#100326
ClickHouse schema migrations in tree252snubaFiles under snuba/snuba_migrations2026-09snuba_migrations
Commits in the monolith111,408sentryFirst commit 2008-05-122026-09-30getsentry/sentry
Merged RFCs64rfcsRepository opened 2022-07-21; proposal numbers reach 1572026-09rfcs/text
Observed incident identifiersINC-181 to INC-2179snuba, arroyo, relayRange in public commit subjects; scheme not published2022-07 to 2026-05commit history
Server published as a Python package371 releases, endedPyPI2012-01-04 to 2023-07-25, final version 23.7.12023-07pypi.org/project/sentry
Python SDK generations187 then 351 releasesPyPIraven ended 2018-12-19; sentry-sdk began 2018-07-262026-09-28pypi.org/project/sentry-sdk
Browser SDK generations88 then 750 versionsnpmraven-js ended 2019-06-04; @sentry/browser began 2017-12-072026-09-29npmjs.com/@sentry/browser

Three of these deserve interpretation. The container count is the honest measure of what a decade of storage specialisation costs to operate: the deployment did not get more complex gradually, it jumped from 7 to 23 in a single release in November 2019 and then grew steadily as each new product class arrived, profiling in 2022, replays in 2022, crons in 2023, feedback in 2023, uptime in 2024. The RAM floor is the same fact expressed in money, and it roughly quadrupled in three years while the product's core promise did not change. Both numbers are about the self-hosted artefact rather than the hosted service, so treat them as a lower bound on the real operational surface.

The shim lifetimes are the numbers to carry into your own planning, and the arithmetic is simple subtraction of two commit dates, so check it yourself. Three shims, three very different outcomes: 4 months when the migration was small and recent, 66 months when the engine being retired had two replacements in flight, and 91 months and still counting for the one that spans two data models. The median of the three is the one to plan against, and the median is 66 months. That is not a migration. That is a subsystem you now own.

Figure 6 · Engines and shims, measured in months

2014201520162017201820192020202120222023202420252026Node store, Riak Node store multi Issue search, Solr Node store, Cassandra Issue search, Elastic Tag store v2 Tag store multi Redis and ClickHouse EnginesShimsLifetime in the tree, engines above, shims below
2014201520162017201820192020202120222023202420252026Node store, Riak Node store multi Issue search, Solr Node store, Cassandra Issue search, Elastic Tag store v2 Tag store multi Redis and ClickHouse EnginesShimsLifetime in the tree, engines above, shims below
The bars in the lower group are the transition mechanisms. Two of the three are longer than the engine retirements they were written to enable. Dates from the git history of getsentry/sentry.
Diagram source
Read these carefully

Every lifetime here is the lifetime of a file in a repository, which is a proxy for the lifetime of a deployed backend and not the same thing. A backend can be dead in production for a year before the directory is deleted, which would make these figures upper bounds, or it can be kept for self-hosted users after the hosted service has moved off it, which makes them incomparable to Sentry's own operations. The incident rate derived from identifier numbering is an inference, not a report. The container counts describe the self-hosted deployment only.

06

The evidence wall

Thirty-one sources, all fetched on 30 September 2026. Every one is a repository artefact or a package registry record, which is both this page's strength and its limit: see the note at the end of this section.

ADR Sentry2022-07

RFC 0002: New Architecture

A living informational RFC naming the constraints on the pipeline: RabbitMQ near its safe limit, pickle preventing non-Python task dispatch, buffers that require pickle and cannot be filtered, and symbolication task durations spanning milliseconds to thirty minutes.

Carry forwardThe serialization format, not the network, is what makes a monolith monoglot.
github.com/getsentry/rfcs/blob/main/text/0002-new-architecture.md
Source Sentry2024-11

Taskbroker README

Describes the answer to RFC 0002's queue problem: Kafka for transport, a SQLite store for inflight activations "to avoid head-of-line blocking, enable out-of-order execution and per-task acknowledgements", and gRPC GetTask and SetTaskStatus for workers.

Carry forwardWhen the log's offset model cannot express per-item completion, the broker needs its own durable state.
github.com/getsentry/taskbroker/blob/main/README.md
Source Sentry2019 to 2026

sentry/tsdb/redissnuba.py

The dual-backend shim itself. A method specification table routes each time-series call to Redis or to Snuba by read or write, with one entry bound to a function called dont_do_this that raises NotImplementedError. Added 2019-02-28; a commit on 2026-08-25 removes "the dead Redis cmsketch frequency tables".

Carry forwardA shim is a routing table over two engines, and routing tables acquire exceptions faster than they lose them.
github.com/getsentry/sentry/blob/master/src/sentry/tsdb/redissnuba.py
Source Sentry2013 to 2019

Backend directories: nodestore, tagstore, search, tsdb

Git history for four interface directories gives the add and delete date of every engine: five node stores, four issue-search backends, four tag stores. Three of the five node stores and two of the four search backends are gone.

Carry forwardA directory per engine behind one base class makes a migration a diff, and makes the retirement auditable years later.
github.com/getsentry/sentry/tree/master/src/sentry/nodestore
Incident fix Sentry2025-01

arroyo#425, inc-1013: DLQ produce failure backlogs the consumer

"Right now, if the produce fails for any reason, the consumer backlogs, allow graceful degradation instead." The shared consumer library's dead-letter path could stop the pipeline it protects.

Carry forwardSpecify the behaviour when the error path is the unavailable component.
github.com/getsentry/arroyo/pull/425
Incident fix Sentry2023-05

arroyo#237, INC-378: join blocks past its timeout

"Pool.join will block for more than timeout seconds if there is a task that takes more than timeout seconds", made worse when the next step rejects messages and batches are reprocessed. Fixed by terminating the pool.

Carry forwardVerify library deadlines under the overload they exist for, not under normal load.
github.com/getsentry/arroyo/pull/237
Incident fix Sentry2026-06

snuba#8002, inc-2173: force a downsampled tier instead of rate limiting

Adds a killswitch that defaults queries to tier 8 rather than tier 1 during ClickHouse overload, which the author acknowledges "will lead to a degraded performance for customers" but prefers to aggressive rate limiting.

Carry forwardIf your data is tiered, resolution is a load-shedding lever. Build it before the incident.
github.com/getsentry/snuba/pull/8002
Incident fix Sentry2026-05

snuba#7925, inc-2141: liveness check scoped to essential clusters

The liveness health check covered all ClickHouse clusters, so a non-essential cluster outage removed serviceable capacity. Narrowed to essential clusters only.

Carry forwardA liveness probe should test only what justifies killing the process.
github.com/getsentry/snuba/pull/7925
Incident fix Sentry2022-07

relay#1355, INC-181: events dropped at the edge

Envelopes with creates_event set true but carrying no event, notably Unreal payloads, were dropped during metrics extraction in point-of-presence relays.

Carry forwardAt a boundary that may discard data, derive from the payload rather than trusting a descriptive flag.
github.com/getsentry/relay/pull/1355
Incident fix Sentry2026-05

snuba#7945, INC-2179: out-of-range timestamps dropped at ingestion

Messages with timestamps outside the retention window are dropped on the EAP items and accepted-outcomes topics; a companion change deletes stored data older than thirty days.

Carry forwardEnforce the partitioning key's assumptions at the ingestion boundary.
github.com/getsentry/snuba/pull/7945
ADR Sentry2022-09

RFC 0042: GoCD succeeds Freight

The most candid document in the corpus. Freight "has been in maintenance mode for a long time" because "primary contributors are no longer here", and a Freight schema migration caused downtime because "no one knows what alembic is". ArgoCD, Tekton and Argo Workflows are each rejected with a stated reason.

Carry forwardAn in-house tool's real expiry date is the departure of the people who wrote its migrations.
github.com/getsentry/rfcs/blob/main/text/0042-gocd-succeeds-freight-as-our-cd-solution.md
ADR Sentry2023-10

RFC 0119: Rust in Sentry

Chooses a single internal bindings package over per-crate PyPI packages, because "maintaining Python Bindings is itself a huge burden", public packages create SemVer obligations, and publishing had been failing. Existing bindings are noted as Sentry-specific with minimal external use.

Carry forwardPublishing your glue code publicly buys you compatibility obligations you did not want.
github.com/getsentry/rfcs/blob/main/text/0119-rust-in-sentry.md
ADR Sentry2026-03

RFC 0157: Distroless base images

Approved decision to replace Debian-based images, on the grounds that standard bases ship "hundreds of OS packages that the application itself never uses", each of which still requires CVE triage.

Carry forwardThe maintenance cost of an unused dependency is triage, which is paid whether or not it is exploitable.
github.com/getsentry/rfcs/blob/main/text/0157-distroless-base-images.md
ADR Sentry2022 to 2026

The RFC corpus itself

Sixty-four merged records from 2022-07-21 to 2026-09-25, with proposal numbers reaching 157. Volume peaks in 2023 and thins sharply afterwards, and the merged set is dominated by SDK and product semantics rather than platform architecture.

Carry forwardRead a public RFC repository for what it omits: the decisions with the largest blast radius here were not recorded in it.
github.com/getsentry/rfcs/tree/main/text
Source Sentry2016 to 2026

self-hosted docker-compose.yml

The single best instrument in this hunt. Mid-year snapshots give the container count across the decade, and the arrival and departure commit of every service: Kafka, ClickHouse, Snuba and Symbolicator on 2019-11-12, Relay on 2020-04-24, Taskbroker on 2025-06-11, object storage on 2025-09-13.

Carry forwardA project's deployment manifest, read along its history, is a dated architecture diagram nobody had to draw.
github.com/getsentry/self-hosted/blob/master/docker-compose.yml
Source Sentry2019-11

self-hosted#220: make on-premise work for Sentry 10

The commit that takes the reference deployment from 7 containers to 23, introducing Kafka, ZooKeeper, ClickHouse, the Snuba services and Symbolicator at once. Dated six days after the switch to the Business Source License.

Carry forwardArchitectural specialisation shows up in the operator's cost first, and in one step rather than gradually.
github.com/getsentry/self-hosted/pull/220
Source Sentry2025-09

sentry#100326: remove more celery related code, part three

"Continue from #100302 and #99677 and remove another chunk of celery related code. Only one consumer and some settings remain." The deletion of celery.py, ten years and five months after it was added.

Carry forwardRemovals arrive in numbered parts; budget the last twenty percent, which is where the odd callers live.
github.com/getsentry/sentry/pull/100326
Source Sentry2025-09

self-hosted#3946: remove the worker and cron containers

The operational counterpart to the Celery deletion. The two containers present since May 2016 leave the reference deployment, replaced by taskbroker, taskworker and taskscheduler which arrived on 2025-06-11.

Carry forwardA queue replacement is finished when the old containers are gone, not when the new ones start.
github.com/getsentry/self-hosted/pull/3946
Source Sentry2024-08

self-hosted#3263: migrate to ZooKeeper-less Kafka

Removes a stateful coordination service from a deployment that had reached 56 containers. Merged one day after the removal of Postgres change data capture and wal2json.

Carry forwardDeleting a dependency is a capacity decision in a deployment measured by container count.
github.com/getsentry/self-hosted/pull/3263
Source Sentry2024-08

self-hosted#3260: remove cdc and wal2json

Change data capture out of Postgres, and the logical decoding plugin behind it, leave the reference deployment with no replacement service appearing in the compose file.

Carry forwardA capability can be removed rather than migrated; check whether the need survived the architecture that created it.
github.com/getsentry/self-hosted/pull/3260
Source Sentry2025-09

self-hosted#3498: S3 node store with SeaweedFS

Event bodies move to an S3-compatible object store in the reference deployment, six years after Bigtable became the hosted node store and a decade after the original Postgres one.

Carry forwardThe last storage layer to move to object storage is usually the one holding opaque blobs, which is the one that should have moved first.
github.com/getsentry/self-hosted/pull/3498
Source Sentry2021 to 2025

self-hosted/install/_min-requirements.sh

The installer's enforced floor: 3,800 MB hard and 7,800 MB soft with 2 cores in March 2021; 15,900 MB and 4 cores in August 2024, trimmed to 14,000 MB six days later; a two-tier scheme from March 2025.

Carry forwardAn installer's minimum requirements file is a dated, honest record of what an architecture costs to run.
github.com/getsentry/self-hosted/blob/master/install/_min-requirements.sh
Source Sentry2018 to 2026

Snuba README

States the origin directly: Snuba "was originally developed to replace a combination of Postgres and Redis to search and provide aggregated data on Sentry errors", and has since absorbed most time-series features. First commit 2018-03-01; a ClickHouse query appears on day one.

Carry forwardThe service that replaces two stores tends to become the store for everything shaped like either.
github.com/getsentry/snuba/blob/master/README.rst
Source Sentry2026-09

Snuba architecture overview

Gives the engine rationale and the consistency bargain: ClickHouse for "a good balance of the real time performance Snuba needs", one writer per table, and "exactly once semantics if we accept eventual consistency" through deduplicating table engines.

Carry forwardExactly-once via deduplication is a read-side property; it costs you freshness, not throughput.
github.com/getsentry/snuba/blob/master/docs/source/architecture/overview.rst
Source Sentry2021 to 2026

Arroyo README

The consumer framework extracted as a library rather than a service: Kafka backends, composable strategies such as RunTask, Filter, Reduce and CommitOffsets, and a stream processor that schedules work and controls consumer progress.

Carry forwardWhen consumers outnumber services, the reuse unit is a library with a strategy interface, not another service.
github.com/getsentry/arroyo/blob/main/README.md
Source Sentry2018 to 2026

Relay README and repository history

"A service that pushes some functionality from the Sentry SDKs as well as the Sentry server into a proxy process." First commit 2018-01-17 as smith, renamed semaphore on 2018-05-09. The repository now carries explicit rules for AI usage in a HOWTOAI.md.

Carry forwardMoving work to the edge is a licensing and trust decision as much as a latency one, which is why this component is the one with its own rules files.
github.com/getsentry/relay/blob/master/README.md
Registry PyPI2012 to 2023

The server stopped being a Python package

371 releases of sentry from 2012-01-04 to 2023-07-25, ending at version 23.7.1. After that the product exists as a fleet of containers rather than something installable with a package manager.

Carry forwardThe date your artefact stops being installable is the date your architecture stopped being a program.
pypi.org/project/sentry
Registry PyPI2011 to 2026

raven and sentry-sdk, the client reset

187 releases of raven ending 2018-12-19; sentry-sdk starts 2018-07-26 and reaches 351 releases by 2026-09-28. The two overlap for five months.

Carry forwardA client rewrite is dated by the overlap window, and five months of overlap implies a hard break rather than a migration.
pypi.org/project/sentry-sdk
Registry npm2013 to 2026

raven-js and @sentry/browser

88 versions of raven-js from 2013-12-20 to 2019-06-04; 750 versions of @sentry/browser from 2017-12-07 to 2026-09-29. The browser SDK reset began eighteen months before the Python one ended.

Carry forwardRelease cadence, counted per platform, shows which client ecosystem sets the pace of your protocol changes.
npmjs.com/package/@sentry/browser
Registry PyPI2021 to 2026

sentry-arroyo releases

205 releases from 2021-06-29 to 2026-09-25, eighteen days after the repository's first commit. A shared consumer framework versioned publicly while being used internally.

Carry forwardA library published on the first sprint is a commitment to compatibility; check whether you wanted that.
pypi.org/project/sentry-arroyo
Registry PyPI2020 to 2026

sentry-relay bindings

93 releases from 2020-01-27 to 2026-09-23. These are the bindings RFC 0119 later describes as a maintenance burden, which is the evidence that the RFC was describing lived experience rather than anticipating a problem.

Carry forwardFour years of public releases is how long it took the binding cost to become a decision.
pypi.org/project/sentry-relay
The shape of this evidence, stated plainly

There are no engineering blog posts, papers or conference talks in this wall, and that is not because none exist. This session's network policy could reach GitHub and the package registries and nothing else, so Sentry's own published narratives, its incident reports and any third-party measurement were unavailable. The consequence is a page built almost entirely from tiers nine to eleven of the evidence hierarchy, which is the strongest material for what was built and the weakest for how it felt to operate. Three distinct hosts is well below what breadth would normally require. Where a reason is quoted here it is quoted from a decision record; where an outcome is described it is reconstructed from a diff. Nothing in this guide should be read as Sentry's account of its own history.

07

Build a miniature, then productionise it

Six rungs. The line between a toy and the real thing is rung four, where you deliberately break the error path.

One interface, two engines

Take any read-write path in a service you own and put an abstract backend in front of it, with the current store as the first implementation and a second store, in a different engine family, as the second. Select by configuration.

Done when: the test suite passes against both backends with no test changes.  Teaches: which of your callers depend on engine behaviour rather than on the interface.

The comparison harness

Add a mode that reads from both and logs divergences with the query arguments, without serving the new answer. Run it on real traffic for a week.

Done when: you can state the divergence rate per method and explain every class of difference.  Teaches: that most divergences are semantic, not bugs, and that the semantics were never written down.

Route per method, and alarm on the old path

Build the routing table Sentry's shim uses: each method tagged read or write and pointed at one engine. Then add the piece the original lacks, a counter per legacy read, exported by method name.

Done when: a dashboard shows legacy reads by method and it is the metric you would use to decide the migration is finished.  Teaches: a shim without that counter cannot tell you it is done, which is why shims do not end.

Break the error path

Make the dead-letter queue, quarantine table or retry topic unavailable while valid traffic flows. Watch whether consumption stops. Then make the failure absorbed rather than raised, and watch what you lose instead.

Done when: an unavailable DLQ degrades throughput and does not halt it, and you can say what data is lost in that state.  Teaches: the choice in arroyo#425, which is that graceful degradation of the error path always costs something and the cost has to be named.

Per-task acknowledgement over a log

Put a small durable store between a partitioned log and a pool of workers so tasks can complete out of order and be acknowledged individually. Feed it a workload whose durations span three orders of magnitude.

Done when: a single thirty-minute task does not delay thousands of millisecond tasks behind it in the same partition.  Teaches: why Sentry put SQLite in a broker in 2024 rather than adding partitions.

Resolution as a load-shedding lever, and a deletion date

Add a switch that answers queries from a coarser or sampled source under load, and measure the error it introduces. Then write the deletion ticket for the shim from rung three: a named owner, a date, and an alert that fires when legacy reads are non-zero after that date.

Done when: the coarse path is exercised in a game day and the deletion ticket exists with a date in it.  Teaches: the two things the ten-year record says are missing, a degradation lever built before the incident and a deletion that is somebody's job.

08

Keep hunting

This page was assembled with git and two package registries, not with a search engine. These are the commands that produced it; they work on any company that develops in the open.

Date the arrival and death of a component

  • git log --reverse --diff-filter=A --format=%ad --date=short -- path/to/backend
  • git log --diff-filter=D --format=%ad --date=short -- path/to/backend
  • git log --all --name-only --format= --diff-filter=A -- 'src/*/nodestore/*' | sort -u

Read a deployment manifest as a dated architecture diagram

  • git rev-list -1 --before=2020-07-01 HEAD
  • git show <sha>:docker-compose.yml | grep -E '^ [a-z0-9_.-]+:$'
  • git log -S' zookeeper:' --format='%ad %s' --date=short -- docker-compose.yml

Find the incidents inside a public repository

  • git log -i --grep='inc-[0-9]' -E --format='%ad | %s' --date=short
  • git log -i --grep='postmortem\|root cause\|data loss' --format='%ad | %s' --date=short
  • git log --format='%ad | %s' --date=short --grep='^Revert'

Find what a project stopped shipping

  • curl -s https://pypi.org/pypi/<pkg>/json | jq -r '.releases | to_entries[] | [.key, (.value[0].upload_time // "")] | @tsv'
  • curl -s https://registry.npmjs.org/<pkg> | jq -r '.time | to_entries[] | [.key, .value] | @tsv'
  • git ls-remote --heads <repo> # unmerged proposal branches

Find the decision records, and what is missing from them

  • grep -m1 -i 'Start Date' text/*.md
  • grep -ho 'rfcs/pull/[0-9]*' text/*.md | sort -u # numbering gaps are rejections
  • git log --format='%ad' --date=format:%Y -- text/ | sort | uniq -c

Cost, read from the installer

  • git log -p --date=short --format='%ad' -- install/_min-requirements.sh | grep -E '^[+-]MIN_'
  • git log --format='%ad | %s' --date=short -- Dockerfile | tail -40
09

References

  1. Sentry, RFC 0002: New Architecture getsentry/rfcs, start date 2022-07-21. Checked 2026-09-30.
  2. Sentry, RFC 0042: GoCD succeeds Freight as our CD solution getsentry/rfcs, start date 2022-09-20. Checked 2026-09-30.
  3. Sentry, RFC 0119: Rust in Sentry getsentry/rfcs, start date 2023-10-25. Checked 2026-09-30.
  4. Sentry, RFC 0157: Distroless base images getsentry/rfcs, start date 2026-03-27, status approved. Checked 2026-09-30.
  5. Sentry, the RFC corpus getsentry/rfcs, 64 merged records, 2022-07-21 to 2026-09-25. Checked 2026-09-30.
  6. Sentry, self-hosted deployment manifest getsentry/self-hosted, repository begins 2016-02-11. Checked 2026-09-30.
  7. Sentry, installer minimum requirements getsentry/self-hosted. Checked 2026-09-30.
  8. Sentry, self-hosted#220: make on-premise work for Sentry 10 Merged 2019-11-12. Checked 2026-09-30.
  9. Sentry, self-hosted#3260: remove cdc and wal2json Merged 2024-08-12. Checked 2026-09-30.
  10. Sentry, self-hosted#3263: migrate to ZooKeeper-less Kafka Merged 2024-08-13. Checked 2026-09-30.
  11. Sentry, self-hosted#3498: use S3 node store with SeaweedFS Merged 2025-09-13. Checked 2026-09-30.
  12. Sentry, self-hosted#3946: remove the worker and cron containers Merged 2025-09-19. Checked 2026-09-30.
  13. Sentry, RedisSnubaTSDB backend getsentry/sentry, added 2019-02-28, last modified 2026-08-25. Checked 2026-09-30.
  14. Sentry, node store backends getsentry/sentry. Checked 2026-09-30.
  15. Sentry, issue search backends getsentry/sentry. Checked 2026-09-30.
  16. Sentry, tag store backends getsentry/sentry. Checked 2026-09-30.
  17. Sentry, sentry#100326: remove more celery related code, part three Merged 2025-09-26. Checked 2026-09-30.
  18. Sentry, Snuba README getsentry/snuba, repository begins 2018-03-01. Checked 2026-09-30.
  19. Sentry, Snuba architecture overview getsentry/snuba. Checked 2026-09-30.
  20. Sentry, Snuba ClickHouse migrations getsentry/snuba, 252 migration files. Checked 2026-09-30.
  21. Sentry, snuba#7925: liveness health check scoped to essential clusters (inc-2141) Merged 2026-05-14. Checked 2026-09-30.
  22. Sentry, snuba#7945: drop messages with out-of-range timestamps (INC-2179) Merged 2026-05-19. Checked 2026-09-30.
  23. Sentry, snuba#8002: tier 1 killswitch, force downsample (inc-2173) Merged 2026-06-08. Checked 2026-09-30.
  24. Sentry, Snuba commit history getsentry/snuba, source of the observed incident identifier range. Checked 2026-09-30.
  25. Sentry, Arroyo README getsentry/arroyo, repository begins 2021-06-11. Checked 2026-09-30.
  26. Sentry, arroyo#237: honour join timeout when the pool is overloaded (INC-378) Merged 2023-05-22. Checked 2026-09-30.
  27. Sentry, arroyo#425: do not fail the consumer when the DLQ produce fails (inc-1013) Merged 2025-01-24. Checked 2026-09-30.
  28. Sentry, Relay README getsentry/relay, repository begins 2018-01-17. Checked 2026-09-30.
  29. Sentry, relay#1355: do not drop event in PoP relays (INC-181) Merged 2022-07-28. Checked 2026-09-30.
  30. Sentry, Taskbroker README getsentry/taskbroker, repository begins 2024-11-01. Checked 2026-09-30.
  31. PyPI, sentry 371 releases, 2012-01-04 to 2023-07-25. Checked 2026-09-30.
  32. PyPI, sentry-sdk 351 releases from 2018-07-26. Checked 2026-09-30.
  33. PyPI, sentry-relay 93 releases from 2020-01-27. Checked 2026-09-30.
  34. PyPI, sentry-arroyo 205 releases from 2021-06-29. Checked 2026-09-30.
  35. npm, @sentry/browser 750 versions from 2017-12-07. Checked 2026-09-30.