Replacing the engine room in public  / field guide
Practitioner field guide · 4 October 2026

Replacing the engine room in public

Sentry spent a decade swapping out the component that receives data, the component that stores it and the component that runs background work, while shipping the whole thing to strangers as a Docker Compose file. This guide reconstructs those five replacements from the artefacts they were executed in, and shows where the people running the free edition held a veto over the architecture.

32 primary sources 1 production system, 2 deployment flavours 5 incident records Evidence through October 2026 Read: 16 min
01

The territory

One company, one repository, ten years, and a second audience that runs the same software on its own hardware.

7 → 51
Containers in the shipped deployment, 2019 tag to 2026 master
2.4 → 14 GB
Documented memory floor for a self-hosted install
111,734
Commits on one master branch, still one repository
20
Public RFCs closed without being merged

Strip the product name out and the problem is this: you operate a service, you also give the same software away for other people to run, and over ten years you need to replace the part that receives the data, the part that stores it and the part that executes background work. Every replacement has to run beside the thing it replaces for a while. Every replacement also has to be packaged, documented and supported for operators you have never met, on hardware you cannot see, who upgrade monthly at best. The question is not which architecture is better. It is which replacement you can actually ship, and what each one costs the second audience.

Sentry is a useful subject for that question because the artefact is the architecture. The product that runs at sentry.io is built from the same public repository that an operator clones, so each internal change leaves a trace in something external: a new service in the Compose file, a new hard requirement in the installer, a line in the changelog telling operators to keep the old component for now. Between the 9.1.2 tag and master, the shipped deployment went from seven containers (smtp, memcached, redis, postgres, web, cron, worker) to fifty-one, and the installer stopped being advice and became a gate: it now refuses to run below 14 GB of RAM and four CPUs, or 7 GB and two if you install the errors-only profile.

Why this now: the longest-running of the five replacements, moving asynchronous work off Celery, crossed into the shipped artefact between June 2025 and September 2025 and was still being extended in the 26.7.0 to 26.9.0 releases, so the whole arc is visible at once, including the part that is not finished.

The finding that surprised me

The free self-hosted edition is not downstream of the architecture. It is a veto on it, and the veto is exercised twice in the record. RFC 0072 rejected running a Kafka schema registry service, in part because it would have to be available "in all regions, open source, dev, CI and single tenant installations", so Sentry shipped schemas as a library in two languages instead. And the Celery replacement, announced to operators in the 25.6.0 release notes as the thing that would take over "the roles of the worker and cron containers", was followed one release later by a changelog entry reading "feat: Continue using celery in self-hosted for now".

What this guide covers: five component replacements inside one product between roughly 2016 and 2026, the decision records behind four of them, the failure reports filed by operators, and the resource floor that the replacements pushed upward. What it does not cover: anything about the scale of sentry.io itself, because no figure for events per second, storage volume or cost was reachable; the client-side SDKs, which is where most of the public decision record actually lives; and the merits of the licence change, which appears here only as a dated commit.

Figure 1 · What arrived in the shipped deployment, and when

adds the streaming pipeline

adds consumers per dataset

removes zookeeper,
cron, worker

Tag 9.1.2 · 7 containers
web, cron, worker,
postgres, redis, memcached, smtp

Tag 20.12.1 · 30 containers
zookeeper, kafka, clickhouse,
snuba, relay, symbolicator

Master, Oct 2026 · 51 containers
18 snuba consumers, pgbouncer,
object store, uptime checker

taskbroker, taskscheduler,
taskworker

adds the streaming pipeline

adds consumers per dataset

removes zookeeper,
cron, worker

Tag 9.1.2 · 7 containers
web, cron, worker,
postgres, redis, memcached, smtp

Tag 20.12.1 · 30 containers
zookeeper, kafka, clickhouse,
snuba, relay, symbolicator

Master, Oct 2026 · 51 containers
18 snuba consumers, pgbouncer,
object store, uptime checker

taskbroker, taskscheduler,
taskworker

Notice that nothing leaves. Across three snapshots of the same Compose file the only removals are Zookeeper and the Celery pair, and both were replacements rather than simplifications. Sources: 9.1.2, 20.12.1, master.
Diagram source
02

How it is actually built

The shape the five replacements converged on, each box traceable to a file in the repository rather than to a diagram in a talk.

The pipeline that exists in 2026 has a consistent shape, and the shape is the result of moving two things outward from the Python application: the hot path and the durability boundary. Receiving, normalising and rate limiting now happen in Relay, a separate Rust process that the README describes as pushing "some functionality from the Sentry SDKs as well as the Sentry server into a proxy process". In processing mode it does not call the application at all; it produces to Kafka. Search and aggregation left the relational store for ClickHouse behind a service called Snuba, whose README states plainly that it "was originally developed to replace a combination of Postgres and Redis to search and provide aggregated data on Sentry errors", and whose architecture document says Kafka is the only ingestion input it has. Asynchronous work is being moved from Celery to taskbroker, a Rust service that is, in its own words, "a Kafka consumer, RPC interface, and inflight task storage".

Kafka is therefore not a queue in this architecture. It is the contract surface between every component, which is exactly what RFC 0072 says in February 2023: "Kafka topics and their schemas are usually not internal to any one service, but part of the contract between services." That RFC exists because schema disagreements had already caused production incidents in post-processing, transaction consumption and replays. Once a broker becomes the interface, the broker's failure modes become everyone's failure modes, which is the thread running through every operator incident in section 4.

The second pattern is less obvious and worth naming, because the sources do not name it: local durability at the edges. Both of the newest Rust services carry their own embedded database. Relay gained a SQLite envelope buffer in version 24.9.0 ("Allow creation of SqliteEnvelopeBuffer from config, and load existing stacks from db on startup") and deleted the older spooling implementation in 24.11.2. Taskbroker keeps inflight tasks in SQLite so that workers can acknowledge individual tasks out of order, which its README gives as the reason it exists: avoiding head-of-line blocking. Neither document references the other, but both are the same move. When the shared broker is the only durable thing in the middle, the processes at either end need somewhere local to put work that cannot move yet.

What did not happen is the part most readers will expect. The application was never split into services. It was split by jurisdiction instead, starting on 31 August 2022 with pull requests that standardised "nomenclature around silos" under a hybrid label. The same code now runs in one of three modes, declared in a single enum whose docstring says the default "assumes that the server is the only 'silo' in its environment and allows access to all tables and endpoints". Deployment topology is described in src/sentry/types/cell.py as cells grouped into localities, "e.g. 'us' contains 'us1', 'us2'", and in monolith mode "there exists only the one monolith 'region', which is a dummy object". A reader should take two things from that file. First, data residency, not scale, is what finally broke the single deployment. Second, the monolith was retained as a synthetic special case, because the self-hosted edition has to keep working, and that decision is why a single enum value in the current code is still spelled CELL = "REGION": the concept was renamed, the wire value could not be.

Figure 2 · The 2026 pipeline, with the two embedded stores marked

envelopes

gRPC GetTask

SDKs

Relay, Rust
normalise, rate limit
SQLite envelope buffer

Kafka topics
schemas as a library

Snuba consumers

taskbroker, Rust
SQLite inflight store

ClickHouse

taskworker, Python

Postgres

Sentry app
monolith, or control + cells

envelopes

gRPC GetTask

SDKs

Relay, Rust
normalise, rate limit
SQLite envelope buffer

Kafka topics
schemas as a library

Snuba consumers

taskbroker, Rust
SQLite inflight store

ClickHouse

taskworker, Python

Postgres

Sentry app
monolith, or control + cells

The durability boundary moved to Kafka, so the two processes at the ends grew local SQLite stores to survive it. Reconstructed from Relay's changelog, taskbroker's README and Snuba's architecture overview.
Diagram source

Read the five replacements side by side and they run through the same five stages, which is the part that transfers to another organisation. The new component first appears beside the old one in the same deployment. Then a switch appears that an operator can set. Then internal traffic moves. Then the old component is dropped from the shipped artefact, which is a different event, usually much later. And at some point the documented resource floor goes up, because the dual-run period is paid for in memory by whoever installs it. The useful discipline here is to ask, for any replacement you are planning, what the observable signal is at each of the five stages, because if you cannot name the signal you cannot tell a migration that is progressing from one that has stalled with both generations running.

Figure 5 · The five stages, and the artefact that records each one

1 Beside
new service appears
in the manifest

2 Switch
an option in the
release notes

3 Internal cutover
option override removed

4 Out of the artefact
container deleted
from the manifest

5 Floor raised
commit to the
requirements file

1 Beside
new service appears
in the manifest

2 Switch
an option in the
release notes

3 Internal cutover
option override removed

4 Out of the artefact
container deleted
from the manifest

5 Floor raised
commit to the
requirements file

Stage four is the one organisations forget to schedule, and the gap between stages three and four is where an architecture is paying for two implementations. Abstracted from the five episodes in the table below.
Diagram source
Replacement episodeAppeared besideOperator switchOld component gone from the shipped manifestState in October 2026
Ingest edge moves to a Rust processRelay present by tag 20.12.1 (compose)Proxy versus processing mode (README)No dated removal of the old server-side path is in reachDone, with the buffer rebuilt in 2024
Search and aggregation move to ClickHouseSnuba and ClickHouse by tag 20.12.1None publishedNo dated commit in reach; the README states the replacementDone; 18 consumers in the manifest
Task execution leaves Celerytaskbroker in 25.6.0 (release)taskworker.enabled in 25.8.0 (release)cron and worker absent from master; override removed in 25.9.0Still porting consumers in 26.7.0 to 26.9.0
Deployment splits by jurisdictionSilo modes from 31 August 2022 (pull requests)Deployment mode, with monolith as defaultNever; monolith is retained as "a dummy object"Permanently dual, by design
Message contracts become an artefactSchema library, RFC 0072, February 2023Not applicableThe registry service was never builtDone; the rejected option stays rejected

The edge process

Rust, deployed separately, and the only part of the pipeline a customer's traffic touches first. Its job grew from proxying to normalising, rate limiting, sampling and buffering to disk.

Evidence: Relay README, changelog 24.8.0 to 24.11.2

The columnar store and its service

ClickHouse is never addressed directly by the application. Snuba owns the schemas, the consumers and a query language, which is what made swapping the store under the product survivable.

Evidence: Snuba README, architecture overview

The task plane

Kafka for arrival, a Rust broker for inflight state, Python workers over gRPC. The design constraint is stated as head-of-line blocking, not throughput.

Evidence: taskbroker README, self-hosted changelog

03

The decisions that matter

Four decisions where the rejected option and its stated reason are both on the record, plus the condition that would flip each one.

Decision: how do services agree on the shape of a Kafka message?

Chosen
  • A schema library published for Python and Rust, with schemas as code artefacts and a get_schema(topic) call (RFC 0072, approved February 2023)
  • No new runtime dependency anywhere
Rejected
  • A schema registry service, Confluent's or otherwise
  • Stated reason: infrastructure overhead, a network dependency, and keeping it available "in all regions, open source, dev, CI and single tenant installations"
Flips when
  • You do not ship your platform to third parties, so a registry has one deployment to live in rather than five
  • Or your producers outgrow a library release cycle, since a library couples schema changes to deploys of every language runtime

Decision: how does Rust get into a large Python application?

Chosen
  • One Sentry-specific bindings package in its own repository, built with PyO3 and maturin (RFC 0119, October 2023)
  • Reasons given: "Only a single Python extension module to build / care about with fixed overhead" and the "ability to move functionality from Python to Rust more fine-grained"
Rejected
  • Per-crate Python wheels published to PyPI
  • Stated reason: it would "still require maintaining a public SemVer API" and multiple extension modules each carry fixed overhead
Flips when
  • The Rust you are writing has users outside your company, at which point the public API you were avoiding is the product

Decision: what executes background work?

Chosen
  • Kafka for arrival plus taskbroker, a Rust service holding inflight tasks in SQLite and serving workers over gRPC
  • Stated purpose: avoid head-of-line blocking, allow out-of-order execution, acknowledge tasks individually
Rejected
  • Celery on Redis, described in the 25.6.0 release notes as what taskbroker "aims to replace"
  • RFC 0002 had already flagged the same area in 2022, including moving off pickle and off RabbitMQ
Flips when
  • Task durations are uniform. Head-of-line blocking is a symptom of mixed-duration work sharing a prefetch window, and a queue per duration class is cheaper than a broker

Decision: how do you serve a jurisdiction without splitting the product?

Chosen
  • One codebase with three deployment modes, and a topology of cells grouped into localities such as 'us' and 'de'
  • Monolith mode retained as "a dummy object" so a single install still works
Rejected
  • A separate build or fork per region, and equally a split into independently deployed services
  • Nothing in the public record proposes microservices for the application itself
Flips when
  • You have no residency requirement, in which case the silo decorators and the RPC layer between control and region are cost with no benefit

Figure 3 · The test that decided each of these

yes

no

no

yes

A component must be replaced

Is it on the hot path
or the durability boundary?

Must it also run in
every self-hosted install?

Leave it; spend the
budget on the pipeline

Build the service
e.g. a registry, a control plane

Ship it as a library
or a single container,
and gate the install

Raise the documented
resource floor

yes

no

no

yes

A component must be replaced

Is it on the hot path
or the durability boundary?

Must it also run in
every self-hosted install?

Leave it; spend the
budget on the pipeline

Build the service
e.g. a registry, a control plane

Ship it as a library
or a single container,
and gate the install

Raise the documented
resource floor

Every decision in the table below resolves through the same two questions, and the second one is the one most organisations do not have to ask. Derived from RFC 0072 and the self-hosted changelog.
Diagram source
DecisionChosenRejectedBecauseEvidence
Where normalisation and rate limiting runA separate Rust process at the edgeKeeping it in the Python serverMoves per-event work off the application and lets the edge drop traffic before it costs anything downstreamRelay README
What backs search and aggregationClickHouse behind SnubaPostgres plus Redis"a good balance of the real time performance Snuba needs, its distributed and replicated nature"Snuba overview
Kafka message contractsA schema library in two languagesA schema registry serviceIt would have to run "in all regions, open source, dev, CI and single tenant installations"RFC 0072
Rust inside PythonOne private bindings packagePer-crate public wheelsAvoids maintaining "a public SemVer API" and per-module overheadRFC 0119
Background executionKafka plus a Rust broker with per-task acksCelery on RedisHead-of-line blocking and all-or-nothing acknowledgementtaskbroker README
Serving a jurisdictionControl silo plus cells, one codebaseA fork or build per regionKeeps one deployable unit; monolith mode preserved for single installstypes/cell.py
Client-side custom metricsAbandonedLocal aggregation in every SDK"We can close this as we haven't moved forward with metrics in the SDKs"RFC PR 115

The last row is the one worth sitting with. A proposal to put metric aggregation into every SDK was opened in October 2023, discussed for a year, and closed unmerged in November 2024 with a single sentence. Twenty of Sentry's public RFCs ended that way. An architect reading a decision record repository should read the closed-unmerged list first, because an approved RFC tells you what a company intended and a withdrawn one tells you what it could not sustain.

04

What broke in production

Four operator-filed incident reports and one release advisory. Sentry's own postmortems are published off GitHub and were unreachable for this session, so what follows is the operator's view of this architecture, which is the only view in reach.

The reports sort into two classes, and the classes are more useful than the individual incidents. The first is coordination: once Kafka became the contract surface, a broker that is slow to answer takes out components that have nothing to do with each other. The second is silence: in all four reports the operator's first signal is an absence of data, not an error, and in two of them the product's own statistics page is what eventually shows the loss. That is the characteristic failure of a pipeline whose durability boundary sits in the middle rather than at the ends.

Figure 4 · How a broker timeout becomes an empty dashboard

Sentry UIConsumerKafkaRelaySDKSentry UIConsumerKafkaRelaySDKenvelopes queue in thelocal SQLite bufferdashboard simplystops updatingenvelope accepted, 200 OKproduceMessageTimedOutpollnothing newquery latest events
Sentry UIConsumerKafkaRelaySDKSentry UIConsumerKafkaRelaySDKenvelopes queue in thelocal SQLite bufferdashboard simplystops updatingenvelope accepted, 200 OKproduceMessageTimedOutpollnothing newquery latest events
The reporter in issue 4240 sees nothing wrong with the application; the failure is a produce timeout three hops earlier, and the only visible symptom is that new events stop appearing. Source: self-hosted issue 4240, March 2026.
Diagram source
Postmortem

The edge accepts traffic it cannot hand on

AssumptionIf the edge process is healthy and returns 200, the data is safe.
What happenedRelay logged "failed to produce message to Kafka (delivery callback) error=Message production error: MessageTimedOut" across multiple topics, and processing stopped after the stack had been up for a while.
Blast radiusAll new events and metrics for that installation. Opened 24 March 2026, still open and labelled "Waiting for: Product Owner" when checked on 4 October 2026.
FixNone published. The local envelope buffer added in Relay 24.9.0 bounds the data loss but does not surface it.
Design ruleA process that accepts work on behalf of a downstream it cannot reach must export the depth and age of its local buffer as a first-class metric. Acceptance without observable backlog is how silent loss happens.
Postmortem

One coordinator, eight unrelated consumers

AssumptionIndependent consumers fail independently.
What happened"Consumer group session timed out (in join-state steady) after 45000 ms without a successful response from the group coordinator", with NOT_COORDINATOR and COORDINATOR_LOAD_IN_PROGRESS errors, across eight consumer groups at once, including subscriptions, segment processing, post-process forwarding, monitors and metrics.
Blast radiusRepeated unhealthy restarts and processing delays. Opened 21 August 2026, open and unanswered by a maintainer when checked.
FixNone published. A related taskbroker change in release 26.8.0 raised kafka_session_timeout_ms, which treats the symptom.
Design ruleConsumer-group coordination is a shared failure domain even when the topics are not. If you count blast radius by topic, you will under-count it; count it by coordinator.
Postmortem

A version bump leaves consumers reading offsets that no longer exist

AssumptionA monthly upgrade of a packaged distribution is a container-image change.
What happenedAfter moving from 22.11.0 to 22.12.0, Snuba consumers crash-looped on "fetch failed due to requested offset not available on the broker: Broker: Offset out of range (broker 1001)".
Blast radiusIngestion stopped for the affected installations; the thread ran from 2 January 2023 and was eventually closed.
FixOperator-side offset resets. The structural answer arrived later as the versioned Kafka schema library proposed in RFC 0072.
Design ruleAny upgrade that moves consumer code must state what it assumes about retained offsets, and the installer should check it. Calendar versioning tells an operator when a release was cut, not what it will do to committed state.
Postmortem

Forty-nine thousand transactions, all counted, all dropped

AssumptionIf the pipeline is counting events, it is storing them.
What happened"Performance page shows zeros for the time period since the update and until now" while the "Stats page shows 49k transactions of which 49k are dropped"; the ClickHouse container logged "Net Exception: Socket is not connected".
Blast radiusEvery transaction for the reporting installation, from an upgrade onward. Opened 10 March 2024 and closed as not planned, with no root cause recorded.
FixNone published.
Design ruleSeparate the accepted counter from the stored counter and show both. A single throughput number cannot distinguish a quiet week from total loss, and the gap between the two counters is the only cheap end-to-end check a pipeline of this shape has.
The release channel as a rollback mechanism

The fifth record is not an incident report but a release note. Version 25.6.2 tells operators: "If you are coming from 25.5.1, you might want to skip 25.6.0 and 25.6.1, and upgrade directly to this version", after a Postgres migration failure. For software other people install, a published skip-list is the cheapest rollback story available, and it is one most teams shipping an on-premises artefact never write down.

05

Numbers you can plan against

Everything quantitative that was reachable, with the date it was true and the file it came from.

MetricValueWhereContextAs ofSource
Containers in the shipped deployment7self-hostedsmtp, memcached, redis, postgres, web, cron, workertag 9.1.2compose
Containers in the shipped deployment30self-hostedZookeeper, Kafka, ClickHouse, Snuba, Relay, Symbolicator presenttag 20.12.1compose
Containers in the shipped deployment51self-hosted18 Snuba consumers, taskbroker trio, pgbouncer, seaweedfs; no Zookeeper2026-10-04compose
Documented memory floor2,400 MBself-hostedStated in the README as advice, not enforcedtag 20.12.1README
Enforced memory floor14,000 MBself-hostedMIN_RAM_HARD, installer refuses below it2026-10-04_min-requirements.sh
Enforced CPU floor4 coresself-hosted2 cores for the errors-only profile2026-10-04_min-requirements.sh
Errors-only memory floor7,000 MBself-hostedContributor's measurement: "fewer resources are required. About 2 times"2025-03-27PR 3634
Date the floor became a gate2024-08-17self-hostedCommit "Mandate minimum requirements for ram/cpu (#3275)", relaxed six days later2024-08file history
Start of the topology split2022-08-31getsentry/sentryFirst silo-mode pull requests, labelled hybrid2022-08PR search
Relicensing dates2023-11-17, 2024-02-27getsentry/sentryFSL-1.0-Apache-2.0, then FSL-1.1, four months apart2024-02LICENSE.md history
Commits on master, single repository111,734getsentry/sentryThe application was never split into repositories2026-10-04repository page
Public RFCs accepted into the text directory71 filesgetsentry/rfcsNumbered 0001 to 0157, mostly client-side concerns2026-10-04text directory
Public RFCs closed unmerged20getsentry/rfcsIncludes custom metrics, combined dynamic sampling, per-category rate limiting2026-10-04closed unmerged
Read these carefully

Every row above is measured from a file in the repository, which makes them reliable about the shipped artefact and silent about the production service: no figure for events per second, storage volume, cost or latency at sentry.io was reachable in this session, and none should be inferred from these numbers. The container counts are mine, derived by counting service keys in three versions of one Compose file, and they count processes rather than hosts. The one genuinely independent measurement here is the resource comparison in pull request 3634, and it is a contributor's screenshot of a running stack, contested in the thread by a reviewer who pointed out that "Kafka and Clickhouse would still take a lot of resources".

06

The evidence wall

Every source behind this page, graded. Filter by kind. The blog, talk, paper and case-study tiers are empty because every host carrying them was blocked for this session; that absence is the biggest limitation of this guide.

Decision record Sentry2022-07

RFC 0002: Sentry Architecture Vision (Living Document)

Opens by naming a scaling limit as the trigger for the decade's architecture work, then lists the workstreams that followed: extracting Celery tasks, leaving pickle, moving from RabbitMQ to Kafka, and phasing out buffers in favour of Kafka and ClickHouse.

Carry forwardA living architecture document that names the constraint, not the target state, is the one that still reads correctly four years later.
github.com/getsentry/rfcs · text/0002-new-architecture.md
Decision record Sentry2023-02

RFC 0072: Centralized Schema Repository for Kafka Topics

Records that schema disagreements had already caused production incidents, chooses a schema library for Python and Rust, and rejects a registry service because of what it would cost to run in every deployment flavour including open source.

Carry forwardWhen a broker becomes the contract surface, the schema question arrives with it; answer it in the artefact your smallest deployment can carry.
github.com/getsentry/rfcs · text/0072-kafka-schema-registry.md
Decision record Sentry2023-10

RFC 0119: Make it easier to use Rust code from Sentry/Python

Chooses one private bindings package built with PyO3 and maturin over per-crate public wheels, with the rejection resting on the cost of maintaining a public SemVer API.

Carry forwardThe cheapest way to adopt a second language inside a monolith is one internal extension module, because the expensive part is the public interface you did not need.
github.com/getsentry/rfcs · text/0119-rust-in-sentry.md
Source Sentry2026-10

getsentry/rfcs, closed-unmerged pull requests

Twenty withdrawn proposals, including custom metrics in SDKs, combined dynamic sampling and per-category abuse rate limiting. The gaps in the numbering of the accepted set are visible here.

Carry forwardRead the withdrawn RFCs before the accepted ones: they show what an organisation tried and could not sustain.
github.com/getsentry/rfcs · closed unmerged
Source Sentry2024-11

RFC pull request 115: Custom Metrics in SDKs

Opened October 2023, closed unmerged thirteen months later with "We can close this as we haven't moved forward with metrics in the SDKs". The proposal required every SDK to aggregate locally over ten-second windows.

Carry forwardAn ingestion path that depends on work inside every client library is the most expensive kind to add and the hardest to withdraw.
github.com/getsentry/rfcs · pull 115
Source Sentry2026-10

getsentry/rfcs text directory

Seventy-one files numbered up to 0157. The large majority concern SDKs, spans, replays and symbolication rather than server architecture, which is where the four decisions in section 3 had to be found.

Carry forwardA public RFC set is shaped by who needs to be convinced; server decisions get made where the arguing happens, which may not be the RFC repository.
github.com/getsentry/rfcs · text
Source Sentry2026-10

getsentry/relay README

Describes the edge process as pushing functionality out of both the SDKs and the server, and documents that in processing mode it produces to Kafka instead of forwarding to an upstream Sentry.

Carry forwardPushing per-event work to a separate edge binary also moves your rate limiting to the only place that can enforce it for free.
github.com/getsentry/relay · README.md
Source Sentry2024-11

Relay changelog, versions 24.8.0 to 24.11.2

Tracks an experimental envelope buffer becoming a SQLite-backed store loaded on startup, then the deletion of the previous spooling implementation and its configuration options.

Carry forwardAn edge that accepts data on behalf of an unavailable broker needs durable local storage; expect to build it twice before the shape is right.
github.com/getsentry/relay · CHANGELOG.md
Source Sentry2026-10

getsentry/snuba README

States in one sentence that the service was built to replace Postgres and Redis for search and aggregation over errors, which dates and motivates the move to a columnar store.

Carry forwardPut a service in front of the new store before you migrate onto it; the indirection is what makes the next store swap possible.
github.com/getsentry/snuba · README.rst
Source Sentry2026-10

Snuba architecture overview

Gives the reason ClickHouse was chosen, states that Kafka topics are the only ingestion input, and notes that no table is written by more than one consumer.

Carry forwardOne writer per table is the constraint that keeps a columnar ingestion pipeline debuggable; it is also what forces a consumer per dataset.
github.com/getsentry/snuba · docs/source/architecture/overview.rst
Source Sentry2026-10

getsentry/taskbroker README

Names the three design purposes of the Celery replacement: no head-of-line blocking, out-of-order execution, per-task acknowledgement, with inflight state in SQLite and workers served over gRPC.

Carry forwardIf your queue problem is described as head-of-line blocking, the fix is per-task acknowledgement, not more workers.
github.com/getsentry/taskbroker · README.md
Source Sentry2025-06

self-hosted release 25.6.0

Announces taskbroker to operators as the service that "aims to replace Celery" and says the next release will have it take over the worker and cron containers.

Carry forwardAnnounce a replacement to your operators one release before you depend on it, and expect to need that slack.
github.com/getsentry/self-hosted · releases/tag/25.6.0
Source Sentry2026-09

self-hosted CHANGELOG

Carries the whole arc of the task migration in four lines across fifteen months, from adding taskbroker, to "Continue using celery in self-hosted for now", to removing the option override, to porting individual consumers to tasks in 2026.

Carry forwardA changelog is the only honest record of how long a replacement really took, because it has to tell operators the truth about what is still dual-running.
github.com/getsentry/self-hosted · CHANGELOG.md
Source Sentry2025-08

self-hosted release 25.8.0

Tells operators to set taskworker.enabled to false if they want their jobs to keep running on Celery, which is the dual-run switch stated as an operator-facing option.

Carry forwardGive the old path an explicit off switch that an operator can set, and the migration stops being a release-day gamble.
github.com/getsentry/self-hosted · releases/tag/25.8.0
Source Sentrytag 9.1.2

docker-compose.yml at tag 9.1.2

Seven services. The entire product shipped as a web process, a cron, a Celery worker, Postgres, Redis, memcached and an SMTP relay.

Carry forwardKeep a copy of your deployment manifest from five years ago; it is the cheapest measure of what your architecture has cost your operators.
github.com/getsentry/self-hosted · 9.1.2/docker-compose.yml
Source Sentrytag 20.12.1

docker-compose.yml at tag 20.12.1

Thirty services, including Zookeeper, Kafka, ClickHouse, the first Snuba consumers, Symbolicator and Relay, alongside the cron and worker pair that were still there.

Carry forwardThe dual-run period shows up as both generations sitting in the same manifest, and that snapshot is what dates the migration.
github.com/getsentry/self-hosted · 20.12.1/docker-compose.yml
Source Sentry2026-10

docker-compose.yml on master

Fifty-one services. Zookeeper, cron and worker are gone; eighteen Snuba consumers, the taskbroker trio, pgbouncer, an object store and an uptime checker have arrived.

Carry forwardCount the services in your shipped manifest once a year. It is the number your smallest customer experiences as your architecture.
github.com/getsentry/self-hosted · docker-compose.yml
Source Sentrytag 20.12.1

self-hosted README at tag 20.12.1

States the memory requirement of the era in six words: at least 2400 MB of RAM, with Docker and Compose version minimums alongside it.

Carry forwardThe requirements line in an old README is a dated, unarguable measurement of architectural weight.
github.com/getsentry/self-hosted · 20.12.1/README.md
Source Sentry2026-10

install/_min-requirements.sh

The current floor as code: 14,000 MB and four CPUs by default, 7,000 MB and two CPUs under the errors-only profile, with a comment reminding the author to update the docs when the values change.

Carry forwardPut the resource floor in a file the installer reads, not in prose. Prose drifts and nobody notices.
github.com/getsentry/self-hosted · install/_min-requirements.sh
Source Sentry2025-08

Commit history of install/_min-requirements.sh

Nine commits across four years. The floor became mandatory on 17 August 2024, was relaxed six days later, and acquired a reduced profile in March 2025.

Carry forwardThe history of the requirements file dates every step-change in architectural cost, and a relaxation six days after a tightening tells you the first number was wrong.
github.com/getsentry/self-hosted · commits for install/_min-requirements.sh
Source community contributor2025-03

Pull request 3634: Minimum requirements for the errors-only profile

A contributor measures the reduced profile at roughly half the resources, a reviewer objects that Kafka and ClickHouse still dominate, and the numbers go in after screenshots of a running stack.

Carry forwardIf your product has a reduced mode, publish its floor separately; a single number makes the whole architecture look heavier than the part most users need.
github.com/getsentry/self-hosted · pull 3634
Source Sentry2026-10

install/check-minimum-requirements.sh

The installer refuses to continue below the floor and separately requires SSE 4.2 on x86_64 for ClickHouse unless the check is explicitly skipped.

Carry forwardAn instruction-set requirement is the kind of dependency a storage choice adds to your install base without appearing anywhere in the architecture diagram.
github.com/getsentry/self-hosted · install/check-minimum-requirements.sh
Source Sentry2026-10

src/sentry/silo/base.py

The enum that lets one codebase act as a monolith, a control silo or a cell, with the monolith documented as the default that "allows access to all tables and endpoints". The member reads CELL = "REGION".

Carry forwardWhen you rename a deployment concept, the old name survives in the values; plan for the wire format to outlive the vocabulary.
github.com/getsentry/sentry · src/sentry/silo/base.py
Source Sentry2026-10

src/sentry/types/cell.py

Defines cells hosted by a region silo and localities grouping them, with 'us' containing 'us1' and 'us2', and keeps the monolith alive as a synthetic region that is "a dummy object".

Carry forwardA residency boundary is cheapest to add as a naming layer above your existing deployment unit, with the single-instance case preserved as a degenerate one.
github.com/getsentry/sentry · src/sentry/types/cell.py
Source Sentry2022-08

Oldest pull requests mentioning silo modes

Work starts on 31 August 2022 with CI jobs for silo modes and a change that standardises "nomenclature around silos", both labelled hybrid.

Carry forwardA topology change begins in the test harness. If CI cannot run both topologies, the migration has not started.
github.com/getsentry/sentry · silo pull requests, oldest first
Source Sentry2024-02

Commit history of LICENSE.md

Relicensed under FSL-1.0-Apache-2.0 in November 2023 and upgraded to FSL-1.1 in February 2024, in the middle of the topology and task work.

Carry forwardLicence terms are part of the architecture of a shipped product: they set who may run the thing whose resource floor you keep raising.
github.com/getsentry/sentry · commits for LICENSE.md
Source Sentry2026-10

getsentry/sentry repository

111,734 commits on master and 45.3k stars, for a product that grew five new runtime components without ever splitting its application repository.

Carry forwardComponent replacement and repository splitting are independent decisions, and this is the case study for doing the first without the second.
github.com/getsentry/sentry
Postmortem self-hosted operator2026-03

Issue 4240: Relay stopped processing new metrics and events

Produce timeouts to Kafka across multiple topics, with ingestion stopping after the stack has been running for a while. Open and labelled as waiting on a product owner when checked.

Carry forwardExport local buffer depth and age from any process that accepts data on behalf of a broker it cannot reach.
github.com/getsentry/self-hosted · issues/4240
Postmortem self-hosted operator2026-08

Issue 4485: consumers repeatedly unhealthy on coordinator timeouts

Eight unrelated consumer groups time out against the group coordinator at once, with NOT_COORDINATOR and COORDINATOR_LOAD_IN_PROGRESS in the logs.

Carry forwardCount blast radius by coordinator, not by topic: consumer-group coordination is shared even when the data is not.
github.com/getsentry/self-hosted · issues/4485
Postmortem self-hosted operator2023-01

Issue 1894: Kafka error after update to 22.12.0 from 22.11.0

Snuba consumers crash-loop on offsets that are no longer available on the broker after a one-month version bump.

Carry forwardState what an upgrade assumes about committed consumer state, and check it in the installer rather than in the release notes.
github.com/getsentry/self-hosted · issues/1894
Postmortem self-hosted operator2024-03

Issue 2876: Sentry stopped accepting transaction data

Every transaction counted and dropped, with ClickHouse socket errors in the logs. Closed as not planned with no root cause recorded.

Carry forwardShow accepted and stored as two counters. The gap between them is the only cheap end-to-end check this pipeline shape has.
github.com/getsentry/self-hosted · issues/2876
Postmortem Sentry2025-06

self-hosted release 25.6.2

Advises operators coming from 25.5.1 to skip two releases entirely and upgrade straight to this one, after a Postgres migration failure.

Carry forwardFor software other people install, a published skip-list is a rollback mechanism, and it costs one line of release notes.
github.com/getsentry/self-hosted · releases/tag/25.6.2
07

Build a miniature, then productionise it

Six rungs that reproduce the five replacements at a scale you can hold in your head. The crossing from toy to real is rung four.

An edge that accepts what it cannot forward

Write a small ingest process in front of your application that validates and forwards to a broker, with a local embedded store for anything it cannot hand on yet.

Done when: you stop the broker, keep posting, restart the edge process, and nothing posted during the outage is missing.  Teaches: why both of Sentry's newest Rust services carry a SQLite file.

Make the loss visible

Add two counters, accepted and stored, and a dashboard panel showing the difference. Then break the store and watch the panel rather than the logs.

Done when: the gap appears within one scrape interval of a failure you injected.  Teaches: the failure mode behind issues 2876 and 4240, where the operator's only signal was an absence.

Put a service in front of the store

Move every query behind a small service that owns the schema and the query language, then swap the underlying store for a different one without touching the application.

Done when: the application's code is unchanged across the swap.  Teaches: what Snuba is for, and why the store underneath it can change.

Dual-run a replacement behind an operator-facing switch

Introduce a second implementation of your task execution, controlled by one setting, and run both in the same deployment with traffic split by task type.

Done when: a third party can flip the setting back with no code change and no data migration.  Teaches: why the changelog line reads "Continue using celery in self-hosted for now" rather than a rollback.

Publish the message contracts as a versioned artefact

Extract your broker schemas into a library, in each language that produces or consumes, and make CI fail on an incompatible change.

Done when: a breaking schema change cannot be merged without a version bump that the other language sees.  Teaches: the trade RFC 0072 made against a registry service.

Ship it to a stranger, with a floor

Package the whole thing as a single manifest with an installer that measures CPU, memory and any instruction-set requirement, and refuses to proceed below a number you commit to a file.

Done when: an operator on an undersized machine gets a clear refusal instead of a slow, broken install.  Teaches: that every component you add is a bill somebody else pays, and that the file is where you find out how big it is.

08

Keep hunting

The queries and file paths that produced this page. The first group works on any company that ships its own product; the second needs only a repository.

The record of a shipped architecture

  • repo path: install/_min-requirements.sh and its commit history
  • same docker-compose.yml at three tags, five years apart, counted
  • CHANGELOG: grep for remove, migrate, continue using, breaking
  • release notes: "you might want to skip" OR "upgrade directly to"

Decisions and the arguments that lost

  • is:pr is:closed is:unmerged in the RFC repository, read every one
  • is:pr "silo" OR "region" OR "hybrid" in:title sort:created-asc
  • RFC text: "alternatives considered" plus "open source" or "single tenant"
  • commits for LICENSE, LICENSE.md and any file named NOTICE

Incidents, when the vendor publishes none you can reach

  • issues: "stopped accepting" OR "stopped processing" OR "shows zeros"
  • issues: "Offset out of range" OR "session timed out" OR "not coordinator"
  • issues: label "Waiting for: Product Owner", sorted by comments
  • issues: "after update to" OR "after upgrading to" plus a version number

Where this page should be extended first

  • the vendor's own engineering blog, for the scale figures missing here
  • conference talks by the authors named in the RFCs
  • the vendor's status page history, for incidents on the paid service
  • job postings for the team that owns the component being replaced

What this page would most benefit from, and could not get: any first-person account from the engineers who ran these migrations, and any figure at all for the volume the pipeline handles. Both exist in public. Neither was reachable from this session, and a reader with ordinary network access should start there.

The lesson to carry into your own design

Sentry replaced the receiving layer, the storage layer and the execution layer of a ten-year-old product without ever splitting the product, and the thing that governed the pace was not scale or language choice. It was the second audience. Every component that had to exist in a stranger's install was shipped as a library, a single container or an option with an off switch, and every component that did not, such as the rejected schema registry, was simply never built. If your software runs anywhere you do not operate, treat the packaged artefact as a first-class constraint on the architecture rather than an output of it: write down the deployable-unit count and the resource floor, review them as you would a latency budget, and schedule the removal of the old component as its own dated piece of work, because nothing else in your process will.

09

References

  1. Sentry, RFC 0002: Sentry Architecture Vision (Living Document) getsentry/rfcs, 21 July 2022. Checked 2026-10-04.
  2. Sentry, RFC 0072: Centralized Schema Repository for Kafka Topics getsentry/rfcs, 1 February 2023. Checked 2026-10-04.
  3. Sentry, RFC 0119: Make it easier to use Rust code from Sentry/Python getsentry/rfcs, 25 October 2023. Checked 2026-10-04.
  4. Sentry, closed-unmerged RFC pull requests getsentry/rfcs. Checked 2026-10-04.
  5. Sentry, RFC pull request 115: Custom Metrics in SDKs getsentry/rfcs, opened 3 October 2023, closed 12 November 2024. Checked 2026-10-04.
  6. Sentry, accepted RFC text directory getsentry/rfcs. Checked 2026-10-04.
  7. Sentry, Relay README getsentry/relay. Checked 2026-10-04.
  8. Sentry, Relay CHANGELOG getsentry/relay, versions 24.8.0 to 26.7.0. Checked 2026-10-04.
  9. Sentry, Snuba README getsentry/snuba. Checked 2026-10-04.
  10. Sentry, Snuba architecture overview getsentry/snuba. Checked 2026-10-04.
  11. Sentry, taskbroker README getsentry/taskbroker. Checked 2026-10-04.
  12. Sentry, self-hosted release 25.6.0 getsentry/self-hosted, June 2025. Checked 2026-10-04.
  13. Sentry, self-hosted release 25.6.2 getsentry/self-hosted, June 2025. Checked 2026-10-04.
  14. Sentry, self-hosted release 25.8.0 getsentry/self-hosted, August 2025. Checked 2026-10-04.
  15. Sentry, self-hosted CHANGELOG getsentry/self-hosted. Checked 2026-10-04.
  16. Sentry, self-hosted docker-compose.yml at tag 9.1.2 getsentry/self-hosted. Checked 2026-10-04.
  17. Sentry, self-hosted docker-compose.yml at tag 20.12.1 getsentry/self-hosted, December 2020. Checked 2026-10-04.
  18. Sentry, self-hosted docker-compose.yml on master getsentry/self-hosted. Checked 2026-10-04.
  19. Sentry, self-hosted README at tag 20.12.1 getsentry/self-hosted, December 2020. Checked 2026-10-04.
  20. Sentry, install/_min-requirements.sh getsentry/self-hosted. Checked 2026-10-04.
  21. Sentry, commit history of install/_min-requirements.sh getsentry/self-hosted, 2021 to 2025. Checked 2026-10-04.
  22. Sentry, install/check-minimum-requirements.sh getsentry/self-hosted. Checked 2026-10-04.
  23. Community contributor, pull request 3634: Minimum requirements for the errors-only profile getsentry/self-hosted, merged 27 March 2025. Checked 2026-10-04.
  24. Sentry, src/sentry/silo/base.py getsentry/sentry. Checked 2026-10-04.
  25. Sentry, src/sentry/types/cell.py getsentry/sentry. Checked 2026-10-04.
  26. Sentry, pull requests mentioning silo, oldest first getsentry/sentry, from 31 August 2022. Checked 2026-10-04.
  27. Sentry, commit history of LICENSE.md getsentry/sentry, November 2023 and February 2024. Checked 2026-10-04.
  28. Sentry, getsentry/sentry repository GitHub. Checked 2026-10-04.
  29. Operator report, issue 4240: Relay stopped processing new metrics and events getsentry/self-hosted, opened 24 March 2026. Checked 2026-10-04.
  30. Operator report, issue 4485: consumers repeatedly unhealthy on coordinator timeouts getsentry/self-hosted, opened 21 August 2026. Checked 2026-10-04.
  31. Operator report, issue 1894: Kafka error after update to 22.12.0 from 22.11.0 getsentry/self-hosted, opened 2 January 2023. Checked 2026-10-04.
  32. Operator report, issue 2876: Sentry stopped accepting transaction data getsentry/self-hosted, opened 10 March 2024. Checked 2026-10-04.

Evidence ledger with one row per claim, including the quote supporting each, ships beside this page as sources.md.