Scaling without splitting  / field guide
Practitioner field guide · 3 October 2026

Scaling without splitting: a decade of Shopify, read from the code it published

Between 2014 and 2026 Shopify kept one Rails application as its unit of deployment and changed almost everything underneath it instead: boundaries moved into a static analyser, resilience into the database drivers, tenant movement into a purpose-built MySQL copier, and performance into two compilers written for the language itself. This guide reconstructs that decade from the repository record, twenty-six graded sources carrying a ledger of thirty-seven claims, and gives an architect the conditions under which each of those choices stops being the right one.

26 graded sources 11 repositories 4 production defect reports Evidence through October 2026 Read: 30 min
01

The territory

One codebase, thousands of contributors, ten years, and no permission to cut it into independently deployed pieces. What has to change instead, and what it costs.

Oct 2014
Circuit breakers and fault injection running in production and in every test environment
2020-09-23
Boundary enforcement extracted from the monolith as a static analyser
2025-04-18
A second JIT compiler upstreamed into CRuby, still crashing most benchmarks
0
Published incident reviews reachable in this corpus

State the problem without naming a technology and it stops being a Rails question. An organisation runs one program. The program is the product, it is deployed as a unit, and the number of people changing it grows faster than the program does. The standard answer for the last fifteen years has been to cut the program into services, which converts a code-structure problem into a network and coordination problem. Shopify declined that trade for its core application and has now lived with the consequences for a decade. The public record of that decade is unusually complete, because almost every mechanism the company built to survive it was extracted and published as its own repository.

That record is what this guide reads. Not the monolith, which is closed, but the ring of artefacts around it: a boundary checker, a circuit breaker, a fault-injection proxy, a live data-movement engine with a formal specification, a deployment tool, two compilers, and a WebAssembly toolchain for code the company does not trust. Each one is a load-bearing part of the answer to the question above, and each one carries, in its own documentation, the argument for why it exists.

The surprise

The company that chose static analysis over network boundaries then removed half of its own boundary rules. Packwerk shipped in 2020 with two checks: dependency direction and constant visibility. In November 2022 the visibility check was pulled out of the core gem entirely and moved to a separate project, and the change shipped as the 3.0.0 major version on 1 March 2023. What stayed in the core is the one rule that can be enforced without knowing anything about intent: which package may reference which. The rule that encoded design opinion left. For a reader planning a modular monolith, that is the more useful fact than any success story, because it says which half of the idea survived contact with a codebase that size.

Three boundaries on the scope. This guide covers the decade from the first Semian commit on 23 September 2014 to the repository state on 3 October 2026. It does not cover the monolith's internal structure, its traffic, or its cost, because none of that is public in a form this session could reach. It makes no comparison with peer companies: that would need their engineering blogs, which were unreachable here, and a comparison built on memory is worth nothing. Where the guide reconstructs rather than reports, it says so in the sentence.

Figure 1 · Where the company intervened

Runtime and data

One deployable

Change time

packwerk
boundary check

Toxiproxy
failure injection

The Rails monolith
(not public)

Semian
in-driver breakers

IdentityCache
read-through cache

Ghostferry
live shard moves

YJIT then ZJIT
compilers in CRuby

krane
deploy verdict

Runtime and data

One deployable

Change time

packwerk
boundary check

Toxiproxy
failure injection

The Rails monolith
(not public)

Semian
in-driver breakers

IdentityCache
read-through cache

Ghostferry
live shard moves

YJIT then ZJIT
compilers in CRuby

krane
deploy verdict

Every published artefact sits below the application rather than beside it: the monolith stays one deployable and the layers under it absorb the change. Reconstructed from the repositories cited in the evidence wall.
Diagram source
02

How it is actually built

The shape that recurs across the published artefacts, and the four places where the design diverges from what a service-oriented answer would have produced.

The architecture that falls out of these repositories has a consistent logic. The unit of deployment is the application. The unit of isolation is the tenant's data, held in a MySQL shard that can be moved while the application is serving. The unit of design is the package, which exists only in the autoloader and in a checker that runs in continuous integration. The unit of failure containment is the connection, guarded inside the driver rather than at a network hop. Nothing in that list is a service, and every one of them is a boundary.

Start with the data plane, because it is where the physical partitioning lives. Ghostferry is a Go library that copies rows out of one MySQL instance and into another while writes continue, by doing two things at once: a SELECT ... FOR UPDATE iterator that copies batches, and a binlog streamer that replays source changes onto the target. Its technical overview states the hard part plainly: "Ghostferry mandates that you stop writes to the dataset you are copying at a stage of execution called cutover", and the stopping of writes happens outside the tool. The repository also ships a sharding package with a ShardedCopyFilter, which is the mechanism by which a subset of rows belonging to one tenant, rather than a whole table, is the thing that moves. A 2020 pull request adding target-side verification was sent for review to a team handle, @Shopify/pods, which is the only public naming of the group that operates this.

The resilience plane is the most unusual choice in the set, and the most deliberate. Semian does not sit in a sidecar or a mesh. It monkey-patches the MySQL, Redis, Postgres and HTTP drivers in the application process, and coordinates concurrency limits across every worker on a host using SysV semaphores. The README explains why a circuit breaker on its own is not enough in a thread-per-request server: with a ten second timeout, "it will still take at least 10s before the circuit is open. In that time every worker is blocked", so "we're at reduced capacity for at least 20s". The bulkhead is sized from an occupancy estimate rather than a guess: if monitoring says there is "a 10% chance they're talking to mysql_shard_0 at any given point in time", then five workers doing so simultaneously has probability 0.001%, so the resource gets five tickets and the sixth worker fails instantly. That is a design where the false-positive rate is chosen, written down, and traded against capacity.

The boundary plane is static. Packwerk reads constant references out of the autoloaded codebase and reports a violation when a package references a constant it has not declared a dependency on. Its own documentation is honest that this is a narrow instrument: "Method calls and objects passed around the application are completely ignored. Packwerk only cares about static constant references", and it is designed "to avoid false positives ... at any cost", accepting false negatives in exchange. Existing violations are not errors. They are recorded in a package_todo.yml file per package, and the usage guide warns that running the command that records them "to resolve a violation should be the very last resort".

The runtime plane is where the strategy becomes expensive. Rather than rewrite hot paths in another language, Shopify built a just-in-time compiler inside CRuby. YJIT, written in Rust, "lazily compiles code using a Basic Block Versioning (BBV) architecture" and ships in upstream Ruby; the company keeps a public fork at Shopify/ruby that both JIT documents name as the place to file bugs and argue about patches. In April 2025 a second compiler, ZJIT, was upstreamed: "a method-based just-in-time (JIT) compiler for Ruby" that "uses profile information from the interpreter to guide optimization". Two compilers with different architectures, maintained by one company for one application, is the clearest statement in this corpus of where the performance budget went.

Figure 2 · The reference shape of a kept monolith

ticket or fail fast

SELECT FOR UPDATE

binlog replay

cutover: writes stopped outside the tool

Request

Monolith worker
packages checked in CI

Driver wrapped
by Semian

Tenant shard
MySQL

Memcached
IdentityCache

Ghostferry

Target shard

ticket or fail fast

SELECT FOR UPDATE

binlog replay

cutover: writes stopped outside the tool

Request

Monolith worker
packages checked in CI

Driver wrapped
by Semian

Tenant shard
MySQL

Memcached
IdentityCache

Ghostferry

Target shard

The request never crosses a service boundary, so every boundary it does cross is either a check that ran earlier (packwerk, in CI) or a guard inside the process (Semian, in the driver). Reconstructed from Semian, packwerk and Ghostferry.
Diagram source

What is shared

One deployment artefact, one autoloader, one connection-pool policy per host, one set of drivers. Sharing these is what makes the in-process mechanisms possible: a semaphore can coordinate workers only because they are the same program.

What is partitioned

Tenant data, by shard, with a copier that moves a filtered subset of rows rather than a table. The partition is in the data layer only; no code is partitioned with it, which is why a shard move is a data problem and not a deployment problem.

What is pushed downward

Performance work lands in the Ruby VM, framework work lands in Rails upstream, and sandboxing for third-party code lands in WebAssembly. The application keeps its shape and the layers beneath it absorb the change.

Figure 5 · What got published, and when

2013-2014go-lua, a sandboxedinterpreterSemian and Toxiproxyin production2017-2018krane, deploy verdictsGhostferry, live shardmoves with a TLA+model2020-2021packwerk extractedfrom the monolithJavy, JavaScript toWebAssembly2022-2023visibility checksremoved frompackwerkSarama handed toIBM, Javy to theBytecode Alliance2025-2026ZJIT upstreamed intoCRubyVM heap sizing fixedupstream fromproduction evidenceTen years of published mechanisms
2013-2014go-lua, a sandboxedinterpreterSemian and Toxiproxyin production2017-2018krane, deploy verdictsGhostferry, live shardmoves with a TLA+model2020-2021packwerk extractedfrom the monolithJavy, JavaScript toWebAssembly2022-2023visibility checksremoved frompackwerkSarama handed toIBM, Javy to theBytecode Alliance2025-2026ZJIT upstreamed intoCRubyVM heap sizing fixedupstream fromproduction evidenceTen years of published mechanisms
Resilience came first and the compilers came last, which is the order the constraints arrived in: capacity, then boundaries, then the cost of the runtime itself. Dates are first commits, merges and release tags from the repositories cited in the references.
Diagram source
03

The decisions that matter

Each fork, the option that lost, the stated reason, and the condition under which the loser becomes the right answer.

Decision: how do you stop a large codebase turning into a ball of mud?

Chosen
  • A static checker over constant references, run in CI, with existing violations recorded in a per-package todo file
  • Stated reason: "Ruby does not provide a good solution to enforcing boundaries between code"
Rejected
  • Extraction into separately deployed services
  • Runtime enforcement: packwerk ignores method calls and runtime object graphs entirely
Flips when
  • You need independent deployment, independent failure or independent scaling, none of which a checker can give you
  • Or when the todo file stops shrinking, at which point you are paying for the check and getting no boundary

Decision: where does the circuit breaker live?

Chosen
  • Inside the process, by patching each driver, with host-wide bulkheads in SysV semaphores
  • Failure is always an exception that inherits from the driver's own base error, so existing rescues keep working
Rejected
  • Tuning timeouts, which the README dismisses: the alternatives "all still revolved around timeouts, and those are extremely hard to get right"
  • An out-of-process proxy, which cannot see how many of this host's workers are blocked
Flips when
  • The fleet is polyglot: a per-language monkey-patch has to be rebuilt per language, which is the argument for a mesh
  • Or when workers are threads with cheap blocking, which removes the capacity argument the design rests on

Decision: how do you move a tenant's rows to another database with the shop still open?

Chosen
  • A purpose-built copier: batched SELECT ... FOR UPDATE plus binlog replay, with a formal model committed beside the code
  • Driven by the move to managed MySQL, where "we cannot access the filesystems of the cloud databases and we had trouble with setting up replication via proprietary interfaces"
Rejected
  • mysqldump and Percona XtraBackup plus replication, which need filesystem access or hold a long transaction
  • Whole-table granularity, which cannot express "this tenant"
Flips when
  • You own the filesystem and can take the dataset read-only for the length of a restore, which makes the standard tools correct and far cheaper
  • Or when your tables have foreign keys, non-integer primary keys, or need schema changes mid-run, all of which the tool refuses

Decision: where do you buy performance when the application is already written?

Chosen
  • Build a compiler for the language, twice: YJIT with basic block versioning in 2021, then ZJIT, method-based and profile-guided, upstreamed 18 April 2025
  • Feed production measurements upstream: a January 2023 Rails change was driven by dumping object shapes "from Shopify's monolith production environment"
Rejected
  • Rewriting hot paths in another runtime, which would split the deployable the strategy exists to keep whole
  • Staying on tuning: a 2026 pull request describes heap tuning through environment variables as working "OK but is high maintaince as it needs to be retuned frequently"
Flips when
  • You cannot staff compiler work for a decade, which almost nobody can; the transferable move is the measurement, not the compiler
  • Or when the hot path is one service's worth of code, which is exactly when extraction is cheaper than a JIT

Figure 3 · Which boundary can you actually afford?

yes

no

yes

no

yes

no

Do two parts need to
deploy independently?

Extract a service
and accept the network

Do they need to
fail independently?

In-process bulkhead
plus breaker per resource

Do they need to
scale on different data?

Partition the data
and build the mover

Static boundary check
in CI with a todo file

yes

no

yes

no

yes

no

Do two parts need to
deploy independently?

Extract a service
and accept the network

Do they need to
fail independently?

In-process bulkhead
plus breaker per resource

Do they need to
scale on different data?

Partition the data
and build the mover

Static boundary check
in CI with a todo file

The decision is not modular monolith against microservices; it is which property you need, because only one of these four boundaries gives independent deployment. Derived from the stated reasons in packwerk's usage guide and Semian's README.
Diagram source
DecisionChosenRejectedBecauseEvidence
Boundary enforcementStatic constant analysis in CIService extraction; runtime checksThe language offers no boundary primitive; false positives were deemed costlier than false negativesUSAGE.md
Which boundary rules to keepDependency direction onlyConstant visibility, moved out in PR 247Visibility encodes design opinion; it shipped as a separate gem from v3.0.0 on 2023-03-01PR 247
Overload protectionIn-driver breaker plus semaphore bulkheadTimeout tuning; external proxyA blocked worker is lost capacity, and timeouts alone leave 20s of degradationSemian
Resilience testingA TCP proxy with programmable faults in CILinux tooling such as ncExisting tools are not cross-platform and "require root, which makes them problematic in test, development and CI environments"Toxiproxy
Tenant movementCustom copier with binlog replaymysqldump, XtraBackup, native replicationManaged MySQL hides the filesystem and the replication interfacePercona Live 2018
Migration verificationInline verifier plus a target-side provenance checkIterativeVerifier, deprecated; CHECKSUM TABLE at large sizesChecksums cost cutover time linear in data size; the inline verifier costs copy time insteadverifiers.md
Runtime performanceTwo JIT compilers in CRuby, in RustRewriting hot paths outside RubyKeeping one deployable is the constraint; the compiler is the variableruby/ruby PR 13131
Untrusted extension codeJavaScript compiled to WebAssembly, 1 to 16 KB per module with dynamic linkingIn-process scripting in the application runtimeA sandbox with a size budget can be run per request; the project moved to the Bytecode Alliance in 2023Javy
Generic infrastructureHand it over: Sarama to IBM (2023-07-17), Javy to the Bytecode Alliance, the benchmark suite to RubyContinuing to maintain it aloneNothing in these is specific to the monolith, so the maintenance is pure costsarama CHANGELOG
04

What broke in production

Shopify publishes no incident reviews that this session could reach. What it does publish is better than nothing and worse than a postmortem: engineer-written defect reports in the repositories, with mechanism, detection and fix.

Read the defect record for the data-movement engine and a taxonomy appears that transfers to any system copying rows between stores while both are live. Three of the four classes are silent by construction, which is the point: an engine that moves a tenant's data cannot fail loudly, because the only thing that would notice is the comparison it is skipping.

Figure 4 · How a type mismatch becomes silent data loss

Inline verifierTarget MySQLBinlog writerSource MySQLInline verifierTarget MySQLBinlog writerSource MySQLJSON compared to stringnever matchesmissed when the pagination keyis not leftmost in a composite PKUPDATE row, JSON columnchangedUPDATE ... WHERE data = '{}'(string)0 rows matched, no errorre-read and compare fingerprintmismatch found, runfails
Inline verifierTarget MySQLBinlog writerSource MySQLInline verifierTarget MySQLBinlog writerSource MySQLJSON compared to stringnever matchesmissed when the pagination keyis not leftmost in a composite PKUPDATE row, JSON columnchangedUPDATE ... WHERE data = '{}'(string)0 rows matched, no errorre-read and compare fingerprintmismatch found, runfails
The replay statement succeeds, affects zero rows, and the run continues: nothing in the path raises. Reconstructed from Ghostferry PR 147.
Diagram source
Defect report

JSON columns silently swallow every update and delete

AssumptionA row can be matched on the target by comparing each column to the value read from the source.
What happenedMySQL never matches a JSON column against a string literal, so the generated UPDATE and DELETE statements matched nothing and returned success.
Blast radiusAny table with a JSON column, on every binlog event replayed. Caught at cutover by the verifiers, except where the pagination key is not the leftmost column of a composite primary key, in which case corruption occurs "WITHOUT the InlineVerifier notifying".
FixWrap the comparison value in CAST(... AS JSON) in every generated statement.
Design ruleEnumerate column types, not tables. Any type whose equality semantics differ from byte equality is a silent-corruption source, and the list is finite: write it down and test each entry.
Defect report

Binlog file 999999 compares greater than 1000000

AssumptionComparing two binlog positions can be done on the filename, because filenames sort the way the numbers do.
What happenedThe upstream library compared filenames as strings, so a target position of mysql-bin.1000000 appeared to be behind mysql-bin.999999 and the binlog streamer terminated early, losing every change after that point.
Blast radiusOnly on instances that have rotated past a million binlog files, which is why it survived for years. The report notes it is "a very very rare condition, but a serious one as no online verifiers will be able to catch this kind of data loss".
FixUpgrade to the library version whose comparison parses the sequence number, after demonstrating the defect with a six-line program in the pull request.
Design ruleA verifier that shares the mover's model of the world cannot catch an error in that model. Budget for one check that is independent of the copier, or accept that an entire failure class is undetectable.
Defect report

An optimisation reordered the result columns

AssumptionRows are read and written in the column order declared by the table schema, so the writer can pair values to columns positionally.
What happenedA rewrite of the shard filter's SELECT introduced a sub-select, and USING(FallbackColumn) moved that column to the front of the result set, so values were written into the wrong columns whenever pagination used a non-primary key.
Blast radiusShard moves paginating on a fallback column. "The inline verifier caught this and crashed Ghostferry correctly", which is the system working as designed.
FixCarry the column order in the row batch instead of inferring it, so query shape and write shape cannot drift apart.
Design ruleAn invariant that lives in two places, here the query and the writer, is not an invariant. Move it into the value that travels between them.
Production report

Boot allocations set the heap size for the rest of the process's life

AssumptionWhat a process allocates while starting is a reasonable predictor of what it needs while serving.
What happenedA monolith's boot sequence is allocation heavy, so the VM's allocatable bytes grew far past the steady-state need, "GC not triggering for a long time, with the adverse effect of a unreasonably high memory usage and infrequent but long GC pauses to sweep dozens of requests worth of garbage".
Blast radiusEvery worker in the fleet, continuously, with the symptom showing up as memory and tail latency rather than as errors. Measured in the patch: 372.1 MiB of allocatable bytes after a synthetic boot, against 57.1 MiB with the fix.
FixRecompute the heap limits at the end of boot, replacing a workaround that tuned environment variables and needed "to be retuned frequently as the application shape changes".
Design ruleAny autotuner that samples during startup is sampling the wrong distribution. Give it an explicit end-of-warmup signal, or it will optimise for a phase that lasts seconds and bill you for it for days.
What the absence of postmortems costs you

Every failure above was found by a developer reading code or a verifier failing a run. None of them tells you what the mechanism did to availability when it broke at scale, how long detection took, or what the operators saw first. For a reader evaluating this architecture, that is precisely the missing evidence: the public record shows that the in-process breaker, the shard mover and the boundary checker exist and are maintained, and shows nothing about the days they did not hold. Design your own rollout to produce that evidence, because nobody else has published it.

05

Numbers you can plan against

Everything quantitative in the corpus, with the context it was measured in and the date it was true.

QuantityValueContextAs ofSource
Cutover downtime for a live moveOrder of seconds; minutes if verification runs in the windowShopify's own runs, presenter notes2018-04Percona Live talk
Copy parallelism in a production run4 parallel table iteratorsScreenshot of a run's control UI2018-04Percona Live talk
Bulkhead tickets per resource5, from a 10% occupancy estimate and an accepted 0.001% false-positive rateWorked example in the library's own READMErepo at 2026-09Semian
Capacity loss from a slow dependency without bulkheadsAt least 20 seconds of degradation on a 10 second timeoutThread-per-request workers, worked examplerepo at 2026-09Semian
Heap allocatable bytes after an allocation-heavy boot372.1 MiB, against 57.1 MiB with the proposed fixSynthetic reproduction of the production behaviour2026-08ruby/ruby PR 18433
Top object-shape edge count in the monolith770 for a single instance variable, out of a 50-entry list all above 140Dump from the production environment2023-01rails/rails PR 47023
WebAssembly module size for sandboxed extension code1 to 16 KB dynamically linked; at least 869 KB statically linkedToolchain documentationrepo at 2026-10Javy
Time from boundary-rule removal to released major versionMerged 2022-11-14, shipped in v3.0.0 on 2023-03-01Repository history and release tag2023-03v3.0.0
Age of the deployment toolFirst commit 2017-01-17, renamed 2019-10-28, still changing 2026-09-29Repository history2026-09krane
Interval between the two JIT compilersYJIT merged upstream in 2021; ZJIT upstreamed 2025-04-18, unable to run most benchmarks at mergeUpstream pull request2025-04ruby/ruby PR 13131
Read these carefully

Measured: the heap figures, the shape counts and the module sizes are measurements published with the code that produced them. Self-reported: the cutover window and the copy parallelism come from a company talk about its own tooling, with no independent replication. Derived: the bulkhead's five tickets follow from an occupancy estimate the reader must make for their own system, and the arithmetic is only as good as that estimate. Unknown: how many shards exist, how long a real tenant move takes end to end, what any of this costs, and how often the cutover window is missed. No reachable source gives a figure for any of them, so plan the first move to measure them yourself.

06

The evidence wall

Every source behind this page, graded. Filter by kind. The ledger with one row per claim, including the quotes, ships beside this file as sources.md.

Decision record Shopifyrepo 2026-08

packwerk USAGE.md

States the problem the tool exists for, the enforcement model, strict mode, and the semantics of the per-package todo file that holds existing violations.

Carry forwardA boundary you cannot enforce today is a backlog with a filename; make the backlog visible and make adding to it an explicit act.
github.com/Shopify/packwerk/blob/main/USAGE.md
Source Shopify2020-09-23

packwerk, first commit: extraction from the Shopify codebase

The tool entered public life as an extraction, which dates the point at which the modular-monolith strategy became a shipped mechanism rather than an intention.

Carry forwardAn extraction commit is the most reliable date stamp a strategy has; look for it before trusting a narrative.
github.com/Shopify/packwerk/commit/1489086
Source Shopify2022-10-31

PR 247: pull privacy concerns out of packwerk

Removes the constant-visibility checker from the core gem, sequences it behind a deprecation release, and states the user-visible cost of the break.

Carry forwardWhen a rule needs intent to evaluate, it belongs in an optional extension, not in the gate everyone runs.
github.com/Shopify/packwerk/pull/247
Source Shopify2023-03-01

packwerk v3.0.0

The release that made the narrower boundary model the default for every user of the gem, four months after the code change merged.

Carry forwardMeasure a tooling decision from its release date, not its merge date; users feel the second one.
github.com/Shopify/packwerk/releases/tag/v3.0.0
Source Shopifyrepo 2026-09

Semian README

The full argument for in-process resilience: driver patching, SysV semaphore bulkheads, the ticket arithmetic, and an explicit warning that the library will surprise operators who have not read it.

Carry forwardSize a bulkhead from an occupancy measurement and state the false-positive rate you are buying; a breaker alone cannot protect a blocking worker pool.
github.com/Shopify/semian/blob/main/README.md
Source Shopifyrepo 2026-08

Toxiproxy README

A programmable TCP proxy used in every development and test environment since October 2014, with the stated goal of proving the absence of single points of failure in tests.

Carry forwardResilience you cannot inject in CI is resilience you are asserting, not testing.
github.com/Shopify/toxiproxy/blob/main/README.md
Decision record Shopifyrepo 2026-09

Ghostferry technical overview

The algorithm, the seven-step cutover, and three hard preconditions: integer primary keys, full row-based replication, and no foreign keys.

Carry forwardWrite the preconditions of a migration tool as refusals in the tool, not as warnings in a runbook.
github.com/Shopify/ghostferry/blob/main/docs/technicaloverview.md
Decision record Shopifyrepo 2026-09

Ghostferry verifiers

A matrix of inconsistency classes against verifiers, with cells reading "Sometimes" and "Probably not", a deprecated verifier, and a note that a MySQL checksum statement is broken for JSON columns.

Carry forwardPublish what your verification does not catch; a verifier built on the mover's own model cannot check that model.
github.com/Shopify/ghostferry/blob/main/docs/verifiers.md
Talk Shopify2018-04-24

Ghostferry at Percona Live, slides with presenter notes

Why the tool exists (managed MySQL hides the filesystem and the replication interface), what TLA+ found before release, and the only published downtime figures.

Carry forwardFormal modelling pays for itself on silent-corruption paths, and its own authors should be the ones saying a finite model is not a proof.
github.com/Shopify/ghostferry/blob/main/docs/_static/percona-talk.pdf
Source Shopifyrepo 2026-09

Ghostferry TLA+ specification

The model is committed next to the Go implementation rather than written once and discarded, and the README admits "proofs remain elusive".

Carry forwardA specification in the repository ages with the code; one in a document does not.
github.com/Shopify/ghostferry/blob/main/tlaplus/ghostferry.tla
Defect report Shopify2020-01-17

PR 147: data corruption on tables with JSON columns

Updates and deletes matched nothing because MySQL does not compare a JSON column to a string, and one verifier configuration cannot detect it.

Carry forwardEnumerate the column types whose equality is not byte equality, and test the mover against each one.
github.com/Shopify/ghostferry/pull/147
Defect report Shopify2021-09-21

PR 307: rare data loss from binlog position comparison

String comparison of binlog filenames terminates the stream early past a million files, in a way no online verifier can catch, demonstrated with a short program.

Carry forwardIdentifiers that look numeric are compared as text by default; the failure waits for the digit count to change.
github.com/Shopify/ghostferry/pull/307
Defect report Shopify2021-06-07

PR 287: corruption when paginating a shard move on a fallback column

A query optimisation changed result-column order and the writer kept pairing values positionally; the inline verifier stopped the run.

Carry forwardPass the column order with the batch. An invariant held in two places is a future corruption.
github.com/Shopify/ghostferry/pull/287
Source Shopify2020-04-21

PR 176: target verification against data corruption

Annotates every statement the mover issues and fails the run on any unannotated write to the target, with the review sent to the @Shopify/pods team.

Carry forwardDuring a migration, provenance is a safety property: anything writing to the target that is not you is corruption.
github.com/Shopify/ghostferry/pull/176
Source Shopifyrepo 2026-09

Ghostferry sharding filter

The shard-aware copy filter, which makes a tenant rather than a table the unit that moves between databases.

Carry forwardIf tenants are your partition, the mover needs a filter, not a table name.
github.com/Shopify/ghostferry/blob/main/sharding/filter.go
Source Ruby core, Shopify2026-10

doc/jit/yjit.md

The first compiler: basic block versioning, written in Rust inside CRuby, with Shopify as the contact address and a company fork named as the discussion venue.

Carry forwardOwning a compiler is a staffing commitment, not a project; read the contact address to see who is actually carrying it.
github.com/ruby/ruby/blob/master/doc/jit/yjit.md
Source Ruby core, Shopify2025-04-18

ruby/ruby PR 13131: ZJIT

The second compiler is upstreamed from a then-private company repository, with side exits and invalidation unimplemented and most benchmarks crashing.

Carry forwardUpstreaming early buys review and compatibility pressure at the cost of shipping something visibly unfinished; say which you need.
github.com/ruby/ruby/pull/13131
Source Ruby core, Shopify2026-10

doc/jit/zjit.md

Method-based compilation guided by interpreter profiles, a different architecture from the first compiler, with bug reports routed to the company fork.

Carry forwardA second system with a different architecture is evidence the first one hit a structural limit, not a performance one.
github.com/ruby/ruby/blob/master/doc/jit/zjit.md
Source Shopify, in rails/rails2023-01-16

rails/rails PR 47023: improve Rails' shape friendliness

Production object-shape data from the monolith is dumped, ranked, and used to justify changes to the framework itself rather than to application code.

Carry forwardThe cheapest fix for a monolith's performance is often upstream, and the evidence that wins the argument is your own production dump.
github.com/rails/rails/pull/47023
Production report Shopify, in ruby/ruby2026-08-22

ruby/ruby PR 18433: recompute allocatable bytes after warmup

A production memory and pause-time problem traced to boot-time allocation shaping the heap, with the previous workaround described as high maintenance.

Carry forwardAutotuners need an end-of-warmup signal; without one they fit the startup phase and bill you during steady state.
github.com/ruby/ruby/pull/18433
Source Shopifyrepo 2026-09

krane README and rename commit

A deployment wrapper whose purpose is a verdict: it runs kubectl underneath and answers "what just happened" and "did it work" for CI.

Carry forwardIf your deploy tool cannot produce a pass or fail, your pipeline is reporting on itself rather than on the system.
github.com/Shopify/krane/blob/main/README.md
Source IBM, formerly Shopify2023-07-17

sarama CHANGELOG, v1.40.0

The release note recording the transfer of a widely used Go Kafka client from Shopify to IBM, with the module path change as a breaking change for every dependent.

Carry forwardHanding over generic infrastructure is a legitimate end state; the import path is where your users pay for it.
github.com/IBM/sarama/blob/main/CHANGELOG.md
Source Bytecode Alliance, formerly Shopify2023-05-16

Javy README and its handover commit

The JavaScript-to-WebAssembly toolchain started in April 2021 and relabelled a Bytecode Alliance project in May 2023, with module sizes that decide the linking model.

Carry forwardA sandbox's per-invocation size budget is an architectural constraint; dynamic linking is what makes per-request execution affordable.
github.com/bytecodealliance/javy/blob/main/README.md
Source Shopifyrepo 2026

IdentityCache and maintenance_tasks

Two more extractions from the same application: an opt-in read-through cache in the ORM invalidated on commit, and an engine that makes data backfills pausable and resumable.

Carry forwardAt monolith scale, backfills and cache invalidation are product surfaces with operators, not scripts.
github.com/Shopify/maintenance_tasks/blob/main/README.md
Source Shopify, Ruby core2026-10

The company fork, and the benchmark suite it gave back

Both JIT documents route contributions through Shopify/ruby, while the benchmark suite built to justify the first compiler now lives as ruby/ruby-bench and benchmarks any implementation.

Carry forwardKeep the fork where you argue, hand over the measurement everyone needs; the split tells you which part was ever company-specific.
github.com/ruby/ruby-bench
Vendor Google Cloud2026

Shopify case study

Marketing material, and the only reachable statement of where the platform runs: Kubernetes, Bigtable, BigQuery and Compute Engine, with an AI assistant as the newest layer on top.

Carry forwardVendor pages establish that something exists and nothing about how it behaves; use them for the stack list and for nothing else.
cloud.google.com/customers/shopify
07

Build a miniature, then productionise it

Six rungs. The line between a toy and the real thing is rung four, where you stop testing that the mechanism works and start measuring what it hides.

Put a boundary check in front of your own codebase

Pick any application with more than one team in it. Declare three packages, generate the violation list, and commit it.

Done when: the check runs in CI and the todo file has a number in it that you can quote.  Teaches: that your real boundaries are the ones the backlog measures, not the ones in the diagram.

Make the backlog shrink, or admit it will not

Set one package to strict mode so no new violations can be recorded against it. Run it for two weeks of normal work.

Done when: either the file shrinks or you have a written reason why that package cannot hold the line.  Teaches: whether a static boundary is a constraint in your organisation or a report.

Fail a dependency on purpose, in a test

Put a programmable proxy between the application and its slowest datastore, add one second of latency, and assert what the request does.

Done when: a test fails because the application waits instead of degrading.  Teaches: how much of your capacity one slow dependency owns.

Size a bulkhead from a measurement

Measure the fraction of workers concurrently in that datastore over a normal day. Compute the ticket count that keeps the false-positive rate where you want it, and enforce it.

Done when: you can state the accepted false-positive rate as a number and show the measurement behind it.  Teaches: that overload protection is a deliberately chosen error budget, not a safety net.

Move a live dataset and try to corrupt it

Copy a table between two databases under write load with binlog replay, then seed the hard cases: a JSON column, a composite primary key whose pagination key is not leftmost, and a write to the target from outside the copier.

Done when: you have at least one corruption your verifier does not catch.  Teaches: which of your failure classes are invisible, which is the only number that matters before a real tenant moves.

Take one measurement from production into the layer below you

Profile the running application, find a cost that belongs to the runtime or the framework rather than to your code, and open the upstream change with the production data attached.

Done when: the upstream discussion is arguing about your numbers.  Teaches: the actual mechanism behind this entire strategy, which is that a shared layer improves for everyone only when somebody brings evidence.

The lesson

Keeping the monolith is not a decision to do less work. It is a decision about where the work goes, and the public record of this decade says it goes downward: into the compiler, the drivers, the data-movement engine and the build. That trade is worth making when your bottleneck is coordination rather than deployment, because every one of those layers is shared and each improvement lands everywhere at once. The bill arrives as correctness risk in places that used to be somebody else's problem. Your boundaries become advisory, so they need a backlog that shrinks. Your resilience lives inside the driver, so it needs a false-positive rate you chose on purpose. Your partition lives in the data, so moving a tenant becomes a silent-corruption problem that needs a verifier independent of the mover. Before you decide not to split, price the verifier, not the migration.

08

Keep hunting

This page was built with no access to engineering blogs, papers, talks or issue trackers. These are the queries that worked anyway, which makes them worth keeping.

Read a company's decade out of its repositories

  • git clone --filter=blob:none https://github.com/ORG/REPO
  • git log --reverse --date=short --format='%ad %s' | head -1
  • git log --date=short --format='%ad %h %s' -S"TERM" -- README.md
  • git log --date=short --format='%ad %s' --all | grep -i "deprecat\|rename\|extract"

Find the argument, not the announcement

  • repo:ORG/REPO privacy in:title
  • repo:ORG/REPO in:title corruption OR mismatch OR "data loss"
  • repo:UPSTREAM/REPO author:ENGINEER "in production" in:body
  • org:ORG "caused an incident" in:body

Documentation that was never meant to be marketing

  • raw.githubusercontent.com/ORG/REPO/BRANCH/docs/verifiers.md
  • find . -name '*.pdf' -o -name '*.tla'
  • repo:ORG/REPO filename:DESIGN.md

What to try when the blogs are unreachable

  • grep -rn -i "postmortem\|outage\|incident" */README.md */docs/*.md
  • git log -1 --date=short --format='%ad' TAG
  • repo:ORG/REPO sort:created-asc TERM in:title

Two notes on method, both learned the hard way on this page. A conference deck committed into a documentation directory is a talk you can read offline, with presenter notes, years after the video host has stopped serving it: look for *.pdf in every repository before concluding the talks are gone. And a pull request that fixes a defect is usually a better account than the release note that announces the fix, because the author is explaining the mechanism to a reviewer rather than to a user.

09

References

  1. Shopify, packwerk README GitHub, repository state 2026-08-25. Checked 2026-10-03.
  2. Shopify, packwerk USAGE.md GitHub, repository state 2026-08-25. Checked 2026-10-03.
  3. Shopify, Extraction of Packwerk from the Shopify codebase GitHub, 2020-09-23. Checked 2026-10-03.
  4. Shopify, PR 247: Pull privacy concerns out of packwerk GitHub, 2022-10-31 to 2022-11-14. Checked 2026-10-03.
  5. Shopify, packwerk v3.0.0 GitHub, 2023-03-01. Checked 2026-10-03.
  6. Shopify, Semian README GitHub, repository state 2026-09-07. Checked 2026-10-03.
  7. Shopify, Toxiproxy README GitHub, repository state 2026-08-25. Checked 2026-10-03.
  8. Shopify, Ghostferry README GitHub, repository state 2026-09-24. Checked 2026-10-03.
  9. Shopify, Ghostferry technical overview GitHub, repository state 2026-09-24. Checked 2026-10-03.
  10. Shopify, Ghostferry verifiers GitHub, repository state 2026-09-24. Checked 2026-10-03.
  11. Shuhao Wu, Ghostferry: the swiss army knife of live data migrations with minimum downtime Percona Live, 2018-04-24, slides and presenter notes committed to the repository. Checked 2026-10-03.
  12. Shopify, Ghostferry TLA+ specification GitHub, repository state 2026-09-24. Checked 2026-10-03.
  13. Shopify, Ghostferry sharded copy filter GitHub, repository state 2026-09-24. Checked 2026-10-03.
  14. Shopify, PR 147: fix possible data corruption on tables with JSON columns GitHub, 2020-01-17. Checked 2026-10-03.
  15. Shopify, PR 176: add target verification against data corruption GitHub, 2020-04-21. Checked 2026-10-03.
  16. Shopify, PR 287: corruption when using FallbackColumn in sharding GitHub, 2021-06-07. Checked 2026-10-03.
  17. Shopify, PR 307: update go-mysql to address a very rare data corruption problem GitHub, 2021-09-21. Checked 2026-10-03.
  18. Ruby core, YJIT documentation GitHub, master at 2026-10-03. Checked 2026-10-03.
  19. Ruby core, ZJIT documentation GitHub, master at 2026-10-03. Checked 2026-10-03.
  20. Ruby core, PR 13131: ZJIT GitHub, 2025-04-18. Checked 2026-10-03.
  21. Ruby core, PR 18433: Process.warmup recompute allocatable_bytes GitHub, 2026-08-22. Checked 2026-10-03.
  22. Rails, PR 47023: improve Rails' shape friendliness GitHub, 2023-01-16. Checked 2026-10-03.
  23. Shopify, the company's public Ruby fork GitHub, checked 2026-10-03.
  24. Ruby core, ruby-bench, formerly Shopify's yjit-bench GitHub, checked 2026-10-03.
  25. Shopify, krane README GitHub, repository state 2026-09-29. Checked 2026-10-03.
  26. Shopify, rename kubernetes-deploy to krane GitHub, 2019-10-28. Checked 2026-10-03.
  27. Shopify, PR 431: removes the risk of sending decrypted EJSON secrets to output GitHub, 2019-02-27. Checked 2026-10-03.
  28. IBM, sarama CHANGELOG, v1.40.0 GitHub, 2023-07-17. Checked 2026-10-03.
  29. Go module index, github.com/IBM/sarama pkg.go.dev, checked 2026-10-03.
  30. Javy, update README to use Bytecode Alliance header GitHub, 2023-05-16. Checked 2026-10-03.
  31. Bytecode Alliance, Javy README GitHub, repository state 2026-10-02. Checked 2026-10-03.
  32. Shopify, IdentityCache README GitHub, checked 2026-10-03.
  33. Shopify, maintenance_tasks README GitHub, checked 2026-10-03.
  34. Shopify, go-lua README GitHub, repository state 2025-07-18. Checked 2026-10-03.
  35. Google Cloud, Shopify case study Vendor page, undated, checked 2026-10-03.