packwerk USAGE.md
States the problem the tool exists for, the enforcement model, strict mode, and the semantics of the per-package todo file that holds existing violations.
Between 2014 and 2026 Shopify kept one Rails application as its unit of deployment and changed almost everything underneath it instead: boundaries moved into a static analyser, resilience into the database drivers, tenant movement into a purpose-built MySQL copier, and performance into two compilers written for the language itself. This guide reconstructs that decade from the repository record, twenty-six graded sources carrying a ledger of thirty-seven claims, and gives an architect the conditions under which each of those choices stops being the right one.
One codebase, thousands of contributors, ten years, and no permission to cut it into independently deployed pieces. What has to change instead, and what it costs.
State the problem without naming a technology and it stops being a Rails question. An organisation runs one program. The program is the product, it is deployed as a unit, and the number of people changing it grows faster than the program does. The standard answer for the last fifteen years has been to cut the program into services, which converts a code-structure problem into a network and coordination problem. Shopify declined that trade for its core application and has now lived with the consequences for a decade. The public record of that decade is unusually complete, because almost every mechanism the company built to survive it was extracted and published as its own repository.
That record is what this guide reads. Not the monolith, which is closed, but the ring of artefacts around it: a boundary checker, a circuit breaker, a fault-injection proxy, a live data-movement engine with a formal specification, a deployment tool, two compilers, and a WebAssembly toolchain for code the company does not trust. Each one is a load-bearing part of the answer to the question above, and each one carries, in its own documentation, the argument for why it exists.
The company that chose static analysis over network boundaries then removed half of its own boundary rules. Packwerk shipped in 2020 with two checks: dependency direction and constant visibility. In November 2022 the visibility check was pulled out of the core gem entirely and moved to a separate project, and the change shipped as the 3.0.0 major version on 1 March 2023. What stayed in the core is the one rule that can be enforced without knowing anything about intent: which package may reference which. The rule that encoded design opinion left. For a reader planning a modular monolith, that is the more useful fact than any success story, because it says which half of the idea survived contact with a codebase that size.
Three boundaries on the scope. This guide covers the decade from the first Semian commit on 23 September 2014 to the repository state on 3 October 2026. It does not cover the monolith's internal structure, its traffic, or its cost, because none of that is public in a form this session could reach. It makes no comparison with peer companies: that would need their engineering blogs, which were unreachable here, and a comparison built on memory is worth nothing. Where the guide reconstructs rather than reports, it says so in the sentence.
The shape that recurs across the published artefacts, and the four places where the design diverges from what a service-oriented answer would have produced.
The architecture that falls out of these repositories has a consistent logic. The unit of deployment is the application. The unit of isolation is the tenant's data, held in a MySQL shard that can be moved while the application is serving. The unit of design is the package, which exists only in the autoloader and in a checker that runs in continuous integration. The unit of failure containment is the connection, guarded inside the driver rather than at a network hop. Nothing in that list is a service, and every one of them is a boundary.
Start with the data plane, because it is where the physical partitioning lives. Ghostferry
is a Go library that copies rows out of one MySQL instance and into another while writes
continue, by doing two things at once: a SELECT ... FOR UPDATE iterator that
copies batches, and a binlog streamer that replays source changes onto the target. Its
technical overview states the hard part plainly: "Ghostferry mandates that you stop writes
to the dataset you are copying at a stage of execution called cutover", and the stopping of
writes happens outside the tool. The repository also ships a sharding package
with a ShardedCopyFilter, which is the mechanism by which a subset of rows
belonging to one tenant, rather than a whole table, is the thing that moves. A 2020 pull
request adding target-side verification was sent for review to a team handle,
@Shopify/pods, which is the only public naming of the group that operates this.
The resilience plane is the most unusual choice in the set, and the most deliberate.
Semian does not sit in a sidecar or a mesh. It monkey-patches the MySQL, Redis, Postgres
and HTTP drivers in the application process, and coordinates concurrency limits across
every worker on a host using SysV semaphores. The README explains why a circuit breaker on
its own is not enough in a thread-per-request server: with a ten second timeout, "it will
still take at least 10s before the circuit is open. In that time every worker is blocked",
so "we're at reduced capacity for at least 20s". The bulkhead is sized from an occupancy
estimate rather than a guess: if monitoring says there is "a 10% chance they're talking to
mysql_shard_0 at any given point in time", then five workers doing so
simultaneously has probability 0.001%, so the resource gets five tickets and the sixth
worker fails instantly. That is a design where the false-positive rate is chosen, written
down, and traded against capacity.
The boundary plane is static. Packwerk reads constant references out of the autoloaded
codebase and reports a violation when a package references a constant it has not declared a
dependency on. Its own documentation is honest that this is a narrow instrument: "Method
calls and objects passed around the application are completely ignored. Packwerk only cares
about static constant references", and it is designed "to avoid false positives ... at any
cost", accepting false negatives in exchange. Existing violations are not errors. They are
recorded in a package_todo.yml file per package, and the usage guide warns
that running the command that records them "to resolve a violation should be the very last
resort".
The runtime plane is where the strategy becomes expensive. Rather than rewrite hot paths
in another language, Shopify built a just-in-time compiler inside CRuby. YJIT, written in
Rust, "lazily compiles code using a Basic Block Versioning (BBV) architecture" and ships in
upstream Ruby; the company keeps a public fork at Shopify/ruby that both JIT
documents name as the place to file bugs and argue about patches. In April 2025 a second
compiler, ZJIT, was upstreamed: "a method-based just-in-time (JIT) compiler for Ruby" that
"uses profile information from the interpreter to guide optimization". Two compilers with
different architectures, maintained by one company for one application, is the clearest
statement in this corpus of where the performance budget went.
One deployment artefact, one autoloader, one connection-pool policy per host, one set of drivers. Sharing these is what makes the in-process mechanisms possible: a semaphore can coordinate workers only because they are the same program.
Tenant data, by shard, with a copier that moves a filtered subset of rows rather than a table. The partition is in the data layer only; no code is partitioned with it, which is why a shard move is a data problem and not a deployment problem.
Performance work lands in the Ruby VM, framework work lands in Rails upstream, and sandboxing for third-party code lands in WebAssembly. The application keeps its shape and the layers beneath it absorb the change.
Each fork, the option that lost, the stated reason, and the condition under which the loser becomes the right answer.
SELECT ... FOR UPDATE plus binlog replay, with a formal model committed beside the code| Decision | Chosen | Rejected | Because | Evidence |
|---|---|---|---|---|
| Boundary enforcement | Static constant analysis in CI | Service extraction; runtime checks | The language offers no boundary primitive; false positives were deemed costlier than false negatives | USAGE.md |
| Which boundary rules to keep | Dependency direction only | Constant visibility, moved out in PR 247 | Visibility encodes design opinion; it shipped as a separate gem from v3.0.0 on 2023-03-01 | PR 247 |
| Overload protection | In-driver breaker plus semaphore bulkhead | Timeout tuning; external proxy | A blocked worker is lost capacity, and timeouts alone leave 20s of degradation | Semian |
| Resilience testing | A TCP proxy with programmable faults in CI | Linux tooling such as nc | Existing tools are not cross-platform and "require root, which makes them problematic in test, development and CI environments" | Toxiproxy |
| Tenant movement | Custom copier with binlog replay | mysqldump, XtraBackup, native replication | Managed MySQL hides the filesystem and the replication interface | Percona Live 2018 |
| Migration verification | Inline verifier plus a target-side provenance check | IterativeVerifier, deprecated; CHECKSUM TABLE at large sizes | Checksums cost cutover time linear in data size; the inline verifier costs copy time instead | verifiers.md |
| Runtime performance | Two JIT compilers in CRuby, in Rust | Rewriting hot paths outside Ruby | Keeping one deployable is the constraint; the compiler is the variable | ruby/ruby PR 13131 |
| Untrusted extension code | JavaScript compiled to WebAssembly, 1 to 16 KB per module with dynamic linking | In-process scripting in the application runtime | A sandbox with a size budget can be run per request; the project moved to the Bytecode Alliance in 2023 | Javy |
| Generic infrastructure | Hand it over: Sarama to IBM (2023-07-17), Javy to the Bytecode Alliance, the benchmark suite to Ruby | Continuing to maintain it alone | Nothing in these is specific to the monolith, so the maintenance is pure cost | sarama CHANGELOG |
Shopify publishes no incident reviews that this session could reach. What it does publish is better than nothing and worse than a postmortem: engineer-written defect reports in the repositories, with mechanism, detection and fix.
Read the defect record for the data-movement engine and a taxonomy appears that transfers to any system copying rows between stores while both are live. Three of the four classes are silent by construction, which is the point: an engine that moves a tenant's data cannot fail loudly, because the only thing that would notice is the comparison it is skipping.
CAST(... AS JSON) in every generated statement.mysql-bin.1000000 appeared to be behind mysql-bin.999999 and the binlog streamer terminated early, losing every change after that point.USING(FallbackColumn) moved that column to the front of the result set, so values were written into the wrong columns whenever pagination used a non-primary key.Every failure above was found by a developer reading code or a verifier failing a run. None of them tells you what the mechanism did to availability when it broke at scale, how long detection took, or what the operators saw first. For a reader evaluating this architecture, that is precisely the missing evidence: the public record shows that the in-process breaker, the shard mover and the boundary checker exist and are maintained, and shows nothing about the days they did not hold. Design your own rollout to produce that evidence, because nobody else has published it.
Everything quantitative in the corpus, with the context it was measured in and the date it was true.
| Quantity | Value | Context | As of | Source |
|---|---|---|---|---|
| Cutover downtime for a live move | Order of seconds; minutes if verification runs in the window | Shopify's own runs, presenter notes | 2018-04 | Percona Live talk |
| Copy parallelism in a production run | 4 parallel table iterators | Screenshot of a run's control UI | 2018-04 | Percona Live talk |
| Bulkhead tickets per resource | 5, from a 10% occupancy estimate and an accepted 0.001% false-positive rate | Worked example in the library's own README | repo at 2026-09 | Semian |
| Capacity loss from a slow dependency without bulkheads | At least 20 seconds of degradation on a 10 second timeout | Thread-per-request workers, worked example | repo at 2026-09 | Semian |
| Heap allocatable bytes after an allocation-heavy boot | 372.1 MiB, against 57.1 MiB with the proposed fix | Synthetic reproduction of the production behaviour | 2026-08 | ruby/ruby PR 18433 |
| Top object-shape edge count in the monolith | 770 for a single instance variable, out of a 50-entry list all above 140 | Dump from the production environment | 2023-01 | rails/rails PR 47023 |
| WebAssembly module size for sandboxed extension code | 1 to 16 KB dynamically linked; at least 869 KB statically linked | Toolchain documentation | repo at 2026-10 | Javy |
| Time from boundary-rule removal to released major version | Merged 2022-11-14, shipped in v3.0.0 on 2023-03-01 | Repository history and release tag | 2023-03 | v3.0.0 |
| Age of the deployment tool | First commit 2017-01-17, renamed 2019-10-28, still changing 2026-09-29 | Repository history | 2026-09 | krane |
| Interval between the two JIT compilers | YJIT merged upstream in 2021; ZJIT upstreamed 2025-04-18, unable to run most benchmarks at merge | Upstream pull request | 2025-04 | ruby/ruby PR 13131 |
Measured: the heap figures, the shape counts and the module sizes are measurements published with the code that produced them. Self-reported: the cutover window and the copy parallelism come from a company talk about its own tooling, with no independent replication. Derived: the bulkhead's five tickets follow from an occupancy estimate the reader must make for their own system, and the arithmetic is only as good as that estimate. Unknown: how many shards exist, how long a real tenant move takes end to end, what any of this costs, and how often the cutover window is missed. No reachable source gives a figure for any of them, so plan the first move to measure them yourself.
Every source behind this page, graded. Filter by kind. The ledger with one
row per claim, including the quotes, ships beside this file as sources.md.
States the problem the tool exists for, the enforcement model, strict mode, and the semantics of the per-package todo file that holds existing violations.
The tool entered public life as an extraction, which dates the point at which the modular-monolith strategy became a shipped mechanism rather than an intention.
Removes the constant-visibility checker from the core gem, sequences it behind a deprecation release, and states the user-visible cost of the break.
The release that made the narrower boundary model the default for every user of the gem, four months after the code change merged.
The full argument for in-process resilience: driver patching, SysV semaphore bulkheads, the ticket arithmetic, and an explicit warning that the library will surprise operators who have not read it.
A programmable TCP proxy used in every development and test environment since October 2014, with the stated goal of proving the absence of single points of failure in tests.
The algorithm, the seven-step cutover, and three hard preconditions: integer primary keys, full row-based replication, and no foreign keys.
A matrix of inconsistency classes against verifiers, with cells reading "Sometimes" and "Probably not", a deprecated verifier, and a note that a MySQL checksum statement is broken for JSON columns.
Why the tool exists (managed MySQL hides the filesystem and the replication interface), what TLA+ found before release, and the only published downtime figures.
The model is committed next to the Go implementation rather than written once and discarded, and the README admits "proofs remain elusive".
Updates and deletes matched nothing because MySQL does not compare a JSON column to a string, and one verifier configuration cannot detect it.
String comparison of binlog filenames terminates the stream early past a million files, in a way no online verifier can catch, demonstrated with a short program.
A query optimisation changed result-column order and the writer kept pairing values positionally; the inline verifier stopped the run.
Annotates every statement the mover issues and fails the run on any unannotated write
to the target, with the review sent to the @Shopify/pods team.
The shard-aware copy filter, which makes a tenant rather than a table the unit that moves between databases.
The first compiler: basic block versioning, written in Rust inside CRuby, with Shopify as the contact address and a company fork named as the discussion venue.
The second compiler is upstreamed from a then-private company repository, with side exits and invalidation unimplemented and most benchmarks crashing.
Method-based compilation guided by interpreter profiles, a different architecture from the first compiler, with bug reports routed to the company fork.
Production object-shape data from the monolith is dumped, ranked, and used to justify changes to the framework itself rather than to application code.
A production memory and pause-time problem traced to boot-time allocation shaping the heap, with the previous workaround described as high maintenance.
A deployment wrapper whose purpose is a verdict: it runs kubectl underneath and answers "what just happened" and "did it work" for CI.
The release note recording the transfer of a widely used Go Kafka client from Shopify to IBM, with the module path change as a breaking change for every dependent.
The JavaScript-to-WebAssembly toolchain started in April 2021 and relabelled a Bytecode Alliance project in May 2023, with module sizes that decide the linking model.
Two more extractions from the same application: an opt-in read-through cache in the ORM invalidated on commit, and an engine that makes data backfills pausable and resumable.
Both JIT documents route contributions through Shopify/ruby, while the
benchmark suite built to justify the first compiler now lives as ruby/ruby-bench
and benchmarks any implementation.
Marketing material, and the only reachable statement of where the platform runs: Kubernetes, Bigtable, BigQuery and Compute Engine, with an AI assistant as the newest layer on top.
Six rungs. The line between a toy and the real thing is rung four, where you stop testing that the mechanism works and start measuring what it hides.
Pick any application with more than one team in it. Declare three packages, generate the violation list, and commit it.
Done when: the check runs in CI and the todo file has a number in it that you can quote. Teaches: that your real boundaries are the ones the backlog measures, not the ones in the diagram.
Set one package to strict mode so no new violations can be recorded against it. Run it for two weeks of normal work.
Done when: either the file shrinks or you have a written reason why that package cannot hold the line. Teaches: whether a static boundary is a constraint in your organisation or a report.
Put a programmable proxy between the application and its slowest datastore, add one second of latency, and assert what the request does.
Done when: a test fails because the application waits instead of degrading. Teaches: how much of your capacity one slow dependency owns.
Measure the fraction of workers concurrently in that datastore over a normal day. Compute the ticket count that keeps the false-positive rate where you want it, and enforce it.
Done when: you can state the accepted false-positive rate as a number and show the measurement behind it. Teaches: that overload protection is a deliberately chosen error budget, not a safety net.
Copy a table between two databases under write load with binlog replay, then seed the hard cases: a JSON column, a composite primary key whose pagination key is not leftmost, and a write to the target from outside the copier.
Done when: you have at least one corruption your verifier does not catch. Teaches: which of your failure classes are invisible, which is the only number that matters before a real tenant moves.
Profile the running application, find a cost that belongs to the runtime or the framework rather than to your code, and open the upstream change with the production data attached.
Done when: the upstream discussion is arguing about your numbers. Teaches: the actual mechanism behind this entire strategy, which is that a shared layer improves for everyone only when somebody brings evidence.
Keeping the monolith is not a decision to do less work. It is a decision about where the work goes, and the public record of this decade says it goes downward: into the compiler, the drivers, the data-movement engine and the build. That trade is worth making when your bottleneck is coordination rather than deployment, because every one of those layers is shared and each improvement lands everywhere at once. The bill arrives as correctness risk in places that used to be somebody else's problem. Your boundaries become advisory, so they need a backlog that shrinks. Your resilience lives inside the driver, so it needs a false-positive rate you chose on purpose. Your partition lives in the data, so moving a tenant becomes a silent-corruption problem that needs a verifier independent of the mover. Before you decide not to split, price the verifier, not the migration.
This page was built with no access to engineering blogs, papers, talks or issue trackers. These are the queries that worked anyway, which makes them worth keeping.
git clone --filter=blob:none https://github.com/ORG/REPOgit log --reverse --date=short --format='%ad %s' | head -1git log --date=short --format='%ad %h %s' -S"TERM" -- README.mdgit log --date=short --format='%ad %s' --all | grep -i "deprecat\|rename\|extract"repo:ORG/REPO privacy in:titlerepo:ORG/REPO in:title corruption OR mismatch OR "data loss"repo:UPSTREAM/REPO author:ENGINEER "in production" in:bodyorg:ORG "caused an incident" in:bodyraw.githubusercontent.com/ORG/REPO/BRANCH/docs/verifiers.mdfind . -name '*.pdf' -o -name '*.tla'repo:ORG/REPO filename:DESIGN.mdgrep -rn -i "postmortem\|outage\|incident" */README.md */docs/*.mdgit log -1 --date=short --format='%ad' TAGrepo:ORG/REPO sort:created-asc TERM in:titleTwo notes on method, both learned the hard way on this page. A conference deck committed
into a documentation directory is a talk you can read offline, with presenter notes, years
after the video host has stopped serving it: look for *.pdf in every repository
before concluding the talks are gone. And a pull request that fixes a defect is usually a
better account than the release note that announces the fix, because the author is
explaining the mechanism to a reviewer rather than to a user.