When owning hardware wins  / field guide
Practitioner field guide · Cost & efficiency · 2026-08-30

When owning hardware wins, and when it owns you

In the summer of 2025, 37signals deleted its AWS account and Stack Overflow unracked its last physical server: the two loudest self-hosting exemplars of the 2010s crossed in opposite directions, weeks apart, and both moves were economically right. This guide reconstructs the rent-or-own decision from the documented record, so a reader can put a defensible number on their own answer and name the conditions that would flip it.

34 primary sources 19 organisations 3 billing incidents, 1 aborted exit Evidence through Aug 2026 Read: ~25 min
01

The territory

The problem, stated without naming any vendor: whether to run steady workloads on computers you rent by the hour, at a premium that buys elasticity, or on computers you own that depreciate whether you use them or not, with the people, spares and lead times that ownership implies.

$74.6M
Dropbox's two-year operating-cost reduction from leaving S3, disclosed in its SEC S-1
8–9%
Share of companies planning full repatriation, against the 83% headline everyone quotes
$72K
Burned in a few hours by one runaway test deployment on elastic billing
29%
Cloud spend practitioners themselves estimate as waste, rising for the first time in five years

Start with the surprise, because it disciplines everything after it. On 30 June 2025, 37signals' final S3 contract lapsed and the company finished moving 18PB onto its own flash arrays, completing an exit it had priced at more than $10M of five-year savings (DHH, 2024; The Register, May 2025). Within weeks, Stack Overflow, the company whose nine on-prem web servers were the standing rebuttal to cloud maximalism for a decade, unracked the last of roughly fifty servers in New Jersey and went entirely to Google Cloud and Azure (Stack Overflow, Aug 2025). Neither company was confused. They had different load shapes, different storage gravity, different teams, and, decisively, different forcing events. The lesson of the whole corpus is that direction is not ideology: every documented move in either direction pairs a forcing event with a unit-economics model, and the same arithmetic points different ways for different shops.

The second thing to fix before going further is the statistics, because the debate runs on numbers that measure different things. The famous 83% figure, from a Barclays CIO survey and amplified by a16z, counts companies moving at least one workload back from public cloud; a company relocating one compliance database and a company exiting entirely land in the same bucket (Channelnomics analysis, 2024). IDC's workload survey puts full repatriation plans at 8–9% of companies, while Gartner has public cloud spend growing past $723B in 2025, up more than 20% in a year (The Stack, Nov 2024). All three numbers are simultaneously true. Repatriation in the record is overwhelmingly workload-level, not company-level, and the interesting question for an architect is not "is the herd leaving" but "which of my workloads is on the wrong side of the line, and what would moving one actually cost".

Figure 1 · Two decades of documented moves, in both directions

2006 to 2008. Early cloud
37signals on EC2 from 2006; Netflix corrupts its own
database in 2008 and turns cloudward

2015 to 2017. First big exit, and its counterpoint
Dropbox builds Magic Pocket and leaves S3; Netflix shuts its last DC;
GitLab proposes metal, then stays; Snap commits $3B of rent

2021. The thesis
a16z prices cloud at about half of software COGS

2022 to 2023. The exits go loud
37signals announces at a $3.2M bill; X empties Sacramento;
Ahrefs prices never-cloud at $400M

2024. The friction falls
Google, then AWS, zero final-exit egress; GEICO rebuilds hybrid on OCP
2025. The crossover
37signals deletes its AWS account;
Stack Overflow unracks its last server
2027. The rule
EU Data Act bans switching charges outright from January 12

2006 to 2008. Early cloud
37signals on EC2 from 2006; Netflix corrupts its own
database in 2008 and turns cloudward

2015 to 2017. First big exit, and its counterpoint
Dropbox builds Magic Pocket and leaves S3; Netflix shuts its last DC;
GitLab proposes metal, then stays; Snap commits $3B of rent

2021. The thesis
a16z prices cloud at about half of software COGS

2022 to 2023. The exits go loud
37signals announces at a $3.2M bill; X empties Sacramento;
Ahrefs prices never-cloud at $400M

2024. The friction falls
Google, then AWS, zero final-exit egress; GEICO rebuilds hybrid on OCP
2025. The crossover
37signals deletes its AWS account;
Stack Overflow unracks its last server
2027. The rule
EU Data Act bans switching charges outright from January 12
Every dated move on this line pairs a forcing event with a cost model; none of them is a fashion choice. Sources: Netflix, Dropbox S-1, GitLab, 37signals, Stack Overflow, Kemp IT Law.
Diagram source

Scope. This guide covers the economics and architecture of moving steady production workloads between rented cloud and owned or colocated hardware, in both directions, plus the middle positions (rented dedicated servers, and self-hosted software on cloud VMs). It deliberately does not cover GPU and AI-training capacity economics, data-sovereignty mandates, desktop virtualisation, or the choice between SaaS and self-hosted applications; and it treats the Broadcom/VMware licensing shock only as a forcing event, not as a hypervisor comparison.

02

What a completed exit actually looks like

Across 37signals, Dropbox, X, OneUptime and GEICO, the successful moves share one shape, and the parts everyone keeps renting are as instructive as the parts they buy.

The common playbook has six stages, and each is attributable. A forcing event opens a decision window: 37signals' $3.2M annual bill and an expiring S3 commitment (DHH, 2022), GEICO's bill rising 2.5× over roughly a decade of cloud-first (Weekly, QCon SF 2024), Dropbox's unit storage cost at hundreds of petabytes. A unit-economics model follows: Ahrefs prices a server-month, Fastmail prices a 2U slot, Dropbox priced cost-of-revenue for the S-1. Then a portability layer: every workload that moved was already, or was first made, an ordinary container or process with open-source datastores underneath. 37signals wrote its own deployment tool, Kamal, whose README promise is exactly this portability: "Deploy web apps anywhere," from bare metal to cloud VMs (basecamp/kamal). Then a landing zone (two colo sites and an ops partner for 37signals; OCP racks and an in-house platform org for GEICO), a phased dual-run cutover (seven apps in about six months at 37signals, per the June 2023 completion post), and finally a remainder that stays rented.

The remainder deserves more attention than it gets. Dropbox, the canonical exit, still served European customers from AWS and kept under 10% of storage there after Magic Pocket, because building datacenters on another continent was a worse trade than renting them (ARCHITECHT, 2017). 37signals kept its CDN and assorted SaaS. Nobody in this corpus repatriated the edge. Inferred, and worth stating plainly: the exits in the record are exits from compute and storage at steady state, never from the cloud as a category.

Figure 2 · Reference architecture of a completed exit

The ground it stands on

Moves to owned metal

Stays rented, in every documented exit

Users

CDN / edge, email,
misc SaaS

Geo + DR remainder
(Dropbox kept AWS intl, under 10%)

Deploy layer: containers over SSH
(Kamal at 37signals)

Commodity compute
(Dell at 37signals, OCP at GEICO)

Storage tier
(Pure arrays 18PB, or Magic Pocket)

Two or more colo sites

Hardware ops:
partner or in-house org

The ground it stands on

Moves to owned metal

Stays rented, in every documented exit

Users

CDN / edge, email,
misc SaaS

Geo + DR remainder
(Dropbox kept AWS intl, under 10%)

Deploy layer: containers over SSH
(Kamal at 37signals)

Commodity compute
(Dell at 37signals, OCP at GEICO)

Storage tier
(Pure arrays 18PB, or Magic Pocket)

Two or more colo sites

Hardware ops:
partner or in-house org

The load-bearing boxes are the bottom two: sites and people. Every documented failure to exit is a failure of one of those, not of the compute layer. Reconstructed from 37signals, Dropbox and GEICO accounts.
Diagram source

Divergence: the storage tier

Three answers exist in production. Buy arrays: 37signals put 18PB of Pure flash in two DCs for about $1.5M plus under $1M of five-year support. Build: Dropbox wrote Magic Pocket, adopted SMR drives first among major operators, and erasure-codes at exabyte scale. Rent anyway: GitLab kept cloud object storage even while planning a metal move, and that dependency helped reverse the plan.

Sources: DHH 2024, Dropbox, The Register 2016

Divergence: the orchestrator

GEICO is building a Kubernetes-and-OpenStack private cloud across OCP hardware, an enterprise platform for hundreds of teams. 37signals went the other way on purpose: containers pushed over SSH, to, in DHH's words, "dodge the complexity of Kubernetes, and avoid any sort of enterprisey service contract entanglements." Fleet size and team count pick the winner, not taste.

Sources: The Stack, DHH 2023

The third position: own software, rent hardware

Zerodha, India's largest broker, self-hosts open-source everything (PostgreSQL, ClickHouse and dozens more) on plain EC2, with about 5% external-vendor dependency, and its CTO estimates at least $3M/yr saved against proprietary subscriptions. The repatriation that matters to them is at the software layer; the hardware stays rented for compliance and convenience.

Source: CNCF case study

Figure 3 · Who moved which way, and what forced it

Metalward to cloud

Netflix, 2008-16
trigger: own-DC failure, then scale

GitLab, 2016-17, exit aborted
trigger: storage ops burden foreseen

Stack Overflow, 2023-25
trigger: lease cliff, small SRE team

VMware estates, 2024 on
trigger: licensing shock; ~75% pick cloud

Cloudward to metal

Dropbox, 2015-16
trigger: unit cost at exabyte scale

X / Twitter, 2023
trigger: company-wide cost purge

37signals, 2022-25
trigger: $3.2M steady-state bill

GEICO, 2023 on
trigger: bill up 2.5x on lift-and-shift

Metalward to cloud

Netflix, 2008-16
trigger: own-DC failure, then scale

GitLab, 2016-17, exit aborted
trigger: storage ops burden foreseen

Stack Overflow, 2023-25
trigger: lease cliff, small SRE team

VMware estates, 2024 on
trigger: licensing shock; ~75% pick cloud

Cloudward to metal

Dropbox, 2015-16
trigger: unit cost at exabyte scale

X / Twitter, 2023
trigger: company-wide cost purge

37signals, 2022-25
trigger: $3.2M steady-state bill

GEICO, 2023 on
trigger: bill up 2.5x on lift-and-shift

Both lanes are full and both lanes are rational. Read the triggers, not the logos: steady storage-heavy load moves left; lease cliffs, small teams and licensing shocks move right. Sources as in the evidence wall.
Diagram source

One mechanism underlies the whole cloudward-to-metal lane, and it is worth naming because the papers wrote it down before the bloggers did. Owned hardware is cheap only at high utilisation, and high utilisation is engineered, not assumed. Google's Borg paper attributes its cluster economics to "admission control, efficient task-packing, over-commitment, and machine sharing" (EuroSys 2015); Facebook's f4 paper drove effective storage replication from 2.8× down to 2.1× with a two-datacenter XOR scheme, saving 53PB of raw disk against 65PB of logical data as of 2014 (OSDI 2014). That is what the rent premium pays someone else to do. The stay-side mirror of the same idea is in AWS's own Karpenter consolidation design, which deletes or replaces underutilised nodes against their price (design doc). Whichever side of the line you sit on, somebody has to do the packing; the decision is who, and at what margin.

03

The decisions that matter

The two mirrored decisions of summer 2025, then the recurring forks with the conditions that flip each one.

37signals, 2022–25: leave, or optimise in place?

Chosen
  • Full exit to owned Dell hardware in two colo sites, run with an ops partner (Deft) and deployed with Kamal
  • "Renting computers is (mostly) a bad deal for medium-sized companies like ours with stable growth" (DHH, Oct 2022)
  • Result: bill $3.2M to $1.3M in year one, then S3 gone mid-2025; savings projection raised from $7M to over $10M across five years
Rejected
  • Staying and optimising: "The savings promised in reduced complexity never materialized" after years of trying
  • Kubernetes as the landing zone, explicitly declined
  • Private-cloud enterprise stacks and their service contracts
Flips when
  • Load stops being steady: spiky or fast-doubling demand re-prices elasticity
  • No ops partner or team will own hardware: the premium buys real labour
  • Your storage is small: the S3 line was the last $1.3M and took a dedicated hardware purchase to replace

Stack Overflow, 2023–25: renew the datacenter, or leave metal?

Chosen
  • Full exit to cloud: Teams to Azure in 2023 as the rehearsal, the public sites to Google Cloud by July 2025
  • Stated reason: engineers were spending themselves on "physical servers, cabling, racking, replacing failed disks, and everything else in between" (Part 1, 2025)
  • Forcing event: the New Jersey datacenter itself shut down; out by 2025-07-31 with no renewal option
Rejected
  • A new colo lease plus a hardware refresh for ~50 aging servers
  • Keeping a Colorado DR site alive (decommissioned that June)
  • Continuing to staff physical ops inside a shrinking engineering org
Flips when
  • The fleet is large enough that refresh amortises: fifty servers cannot carry a storage team; fifty racks can
  • Traffic is growing rather than declining, so bought capacity gets used
  • Someone wants to own hardware as a career path, not a chore

Read together, the two blocks say the quiet part: the decision inputs are load shape, storage gravity, fleet size, team appetite and the contract calendar. Vendor pricing appears in both models, but neither company moved because a price changed; they moved when a commitment expired or a building closed. Corroborated across the corpus: Dropbox timed against S3 contracts, 37signals against its S3 commitment, Stack Overflow against a lease, and the VMware estates now heading for cloud are timing against renewal dates that arrive with 71% of customers reporting outsized price rises (ETR, Dec 2025). Your real decision windows are dated documents, not architecture reviews.

Figure 4 · The decision, as the record actually resolves it

no

yes

no

yes

no

yes

nobody yet

committed

Load steady 12+ months out?

Bill above ~$1M/yr, rackable?

Workloads portable?

Who owns hardware ops?

Stay. Elasticity is
what you are paying for

Optimise in place:
commitments + consolidation

Build portability first.
That is the real project

Middle path: rented dedicated,
or colo with an ops partner

Phased exit on the contract calendar.
Keep CDN, DR and geo rented

no

yes

no

yes

no

yes

nobody yet

committed

Load steady 12+ months out?

Bill above ~$1M/yr, rackable?

Workloads portable?

Who owns hardware ops?

Stay. Elasticity is
what you are paying for

Optimise in place:
commitments + consolidation

Build portability first.
That is the real project

Middle path: rented dedicated,
or colo with an ops partner

Phased exit on the contract calendar.
Keep CDN, DR and geo rented

Terminal nodes are actions. The tree encodes the flips-when conditions from the documented cases; it does not price your workload for you, and rung 4 of the build ladder exists because nothing but your own numbers can.
Diagram source
DecisionChosen (by whom)RejectedBecauseFlips when
Storage tier on metalBuy arrays (37signals: 18PB Pure, ~$1.5M + support) Building a blob store10PB does not amortise a storage-systems team At exabyte scale it flips hard: Dropbox built Magic Pocket and banked $74.6M in two years
Orchestration on metalContainers over SSH, no Kubernetes (37signals, OneUptime went the k8s route instead) Kubernetes/OpenStack platformA handful of apps and one ops team do not need a platform org Hundreds of teams and mixed workloads: GEICO is building exactly that platform on OCP
Who racks the boxesOps partner (37signals with Deft) Hiring a datacenter teamKeeps headcount flat; DHH reports no ops-team growth Fleet grows past what a partner contract covers, or the partner becomes the single point of failure
Exit egress timingAfter the 2024 waivers (37signals: ~$250K waived in a 60-day window) Paying metered egressGoogle (Jan 2024) then AWS (Mar 2024) zeroed final-exit transfer, EU-Data-Act driven Never flips back inside the EU: switching charges are banned outright from 2027-01-12
Commitment posture while rentingDiscounts with exit windows (37signals rode its S3 term to the exit date) Long commitments without options (Snap: $2B + $1B over five years, with pay-the-difference shortfall clauses) Commitments price like debt: cheaper, and binding If a plausible exit or big re-architecture sits inside the term, shorter and dearer wins
Which workloads move firstSteady, storage-heavy, stateless-friendly apps (all exits) Spiky, geo-distributed, managed-service-coupled onesThey carry the premium without using the elasticity A workload that bursts 16x on a viral week (Cara) belongs on elastic capacity, with a budget cap

One more decision hides in a rejected pull request. Kamal's tracker keeps receiving orchestrator-shaped feature proposals, and keeps declining them: a custom-SSL-path implementation was argued down in favour of reading certificates from secrets and closed unmerged in June 2025 (PR #969, superseded by the merged #1531), and a standby-container failover mode was closed unmerged in June 2026 (PR #1810). The tool's boundary is a decision record in itself: the moment you need pre-warmed failover pools and cluster-level scheduling, you have re-derived the platform you left, and you should either accept that platform or stop needing it.

04

What broke, on both sides of the line

Three failure classes account for the published incidents: unbounded elastic billing, the lift-and-shift premium, and adopting an ops discipline nobody owned. The fourth finding is an absence, and it matters as much.

Figure 5 · Anatomy of a billing runaway (Milkie Way, March 2020)

Billing meterFirestoreCloud Run (test app)EngineerBilling meterFirestoreCloud Run (test app)Engineer116B reads + 33M writes16,022instance-hours in 24h$72K in a few hoursdeploy scraping experiment1reads, peaking near 1Brequests/min2scale-out to ~1,000instances per service3usage metered, hours behind4budget alert arrives after the money is gone5
Billing meterFirestoreCloud Run (test app)EngineerBilling meterFirestoreCloud Run (test app)Engineer116B reads + 33M writes16,022instance-hours in 24h$72K in a few hoursdeploy scraping experiment1reads, peaking near 1Brequests/min2scale-out to ~1,000instances per service3usage metered, hours behind4budget alert arrives after the money is gone5
The failure is not the scale-out; it is that the billing meter is the only backpressure in the loop, and it reports hours late. Reconstructed from the Milkie Way postmortem.
Diagram source
Postmortem

Elastic billing: $72K in hours

AssumptionA test project with a small budget cap cannot spend real money.
What happenedA Cloud Run experiment scaled itself against Firestore: ~1,000 instances, reads peaking near 1B requests/minute, 116B reads total; GCP budgets alert but do not stop spend, and billing data lagged by hours.
Blast radius$72K liability in a few hours against a startup's ~$100K of funds, March 2020.
FixQuotas set to bounded values everywhere; billing treated as an incident source with its own monitoring.
Design ruleOn elastic pricing, spend is a failure domain. A budget that only emails is an alert, not a limit; know the platform's actual kill switch before the first deploy.
Postmortem

Elastic billing: one cache setting, $11K

AssumptionThe CDN absorbs download traffic, so origin egress stays negligible.
What happenedNew Pwned Passwords files exceeded Cloudflare's 15GB max-cacheable-object setting; every download silently went to Azure origin, metered as egress.
Blast radiusRoughly $11K of unexpected charges on the December 2021 invoice; detection was the invoice itself.
FixCache limits raised and egress dashboards watched as first-class signals.
Design ruleAny config boundary that changes where bytes are served from is a cost control, and it fails silent. Alert on origin egress volume, not on the bill.
Postmortem

Elastic billing: virality at $96K/month

AssumptionServerless pricing scales down to zero, so it must scale up affordably too.
What happenedCara grew from 40K to 650K users in a week; function invocations peaked at 56M/day on per-invocation pricing.
Blast radiusA ~$96K–$98K Vercel month for a free, artist-run app, June 2024.
FixPublic appeals, plan renegotiation, and re-architecture pressure toward pooled compute.
Design ruleSuccess is a load pattern. Before launch, price the best case, not only the expected one; per-request billing turns a growth spike into a solvency event.
Incident review

Lift-and-shift: 2.5× the bill, less reliability

AssumptionMoving existing workloads into cloud, as they are, captures cloud economics.
What happenedGEICO reports that about ten years into cloud-first the migration was still unfinished, bills had risen 2.5×, availability had suffered, and the firm sat "at the cumulative mercy of our clouds" with no consistent data strategy or hybrid stack.
Blast radiusMore than $300M/yr of cloud spend by 2021, across multiple providers.
FixHybrid rebuild: repatriating storage-heavy and steady workloads onto OCP hardware with an open-source platform, keeping cloud where it earns its premium.
Design ruleCloud pricing assumes you re-architect for elasticity. If the org will not, the premium buys nothing; price lift-and-shift as rent with no upside.
Decision failure

The exit that was reviewed out of existence

AssumptionBare metal plus Ceph would beat cloud on performance and cost for GitLab.com.
What happenedGitLab published its late-2016 hardware proposal, took hundreds of comments and emails "filled with advice and warnings" about what running its own storage would really take, and reversed in December 2016. The drafted announcement, "Why We're Choosing Bare Metal," survives as a closed issue; the post that shipped in March 2017 is titled "Why we are not leaving the cloud."
Blast radiusNone in production, which is the point: the failure was caught at design review, in public.
FixStay in cloud; invest in making the cloud deployment fast and stable instead.
Design ruleRepatriation trades a bill for an org chart. Cost the org chart first, and let people who run the target stack review the plan before any hardware is ordered.
Contract cliff

Commitments with pay-the-difference clauses

AssumptionCommitted-use discounts are free money against spend you will incur anyway.
What happenedSnap's 2017 S-1 disclosed $2B over five years committed to Google Cloud plus $1B to AWS, and stated: "If we fail to meet the minimum purchase commitment during any year, we are required to pay the difference," listing the dependency as an IPO risk factor.
Blast radiusNot an outage: a floor under costs and a ceiling over architectural freedom for five years, disclosed to investors as risk.
FixStructural, industry-level: by 2024 the exit-egress waivers and the EU Data Act began regulating the sharpest exit frictions away.
Design ruleA commitment is a short position on your own change. Before signing, check what re-architecture or exit inside the term would cost, and negotiate the option in.
The absence that shapes the risk

No published postmortem in this corpus attributes a production incident to an exit cutover, or to post-exit hardware operations. 37signals moved seven apps and 10PB with no incident write-up; Stack Overflow's move produced retrospectives, not apologies. Either the cutovers work, or their failures go unpublished; both readings point the same way: the risk you cannot price from public evidence is year-two operations, not migration weekend. Only OneUptime has published a multi-year follow-up (two years later, savings grown to $1.2M/yr), and one account is anecdote, not a base rate.

05

Numbers you can plan against

Everything quantitative in one place, dated. Measured means a primary account with arithmetic behind it; claimed means self-reported without methodology; survey means aggregated self-reports.

MetricValueAtStatusAs ofSource
Two-year opex reduction from storage exit$74.6MDropboxMeasured (SEC filing)2018S-1
Year-one components of the above−$92.5M / +$53MDropboxMeasured: third-party DC spend cut vs own-DC cost added, netting $39.5M2016GeekWire
Annual bill before/after exit year one$3.2M → $1.3M37signalsMeasured (self-reported with line items)2024DHH
Exit hardware spend~$700K37signalsMeasured; paid back inside the first year's savings2023DHH
18PB flash, bought + 5yr support~$1.5M + <$1M37signalsMeasured (self-reported)2024DHH
Final-exit egress waived~$250K37signals / AWSReported; 60-day window under the 2024 policy2025AWSInsider
Colo server-month, all-in$1,550AhrefsClaimed; own model, includes space, power, transit, network2023Ahrefs
Cloud-equivalent server-month$17,557AhrefsClaimed; EC2 list-price equivalent of hardware clouds do not sell as one instance (2TB RAM, 16×15TB, 2×100Gbps)2023The Register
The $400M headline, checked850 × ($17,557−$1,550) × 30 ≈ $408MAhrefsDerived here from their own per-server figures; matches their $39.5M vs $447.7M totals2023The Register
Storage server, 24×61TB NVMe, 2U~$190K + ~$3K/yr coloFastmailMeasured (self-reported)2024Fastmail
Managed k8s bill pre-exit>$38K/mo (28 nodes)OneUptimeMeasured (self-reported); savings $230K/yr at exit, >$1.2M/yr two years on2023/2025OneUptime
Cloud + storage cost cuts−60% / −75%X (Twitter)Claimed; no methodology published; Sacramento exit put at $100M/yr, 5,200 racks, 148K servers2023DCD
Bill growth under cloud-first2.5× over ~10yrs; >$300M/yrGEICOClaimed in talks; consistent across two venues2021/2024The Stack
Committed cloud spend at IPO$2B + $1B / 5yrsSnapMeasured (S-1 disclosure), with shortfall clause2017Recode
Cloud share of software COGS~50%50 public SaaS cosDerived by a16z from filings; contested, directionally accepted2021a16z
Self-estimated cloud waste / over-budget orgs29% / 17%Flexera respondentsSurvey; waste rose for the first time in five years2026Flexera
Companies planning full repatriation8–9%IDC surveySurvey; against 83% moving ≥1 workload (Barclays)2024The Stack
Public cloud end-user spend$723.4B (from $595.7B)WorldwideForecast (Gartner)2025The Stack
EU switching chargescost-based now, $0 from 2027-01-12All providers serving EURegulated fact (Reg. 2023/2854, applicable 2025-09-12)2025Kemp IT Law
Effective storage replication factor2.8× → 2.1×Facebook f4Measured (paper); 53PB raw saved on 65PB logical2014OSDI 14
Read these carefully

The Ahrefs and X rows are the ones most often quoted and least verifiable: Ahrefs compares against list-price on-demand for a hardware shape no cloud sells as one instance, and X published percentages with no baseline. The Dropbox row is the only figure in this table that passed an audit. The a16z 50% figure is a derivation from public filings by an investor with a thesis; treat it as an upper anchor. Undated cloud list prices change quarterly, which is why none appear here; re-derive any comparison with current pricing before quoting it in a review.

06

The evidence wall

Every source behind this page, graded by kind. Filter it; the postmortems and the two closed-unmerged pull requests are the rows generic coverage never cites.

Postmortem Milkie Way2020-12

We Burnt $72K testing Firebase + Cloud Run

A test deployment recursively scaled to ~1,000 Cloud Run instances driving ~1B Firestore reads/minute; budgets alerted but did not stop spend, and billing lagged hours. $72K in an afternoon, March 2020.

Carry forwardElastic billing is a failure domain with alert-only backpressure. Bound every quota before the first deploy.
blog.tomilkieway.com/72k-1
Postmortem Troy Hunt / HIBP2022-01

How I Got Pwned by My Cloud Costs

Files grew past Cloudflare's 15GB cacheable-object cap; downloads silently fell through to Azure origin egress. ~$11K surprise on the December 2021 invoice, found on the invoice.

Carry forwardConfig that changes where bytes are served from is a cost control that fails silent. Alert on origin egress, not the bill.
troyhunt.com
Postmortem Cara / InfoQ2024-06

A viral week on per-invocation pricing

40K to 650K users in a week; 56M function invocations/day at peak; a ~$96K–$98K Vercel month for a free app. Incident review assembled from the founder's public accounts.

Carry forwardPrice the best case before launch. Success on per-request billing is a solvency event, not a scaling event.
infoq.com
Source Basecamp2023–

basecamp/kamal: the exit's portability layer

"Deploy web apps anywhere." Containers pushed over SSH with zero-downtime proxy swaps; deliberately not an orchestrator. The configuration surface is a candid map of what leaving managed platforms makes you own.

Carry forwardPortability is a tool you can read, not a slogan. If your workload cannot be described in a Kamal-sized config, it is not portable yet.
github.com/basecamp/kamal
Source Basecamp2025-06

Kamal PR #969, closed unmerged

Custom SSL certificate paths, argued down by maintainer djmb ("I think we should read these directly from the Kamal secrets") and superseded by the merged secrets-based #1531. A complete rejected-design argument, on the record.

Carry forwardRead a tool's closed-unmerged PRs before adopting it; they are the boundary of what it will ever do for you.
github.com/basecamp/kamal/pull/969
Source Basecamp2026-06

Kamal PR #1810, closed unmerged

Standby containers "in a Created (stopped) state" for near-instant failover, proposed March 2026 (issue #1809), closed June 2026 without a public design rationale. Orchestrator-shaped asks keep arriving; the boundary holds.

Carry forwardWhen your needs turn into pre-warmed pools and scheduling, you have re-derived the platform you left. Decide, don't drift.
github.com/basecamp/kamal/pull/1810
Source GitLab2016

The announcement that never shipped

Blog-post issue #289, "Why We're Choosing Bare Metal & How We Knew it was Time to Leave the Cloud," drafted during the 2016 metal proposal and left behind when the decision reversed in public.

Carry forwardPublish the proposal before the purchase order. GitLab's community review was the cheapest incident it never had.
gitlab.com issue #289
Design doc AWS / Karpenter2022

Consolidation design (merged RFC)

The stay-side answer to exit math: delete nodes whose pods fit elsewhere, replace nodes with cheaper ones, reasoning explicitly about "what the node we are considering replacing costs."

Carry forwardRun consolidation-class optimisation for 90 days before any exit decision; it captures the same idle capacity an exit would.
karpenter designs/consolidation.md
Specification FinOps Foundation2023–

FOCUS: a common schema for cost data

"A common schema for technology cost and usage data across cloud, SaaS, data center, and other technology categories." The spec exists because rent-vs-own comparisons fail at the data layer first.

Carry forwardNormalise the bill before modelling the exit; a TCO built on unnormalised line items inherits every vendor's flattery.
github.com FOCUS_Spec
Case study Dropbox / SEC2018-02

The S-1: the only audited exit economics

$39.5M cost-of-revenue reduction in 2016 (third-party datacenter spend down $92.5M, own-datacenter costs up $53M), a further $35.1M in 2017: $74.6M over two years, disclosed under securities law.

Carry forwardModel exits net, not gross: Dropbox's headline hides $53M of new owned-infrastructure cost in year one alone.
sec.gov S-1
Case study Ahrefs2023-03

Never-cloud, priced per server-month

850 colo servers at $1,550/month all-in against $17,557 for an EC2-equivalent: $39.5M vs $447.7M over 30 months. The hardware shape (2TB RAM, 16×15TB drives, 2×100Gbps) is exactly what clouds price worst.

Carry forwardThe gap is largest where your hardware shape diverges from cloud instance shapes. Check divergence before believing any multiple.
tech.ahrefs.com
Case study Zerodha / CNCF2023

Self-hosted software on rented hardware

India's largest broker runs self-hosted FOSS on plain EC2 with ~5% external-vendor dependency; CTO Kailash Nadh puts the saving at ≥$3M/yr against proprietary subscriptions. Repatriation at the software layer only.

Carry forwardManaged-service spend and hardware spend are separable decisions. You can repatriate one without touching the other.
cncf.io/case-studies/zerodha
Case study Snap / Recode2017-03

$3B of committed cloud spend, as an IPO risk factor

$2B/5yr to Google plus $1B/5yr to AWS, with the S-1 admitting "If we fail to meet the minimum purchase commitment during any year, we are required to pay the difference."

Carry forwardRenting has lock-in too; it is written in commitment schedules rather than depreciation tables.
recode.net
Eng blog 37signals2024-10

Cloud-exit savings will top ten million over five years

Bill $3.2M to $1.3M in year one; the remainder was one S3 contract, retired mid-2025 with 18PB of bought flash (~$1.5M plus support). The projection that started at $7M was raised, not walked back.

Carry forwardA credible exit model updates in public, line item by line item, against its own earlier predictions.
world.hey.com/dhh
Eng blog Stack Overflow2025-08

Moving the public sites to the cloud, Part 1

Sixteen years of racking, cabling and disk swaps by a small SRE team, ended by a datacenter that closed underneath them: out by 2025-07-31, no renewal option, public sites to Google Cloud.

Carry forwardA lease expiry is an architecture decision arriving on someone else's schedule. Keep the calendar where the architects can see it.
stackoverflow.blog
Eng blog Stack Overflow2025-12

The Great Unracking

Roughly fifty servers decommissioned in New Jersey; the Colorado DR site had gone that June. The company that taught a generation to run lean on metal now owns no hardware at all.

Carry forwardOn-prem excellence is a staffing posture, not an asset. When the people and the traffic go, the racks are pure liability.
stackoverflow.blog
Eng blog GitLab2017-03

Why we are not leaving the cloud

The public reversal of the 2016 bare-metal proposal, after hundreds of comments and emails warned what owning a Ceph estate would take. The rare exit postmortem written before the exit.

Carry forwardThe best time to catch a bad infrastructure decision is while it is still a blog draft.
about.gitlab.com
Eng blog Fastmail2024-12

Why we use our own hardware

Twenty-five years on owned metal; current fleet: 2U servers with 24×61TB NVMe at ~$190K each, ~$3K/yr per 2U for space, power and cooling. The stayer's argument is knowing your load curve for decades.

Carry forwardLong-lived, well-understood workloads are where ownership compounds; the discount is paid for in institutional knowledge.
fastmail.com/blog
Eng blog Dropbox2016-05

Inside the Magic Pocket

The custom multi-exabyte blob store behind the exit: replication plus an erasure coding variant, later SMR drives adopted first among major operators. The $74.6M did not come from racking servers; it came from storage systems engineering.

Carry forwardAt extreme scale the exit's payoff is proportional to how much storage engineering you can field, not to server prices.
dropbox.tech
Eng blog OneUptime2025-10

AWS to bare metal, two years later

The only multi-year exit retrospective in the record: a $456K/yr managed-k8s bill became $230K/yr of savings at exit, growing past $1.2M/yr as the fleet grew, with the operational questions answered in public.

Carry forwardDemand the two-years-later account, from vendors and from your own team; the exit-week post is marketing on both sides.
oneuptime.com/blog
Eng blog Netflix2016-02

Completing the cloud migration

Triggered by a 2008 corruption in its own datacenter that stopped DVD shipping for three days; finished January 2016 when the last datacenter closed. Seven-plus years for one committed company: the honest duration anchor for any full migration.

Carry forwardFull migrations are measured in years per direction. Multi-year commitments signed mid-flight are how you get Snap's clause.
about.netflix.com
Eng blog Duckbill Group2021/2025

The skeptic's ledger: repatriation isn't a thing, then gets complicated

Corey Quinn's standing argument: Dropbox is always the example because a second one is hard to name; his later update concedes nuance while holding that wholesale exits stay rare. The strongest published counterweight to exit triumphalism.

Carry forwardWeigh survivorship: exits publish victory laps, stays publish nothing, and the base rate lives in the surveys, not the headlines.
lastweekinaws.com
Paper Google2015

Borg (EuroSys 2015)

High utilisation from "admission control, efficient task-packing, over-commitment, and machine sharing" across cells of tens of thousands of machines. The mechanism that makes owned hardware cheap, written down by its largest operator.

Carry forwardUtilisation is engineered. An exit model that assumes 70% packing without naming who builds the packer is fiction.
research.google
Paper Facebook2014

f4: warm BLOB storage (OSDI 2014)

Erasure coding plus a cross-datacenter XOR scheme cut effective replication from 2.8× to 2.1×, saving 53PB of raw storage against 65PB logical at publication. Cost-driven architecture at the owner's limit.

Carry forwardThe cloud's margin sits on exactly this kind of engineering. Owning hardware means either doing some of it or leaving that money on the table.
usenix.org (PDF)
Talk GEICO / InfoQ2024-11

Rebecca Weekly at QCon SF 2024 (44:39, transcript)

The enterprise verdict on cloud-all-in: unfinished after a decade, bills up 2.5×, reliability down, "at the cumulative mercy of our clouds." The rebuild is hybrid on OCP hardware with an open-source platform.

Carry forwardThe alternative to a bad cloud posture is rarely metal alone; it is a deliberate hybrid with a data strategy.
infoq.com/presentations
Talk Lex Fridman Podcast2025-06

DHH on the exit, timestamped (#474, from 3:24:04)

The exit's author on the record, in a transcribed long-form interview: early AWS adoption from 2006 ("in the cloud before it was cool"), the premium for elasticity never used, and the decision mechanics behind the posts.

Carry forwardThe strongest exit advocates were early cloud adopters; the argument is about load shape and margin, not nostalgia.
lexfridman.com (transcript)
Vendor a16z2021-05

The Cost of Cloud, a Trillion Dollar Paradox

"You're crazy if you don't start in the cloud; you're crazy if you stay on it." Cloud at ~50% of COGS across 50 public SaaS companies, ~$100B of market value suppressed. An investor thesis, and the debate's reference point.

Carry forwardUse it as the upper anchor on what staying costs, not as a measurement; its critics (Duckbill, CloudZero) are in this wall too.
a16z.com
Vendor Flexera2026-03

2026 State of the Cloud

Self-estimated waste up to 29%, the first rise in five years, driven by AI workloads; 17% of orgs over budget; fewer than half using any one commitment discount per provider.

Carry forwardMost organisations have not exhausted stay-side savings. Exit math against an unoptimised bill flatters the exit.
flexera.com
Vendor ETR / everyWAN2025-12

The VMware reckoning, surveyed

71% of surveyed VMware customers report price rises well beyond the software market since the acquisition; 42% plan partial migration to Microsoft within a year; of those leaving, close to 75% pick public cloud over another hypervisor.

Carry forwardLicensing shocks push workloads cloudward, not metalward: the counterflow is bigger than the exits and runs the other way.
research.etr.ai
07

Build the decision, then earn the exit

Seven rungs from a spreadsheet you can build this week to an exit (or a documented decision to stay) your CFO and your on-call rota both sign.

Reconstruct your real bill

Normalise one month of billing into unit economics per workload (FOCUS gives you the schema). Find the three line items that dominate; in the documented exits they were steady compute, storage at rest, and egress.

Done when: you can price one request, one user and one stored TB.  Teaches: where the money actually is, which is rarely where the dashboard points.

Run one real service on one rented dedicated box

Rent a dedicated server for the price of a lunch a month, deploy a production-shaped service with Kamal or plain containers, and replay production-shaped load at it.

Done when: p95 matches your cloud baseline and you know the box's actual ceiling.  Teaches: what the managed layer was doing for you, item by item.

Kill it on purpose

Fail the disk, reboot mid-traffic, restore the datastore from backup onto a second box, and time every step. This is the rung that separates tourists from residents.

Done when: a written restore runbook beats your stated RTO.  Teaches: the operational surface you would be adopting, before it is yours at 3am.

Price the exit honestly, both columns

Three-year TCO for your top workload: hardware plus refresh, colo, people or partner, support contracts, the rented remainder, and the migration itself. Beside it, the optimised stay: commitments, rightsizing, consolidation.

Done when: two numbers a CFO accepts, with the crossover condition named.  Teaches: your own flips-when, which no blog post can supply.

Optimise in place for 90 days first

Apply the stay-side machinery before deciding: commitment coverage, consolidation or its equivalent, storage tiering, egress audit. Flexera's respondents call 29% of spend waste; claim yours either way.

Done when: waste is under ~10% or you hold evidence the premium is structural.  Teaches: whether the exit's savings were actually available without moving.

Pilot one repatriation with a rollback

Move one steady, storage-heavy workload to the landing zone with dual-run, reconciliation, and a written, rehearsed rollback. 37signals moved seven apps this way; Stack Overflow rehearsed on Teams before touching the public sites.

Done when: 30 days on the new floor with error budget intact and unit cost measured.  Teaches: cutover mechanics and the true shape of day-two work.

Put the decision on the calendar, permanently

Build the renewal-and-lease calendar with decision windows and trigger conditions (bill growth, utilisation, egress share, headcount). Every documented move in this guide happened at a contract boundary; engineer your next one instead of meeting it by surprise.

Done when: the next renewal arrives with an option you designed.  Teaches: that rent-or-own is a standing decision with dated inputs, not a one-time debate.

08

Keep hunting

The queries that found this material, grouped by what they surface. The numbers will stale; these will not.

Exit and stay accounts

  • "why we're leaving the cloud" OR "we have left the cloud"
  • "why we use our own hardware"
  • "moving from AWS to bare-metal" saved
  • "by not going to the cloud" saved

The counterflow

  • "why we are not leaving the cloud"
  • "the great unracking" stackoverflow
  • "journey to the cloud" site:stackoverflow.blog
  • VMware Broadcom migration survey "public cloud" percentage

Money incidents

  • "we burnt" OR "we burned" cloud bill postmortem
  • "how I got pwned by my cloud costs"
  • serverless bill "$96,000" OR "96k" vercel
  • "budget alert" lag "billing" postmortem quota

Decisions, filings, regulation

  • is:pr is:closed is:unmerged repo:basecamp/kamal
  • S-1 "cost of revenue" infrastructure savings site:sec.gov
  • EU Data Act switching charges "12 January 2027"
  • path:designs consolidation repo:aws/karpenter-provider-aws
09

References

Checked 2026-08-30. This session ran under an egress allowlist: github.com URLs were fetched directly; all others were verified live through search-engine retrieval, several with exact-phrase confirmation. The companion sources.md ledger records the claim and supporting quote taken from each.

  1. DHH, Why we're leaving the cloud37signals / HEY World, 2022-10-19. Checked 2026-08-30.
  2. DHH, We stand to save $7m over five years from our cloud exit37signals / HEY World, 2023. Checked 2026-08-30.
  3. DHH, We have left the cloud37signals / HEY World, 2023-06. Checked 2026-08-30.
  4. DHH, Our cloud-exit savings will now top ten million over five years37signals / HEY World, 2024-10. Checked 2026-08-30.
  5. 37signals, Leaving the Cloud (hub page)basecamp.com, ongoing. Checked 2026-08-30.
  6. DCD, 37signals claims it saved almost $2m last yearDatacenter Dynamics, 2025. Checked 2026-08-30.
  7. The Register, 37signals is completing its on-prem move2025-05-09. Checked 2026-08-30.
  8. AWSInsider, Cloud data egress fee tussle plays out with $250k AWS compConverge360, 2025-05-09. Checked 2026-08-30.
  9. Dropbox, Form S-1SEC EDGAR, 2018-02-23. Checked 2026-08-30.
  10. GeekWire, Dropbox saved almost $75 million over two years2018-02-23. Checked 2026-08-30.
  11. Techmeme snapshot of the S-1 savings coverage2018-02-23. Checked 2026-08-30.
  12. Dropbox, Inside the Magic Pocketdropbox.tech, 2016-05-06. Checked 2026-08-30.
  13. Dropbox, How we optimized Magic Pocket for cold storagedropbox.tech, 2019. Checked 2026-08-30.
  14. ARCHITECHT, Dropbox's AWS migration was as much about controlMedium, 2017. Checked 2026-08-30.
  15. Rebecca Weekly, Evaluating and Deploying State-of-the-Art Hardware (QCon SF 2024)InfoQ recording + transcript, 44:39. Checked 2026-08-30.
  16. The Stack, $300 million hyperscaler bill triggered cloud repatriation2024. Checked 2026-08-30.
  17. The Stack, GEICO repatriates work from the cloud2024. Checked 2026-08-30.
  18. Stack Overflow, Moving the public sites to the cloud: Part 12025-08-28. Checked 2026-08-30.
  19. Stack Overflow, The Great Unracking2025-12-24. Checked 2026-08-30.
  20. Stack Overflow, Journey to the cloud part I (Teams to Azure)2023-08-30. Checked 2026-08-30.
  21. Stack Overflow, Are clouds having their on-prem moment?2023-02-20. Checked 2026-08-30.
  22. GitLab, Why we are not leaving the cloud2017-03-02. Checked 2026-08-30.
  23. The Register, GitLab to dump cloud for its own bare metal Ceph boxen2016-11-14. Checked 2026-08-30.
  24. GitLab blog-posts issue #289 (the unshipped announcement)gitlab.com, 2016. Checked 2026-08-30.
  25. Hacker News discussion of the GitLab reversal2017-03. Checked 2026-08-30.
  26. Netflix, Completing the Netflix Cloud Migration2016-02. Checked 2026-08-30.
  27. DCD, X/Twitter claims $100m in annual savings after exiting Sacramento2023-10. Checked 2026-08-30.
  28. Ahrefs, How Ahrefs Saved US$400M in 3 Years by NOT Going to the Cloud2023-03. Checked 2026-08-30.
  29. The Register, Software vendor says own hardware $400m cheaper than cloud2023-03-13. Checked 2026-08-30.
  30. Fastmail, Why we use our own hardware2024-12-22. Checked 2026-08-30.
  31. CNCF, Zerodha case study2023. Checked 2026-08-30.
  32. OneUptime, Moving from AWS to bare metal saved us $230,000/yr2023-10-30. Checked 2026-08-30.
  33. OneUptime, AWS to Bare Metal Two Years Later2025-10-29. Checked 2026-08-30.
  34. Milkie Way, We Burnt $72K testing Firebase + Cloud Run (Part 1)2020-12. Checked 2026-08-30.
  35. Troy Hunt, How I Got Pwned by My Cloud Costs2022-01. Checked 2026-08-30.
  36. InfoQ, Vercel serverless scale expenses (Cara)2024-06. Checked 2026-08-30.
  37. Recode, This is what Snap is paying Google $2 billion forVox Media, 2017-03-01. Checked 2026-08-30.
  38. SiliconANGLE, AWS follows Google Cloud in canceling egress fees2024-03-05. Checked 2026-08-30.
  39. Kemp IT Law, The End of Switching Charges2024/2025. Checked 2026-08-30.
  40. Wang & Casado, The Cost of Cloud, a Trillion Dollar Paradoxa16z, 2021-05. Checked 2026-08-30.
  41. Corey Quinn, Cloud Repatriation Isn't a ThingDuckbill Group, 2021. Checked 2026-08-30.
  42. Corey Quinn, Cloud Repatriation is Getting ComplicatedDuckbill Group, 2024/2025. Checked 2026-08-30.
  43. Channelnomics, Breaking Down the 83% Public Cloud Repatriation Number2024. Checked 2026-08-30.
  44. The Stack, Gartner forecast dampens cloud repatriation outlook2024-11-20. Checked 2026-08-30.
  45. Flexera, 2026 State of the Cloud Report2026-03. Checked 2026-08-30.
  46. ETR, VMware Customers Face A Cost Reckoning2025-12. Checked 2026-08-30.
  47. everyWAN, Two years of Broadcom: the mass VMware exodus that never happened2026. Checked 2026-08-30.
  48. basecamp/kamal (repository)GitHub, checked at HEAD 2026-08-30 (fetched directly).
  49. kamal PR #969, custom SSL certificates (closed unmerged)GitHub, closed 2025-06-19. Checked 2026-08-30 (fetched directly).
  50. kamal PR #1810, standby containers (closed unmerged)GitHub, closed 2026-06-16. Checked 2026-08-30 (fetched directly).
  51. kamal issue #1809, deploy-without-start proposalGitHub, opened 2026-03-26. Checked 2026-08-30 (fetched directly).
  52. Karpenter consolidation design documentGitHub, merged RFC, 2022. Checked 2026-08-30 (fetched directly).
  53. FinOps Foundation, FOCUS specificationGitHub, active. Checked 2026-08-30 (fetched directly).
  54. Verma et al., Large-scale cluster management at Google with BorgEuroSys 2015; PDF at research.google.com/pubs/archive/43438.pdf. Checked 2026-08-30.
  55. Muralidhar et al., f4: Facebook's Warm BLOB Storage SystemOSDI 2014, USENIX. Checked 2026-08-30.
  56. Lex Fridman Podcast #474, DHH (transcript; cloud exit from 3:24:04)2025-06. Checked 2026-08-30.