intermediate 3 min answer

Dropbox's customer data grew from roughly 40 PB to roughly 500 PB before it moved the bulk of file storage onto its own hardware, and its 2018 S-1 says that migration completed in the fourth quarter of 2016. Using 2026 commodity object-storage list prices, roughly what annual bill sits behind 500 PB, how does that compare with the saving Dropbox actually reported, and which assumption dominates the error?

dropboxbuild-vs-buyunit-economicsstoragelist-price
Show the full answer Hide the answer

The assumptions, stated

  • 500 PB of stored customer data, counted as logical bytes, 500 million GB on a decimal basis.
  • S3 Standard in us-east-1 at 2026 list prices: $0.023 per GB-month for the first 50 TB, $0.022 for the next 450 TB, $0.021 above 500 TB. At this volume effectively all of it bills at $0.021.
  • Storage only. Requests and egress are excluded, and for a sync product egress is not small.

The arithmetic

500,000,000 GB x $0.021 = about $10.5 million a month, so on the order of $125 to 135 million a year at list. Treat it as one significant figure.

Now the reported number. Dropbox's S-1 describes an "Infrastructure Optimization" initiative that moved the vast majority of user data off third-party providers onto its own co-located infrastructure, and reports infrastructure cost decreases of $39.5 million in 2016 and \(35.1 million in 2017 - **\)74.6 million over two years, roughly $37 million a year.**

So the realised annual saving was on the order of a quarter to a third of the list-price bill.

Which assumption dominates the error

Not the capacity. Petabytes stored is known to within ten percent and the price tiers are published. The dominant term is the effective negotiated price per GB-month, which at this volume is far below list and is never disclosed. A 60% enterprise discount moves the estimate from $130 million to $52 million, which is a bigger swing than any other assumption in the model. Second-order: whether the 500 PB is logical or raw - building it yourself means paying for your own redundancy, around 1.3 to 1.5x raw with erasure coding and 2 to 3x with replication - and the fact that the in-house programme has its own permanent run cost in data centres, hardware refresh and an infrastructure organisation. The $74.6 million is net of that.

What the number rules in or out

Divide the saving by a fully-loaded engineer at order $300,000 to \(400,000 a year. **\)37 million funds something like 90 to 120 engineers permanently**, which is what an exabyte-scale storage organisation costs, and storage was Dropbox's cost of goods sold rather than an overhead. The arithmetic closes.

Run the same division on your own bill. A $2.4 million annual storage spend with a 40% theoretical saving yields about $1 million a year, which funds three engineers - not a storage team, and not a hardware refresh cycle. The decision rule: compute the saving at your negotiated price, divide by fully-loaded engineer cost, and do not build unless the quotient comfortably exceeds the headcount the thing needs to be run safely forever. Under roughly ten engineers' worth of saving the answer is no, because the team you can fund is smaller than the on-call rota the system needs.

Common weak answers

  • Quoting list price as the prize. It overstates the opportunity by two to three times, and it is the single most common error in repatriation business cases.
  • Treating the saving as free cash. It is gross; the in-house system's own cost comes out of it, and the one-off programme cost comes out of several years of it.
  • Ignoring the head start. Dropbox already ran its own metadata tier and data centres, so the marginal programme was far smaller than a greenfield build. A company starting from pure cloud is pricing a different project.