On 8 August 2024 Hugging Face announced it had acquired XetHub, a 14-person storage startup, describing it as its largest acquisition to that date. Continuing on Git LFS, or building content-defined chunking in-house, were the alternatives. What cost model separates the three, and where would copying this be a mistake?
Show the full answer Hide the answer
The situation they were in
The Hub distributes very large model and dataset artifacts. Git LFS stores and transfers whole files, so an iterative workflow that rewrites a multi-gigabyte checkpoint repeatedly creates a new full copy in storage and a new full transfer on the wire each time. Cost grows with the number of edits rather than with the amount of distinct content, which is the wrong curve for a platform whose users iterate.
The chosen replacement, Xet, uses content-defined chunking, splitting files at boundaries decided by the data so that an insertion shifts one chunk rather than every byte after it. Hugging Face's published documentation describes chunks of roughly 64 KB, and XetHub's own benchmarking claimed around a 50% improvement in storage and transfer against Git LFS on iterative development workloads. Treat that figure as a benchmark on a particular workload, not as a constant.
What they chose, and the model that separates the options
Acquisition, announced 8 August 2024, with 14 employees joining and the price undisclosed.
All three options need the same five terms over a three-year horizon:
- Storage, at the deduplicated size rather than the raw size. This is the term the change is supposed to bend, and it decouples stored bytes from uploaded bytes.
- Transfer and egress bytes. For a distribution business this is usually the largest line, and dedup plus client-side chunk caching reduces it on every repeat download, not only on upload.
- Request and metadata operations. A chunked system replaces a few large object operations with very many small ones. At standard object-store list prices in 2026, roughly $0.005 per 1000 writes and $0.0004 per 1000 reads, a naive 64 KB scheme turns one 10 GB upload into on the order of 160000 operations. Batching chunks into larger blocks is not an optimisation, it is what makes the model work at all, and a cost model that omits this term will recommend the design that bankrupts you.
- Compute for chunking, hashing and the global dedup index, which is real and grows with upload volume rather than with stored bytes.
- Engineering time to parity, and the calendar time before any saving starts.
Build, buy and acquire differ almost entirely in that last term. Building it is not a storage project; it is a content-addressed store with a global dedup index, a client library, and a migration of live repositories. Acquisition converts an uncertain multi-year schedule into a one-off price and a team that has already made the design mistakes.
Why it fit their constraints
Storage and transfer are cost of goods sold for an artifact-distribution business, so bending that curve changes the business rather than the infrastructure bill. A one-off price against a recurring curve is the case where acquisition arithmetic works, and it works better the longer the horizon and the faster the underlying volume grows.
Where copying this would be a mistake
- If artifacts are not rewritten, dedup saves close to nothing. The saving comes from repeated near-identical versions. Measure the rewrite ratio before modelling anything.
- If storage and transfer are a small share of spend, the whole exercise is a rounding error however elegant the mechanism.
- Acquiring a team is available to a company with capital and a hiring brand. For almost everyone else the same decision reads as "buy a managed product, or do nothing".
- For small files, chunk-level dedup is a bad trade, because the operation and metadata terms dominate and there is little repeated content to find.
Common weak answers
- "Deduplication saves 50%." That was a benchmark on iterative workloads. Quoting it as a planning number is how a cost model becomes fiction.
- Modelling storage GB-month alone. Transfer and operations usually decide the answer.
- Treating the acquisition price as the cost. The recurring cost is the team and the system they now run, and it does not stop when the migration ends.