The Hugging Face Hub stored large model and dataset files with Git LFS which versions whole files - appending 1 MB to a 5 GB file re-uploads 5 GB. In 2025 the Hub began moving repositories onto Xet storage which splits content into roughly 64 KB content-defined chunks aggregated into roughly 64 MB blocks so only changed chunks travel. The API developers call did not change. What forced the change and what had to stay stable and where would copying it be a mistake?
Show the full answer Hide the answer
The situation they were in
The artefacts outgrew the storage model. Hugging Face's own write-up on moving from files to chunks gives the shape of the problem: average Parquet and CSV files on the Hub in the 200 to 300 MB range, average safetensors files around 1 GB, and GGUF files exceeding 8 GB. Git versions at file granularity, so any change to a file means re-uploading the whole asset. Iterating on a model means republishing gigabytes to change metadata or one shard.
What they chose
Content-defined chunking under the same interface. The Hub's Xet documentation describes deduplication at roughly 64 KB chunks with boundaries derived from a rolling hash over content, so an insertion does not shift every subsequent boundary, and chunks aggregated into roughly 64 MB blocks for transfer. Their published example: appending 1 MB to a 5 GB file goes from re-uploading 5 GB to pushing only the new data. Their CORD-19 benchmark reports 8.9 GB stored in an LFS-backed repository against 3.52 GB in a Xet-backed one.
The platform decision is the interesting part: the storage substrate changed and the consumer contract did not. The migration was incremental — their account of the February 2025 migration describes moving a target set of repositories totalling 4.5 TB and shifting about 6% of the Hub's download traffic onto the new infrastructure — and the client library kept the same calls, so repositories worked whether or not a user had a Xet-aware client.
Why it fit their constraints
Two properties made it work. First, the interface teams depended on was already a library call rather than a storage URL, which is what gave the platform freedom to replace what was behind it. Second, the benefit is opt-in per client: an old client still works and simply does not get chunk-level savings. That converts a fleet-wide cutover into a long tail, and the number to watch becomes share of traffic on the new path rather than a cutover date.
What it cost them
A content-addressed store, a chunking client and a cache layer are now theirs to operate, and their write-up notes a block-format problem found after migration whose fix reduced GET latency by about 35% — the kind of defect that only appears under real traffic. Running two storage paths at once also means two sets of failure modes, two cost lines and a migration that ends when the last repository moves rather than when the announcement goes out.
Where copying it would be a mistake
Chunk-level deduplication pays when artefacts are large and change partially. If your artefacts are small, or each version differs from the last entirely — compiled binaries, encrypted blobs, re-quantised weights — the chunk index finds nothing to reuse and you have bought a content-addressed store, a custom client and a cache to operate for no saving. Plain object storage with immutable versioned keys is cheaper and has no client to upgrade.
The transferable lesson is narrower than the technology: a platform can replace its substrate only if consumers were never coupled to it, and the way to know is to ask whether an improvement can be shipped without asking anyone to change their code. If the answer is no, the first piece of work is the interface, not the storage.
Common weak answers
- "They should have told everyone to upgrade the client." A mandate on a community of roughly two million developers is not a migration plan; the opt-in path is what made incremental rollout possible.
- "Use a CDN." A cache lowers the cost of repeated downloads and does nothing about re-uploading 5 GB to change 1 MB.
- "Git LFS was the wrong choice originally." It was the correct cheap choice at the time. The lesson is to keep the interface stable enough that outgrowing a substrate stays survivable.