Figma's database stack grew roughly 100x between 2020 and 2024. Rather than shard, the team first moved groups of related tables onto their own databases, and only later sharded horizontally. What did the cheaper structural change buy, what did it leave untouched, and when does that bill arrive?
Show the full answer Hide the answer
What it bought
Figma described the path in two posts: an April 2023 one on vertical partitioning, and a March 2024 one on horizontal sharding. The first move was to split groups of related tables, for example the ones behind files or behind organisations, onto their own Postgres databases. By the end of 2022 that was about a dozen vertical partitions behind caches and read replicas.
Vertical partitioning removes the largest separable workload from the hottest instance, and it does so without touching the shape of any query. A table group that moves to its own host takes its writes, its buffer-cache footprint, its vacuum load and its connection count with it. The application change is a connection string per group. There is no new routing layer, no shard key to choose, and no rewrite of anything that reads a single table.
The runway it bought is roughly proportional to the number of separable groups and to how lopsided the load is. When one group is 40% of the write volume, moving it gives back 40% of the headroom in weeks.
What it cost
Three things, all deferred rather than avoided.
- Cross-partition joins and transactions stop existing. Anything that spanned two groups becomes two queries plus application-side logic, or a denormalised copy that now needs keeping fresh. This is the real work, and it is paid at partition time whether or not anyone plans for it.
- The number of things to operate multiplies. A dozen databases means a dozen upgrade schedules, a dozen failover drills and a dozen capacity models.
- It does nothing for a single hot table. Partitioning by group is bounded by the largest group, and the largest group is eventually one table.
When the bill arrives
The bill arrives on the day one table exceeds one machine, and there is no vertical move left to make. That is the point Figma's second project addressed, with a service called DBProxy sitting between the application and the connection pooler, parsing queries and routing them across shards. The documented sequence is worth more than the component: they separated the logical sharding, rolled out using database views, from the physical failover, so the irreversible step was short. The first horizontally sharded table shipped in September 2023 with about ten seconds of partial write availability on the primaries and no impact on the read replicas.
The decision rule
Vertical partitioning wins while the hot set is separable; it flips the moment a single table or a single key range exceeds one host. Two signals tell you which side of that line you are on: whether the top table by write volume is more than about half the instance's load, and whether the next partition you could make would move more than about 20% of the work. When the answer to the first is yes and to the second is no, further partitioning is theatre and the sharding project has already started.
When this is the wrong answer
If you already know the product's dominant entity will be two orders of magnitude larger than anything else within a year, vertical partitioning spends six months to arrive at the same sharding project with a dozen extra databases to shard. If the constraint is read throughput rather than write throughput or working set, replicas are cheaper than either. And if the application still runs cross-group transactions that the business genuinely requires, splitting the groups converts a correctness guarantee into application code you now own, which is a real loss and sometimes a reason to buy a bigger machine instead.