The cursor is a lease, not a bookmark: paginating the unbounded list
How production APIs page through collections too large for one response: offset versus keyset versus continuation tokens, what the token must encode, the consistency contract against a changing collection, and the published failures.
Every API that lists anything eventually meets a collection that does not fit in one response, and the paging scheme it picks becomes a public contract. This guide reconstructs how Kubernetes, Elasticsearch, GitLab, GitHub, Mastodon, PostgreSQL and two package registries actually page, from their own repositories, design reviews and public incident trackers: the anchored seek that replaces the offset skip, the opaque token that carries a position plus a snapshot version, the expiry that turns a cursor into a lease, and the caps the operators drew after production taught them where the cost lives. A reader finishes able to choose a scheme, specify what the cursor encodes, and defend a depth cap and a count line in review.
The expensive requests are not the deep pages: page one is the memory bomb (a 90 GB apiserver spike came from first pages, not deep ones) and the total count is the luxury the big systems quietly withdrew, with GitLab suppressing totals past 10,000 records, GitHub capping search at 1,000 results and Kubernetes never offering a count at all.
What you get out of it
- Every system that survives scale turns the page into an anchored indexed seek; offset's cost grows with depth because the skipped rows are still computed, as PostgreSQL's own manual states.
- A continuation token is a lease on a snapshot, not a bookmark: Kubernetes expires it in about 5 minutes and answers 410 Gone, and Elasticsearch requires an explicit point-in-time to keep pages consistent across a refresh.
- A cursor must encode a position and ordering, never a page number: a GitLab merge request records clients silently skipping or repeating rows because the cursor depended on the previous request's page size.
- The pagination scheme is a public contract: GitLab's 2021 production incident came from tightening an offset limit behind a feature flag, and its guidelines now call removing a pagination type a breaking change.
- Capacity-plan page one, not page N: Kubernetes measures a LIST at roughly five times the store response in temporary memory, and its remedy was to stop paginating bulk reads and stream them at a ~2 MB per-watcher constant.
Scope
Why this, now. Kubernetes spent 2023 to 2026 replacing paginated LISTs with streamed watches (KEP-3157) after list storms kept OOM-killing API servers, and GitLab closed two pagination merge requests as recently as September 2026, so the record of what paging costs at scale has never been fresher.
What it does not cover. GraphQL connection internals, client-side infinite-scroll design, and stream checkpointing; also every engineering-blog, paper and talk source, because this session's network egress reached only code hosts, registries and vendor documentation, which the page and ledger disclose.
Other field guides
Putting the services back together: ten years of Airbnb, read from its own artefacts
Reconstructs a decade of one company's architecture from artefacts rather than announcements: release timestamps on five package registries, archive …
28 sources · 4 organisations · 3 postmortemsMaking the client carry it: ten years of Discord's gateway contract
A fanout platform pays for connected consumers multiplied by events, and it cannot deploy a fix to consumers it does not own. This guide reconstructs…
28 sources · 7 organisations · 4 postmortemsThe reference architecture Netflix retired
Between 2013 and 2016 the industry copied one company's answer to service-to-service communication: discovery, load balancing, circuit breaking and c…
26 sources · 7 organisations · 4 postmortems