Silent change to offset pagination limit impacted a customer
A feature flag tightening an offset limit removed total-count and page fields from an API response; a customer's tooling broke and the flag was rolled back while they rewrote it.
How production APIs page through collections too large for one response, reconstructed from the systems' own repositories, design records and public incident trackers: Kubernetes, Elasticsearch, GitLab, GitHub, Mastodon, PostgreSQL and two package registries. A reader finishes knowing what a continuation token must encode, where offset pagination stops scaling, and why the expensive thing to serve is not the page but the total.
Reading a collection that does not fit in one response, while it changes underneath you, at a server cost that must not depend on how deep the reader has walked.
Stated without any product name, the problem is this: a client needs to read a collection that is larger than any single response the server is willing to produce. The server must answer each page in bounded time and bounded memory, whatever the page's depth. The collection keeps changing while the client walks it, so the two of them need a shared understanding of what a page even means against a moving set. And the scheme they agree on becomes a public contract: clients encode it in retry loops, billing scripts and mirrors, so changing it later is a breaking change whether or not it is called one.
The surprise of this dig is where the cost actually lives. Going in, the expected villain was the deep page, and the deep page is real: Elasticsearch refuses offset reads past 10,000 hits by default, and PostgreSQL's manual says plainly that skipped rows are still computed. But the systems' own records say the two expensive things are elsewhere. The first is page one: an unpaginated or weakly bounded first request is the memory bomb, and it is how a Kubernetes API server reached roughly 90 GB of memory and an OOM kill during one operator's 400-node scale-up (kubernetes#98423, 2021). The second is the count: GitLab serves any page you ask for but stops telling you the total beyond 10,000 records, because computing the total is the part whose cost it cannot bound (REST docs). GitHub's search API simply caps every query at 1,000 results (docs). The page is cheap. Knowing how many pages there are is not.
Scope. This guide covers list and search pagination in APIs and the storage reads behind them: offset, keyset, and token schemes, what the token encodes, the consistency contract against a changing collection, and the published failures. It does not cover GraphQL connection internals beyond where the sources touch them, client-side infinite-scroll design, or stream-processing checkpointing. One environment note, in the open: this session's network egress allowed code hosts, registries and a few documentation domains, so every source here is repository-native, a vendor document, or a public incident record; the engineering-blog, paper and talk tiers were unreachable and are absent, and the guide says below where that absence matters.
The reference shape is a seek, not a skip: every system that survives scale turns the page request into an indexed range read anchored at a position, and the argument is only about where the anchor lives.
The common architecture across every system studied has three parts. First, an anchored range read: the page is produced by seeking to a position in an index and reading forward one page, never by counting past discarded rows. GitLab's engineering guideline states the offset alternative plainly: when a client requests a large page number, the database reads page times page-size rows to return one page, which it calls unsuitable for large tables (pagination guidelines). PostgreSQL's own manual backs the mechanism: the rows skipped by an OFFSET clause are still computed inside the server (queries.sgml). Kubernetes designed its chunking directly on the store's native range read, noting that etcd3 offers consistent chunking with minimal overhead and that SQL stores can implement the same (design proposal, 2017).
Second, a position the client carries. This is where implementations
diverge, and the divergence is the decision table in the next section. Mastodon hands back
entity IDs and lets the client pass max_id or min_id to anchor the
next page (API guidelines).
GitLab's keyset mode returns a ready-made next-page URL carrying id_after.
Kubernetes wraps the position in an opaque continue token that also encodes the
snapshot version the walk started from; the design review records both why it is opaque and a
security objection, that a token embedding raw store keys could leak object names a caller is
not allowed to see
(community#896, 2017).
Third, an expiry story, because a position only means something relative to a version of the collection, and the server cannot keep old versions forever. Kubernetes is the system that says this out loud: a continue token expires after about five minutes by default, after which the server answers 410 Gone and the client must choose between starting over consistently or accepting an inconsistent continuation (API concepts). Elasticsearch makes the same state explicit as a point-in-time the client opens and must keep alive, because otherwise a refresh between pages can reorder results (paginate docs). This is the sentence the title of this guide compresses: a cursor is a lease on a snapshot, with an owner, a term and an eviction, not a bookmark in a book that holds still.
Totals and last-page links require visiting everything the filter matches. GitLab
withholds x-total, x-total-pages and the last-page link past
10,000 records; GitHub caps search at 1,000 results outright; Kubernetes never offers a
count. The npm search endpoint still returns totals, and also caps a page at 250 items on
a bounded result set.
Sources: GitLab REST docs, GitHub docs, npm registry docs
KEP-3157 measures a paginated LIST at roughly five times the store response in temporary memory, and reports losing a test server after sixteen concurrent list-building clients. Its remedy is radical: stop paging for cache-priming reads and stream objects over a watch instead, taking per-watcher cost to a roughly 2 MB constant.
Source: KEP-3157
GitLab's guideline calls removing a pagination type a breaking change, and its own 2021 incident proves the point: tightening an offset limit behind a feature flag broke a customer's integration, and the flag was rolled back while the customer rewrote their tooling.
Sources: guidelines, production#5843
Four forks, each with the condition that flips it. The recurring hinge is whether the client may anchor the next page to data it has already seen, or to a count the server must recompute.
max_id, GitLab's id_after) or an
opaque key-plus-snapshot token (Kubernetes), so the cursor's meaning cannot depend on
any earlier request's parameters.first mid-walk skipped or
repeated policies
(MR 256256,
closed unmerged; the close reason is not publicly visible).index.max_result_window probably will
not perform well and that the fix is a better pagination strategy
(issue 8754, 2018).| Decision | Chosen | Rejected | Because | Evidence |
|---|---|---|---|---|
| Paging scheme at scale | Keyset seek | Offset skip | Offset reads and discards every skipped row | GitLab guidelines |
| Cursor contents | Key plus snapshot version, opaque | Page number in the cursor | A page-number cursor depends on the previous request's page size | MR 256256 |
| Consistency of the walk | Snapshot (resourceVersion, PIT) | Live walk for bulk reads | A refresh between pages reorders results; chunking is meant to return consistent lists | community#896 |
| Depth policy | Visible refusal and capped search | Raising the result window | Depth converts into per-shard memory; a bigger window moves the cliff, not the slope | gitlab#8754 |
| Client default | Chunked list, 500 per page | Unpaginated LIST | Unbounded first pages are the documented memory bomb | kubectl get.go |
| Bulk cache priming | Streamed watch instead of paging | Paginated LIST per informer | Each LIST costs about five times the store response in temporary memory | KEP-3157 |
The public record sorts into four classes: the deep page the server pays for, the unbounded first page, the collection that moves under the walker, and the contract change that breaks the clients.
Class one, depth as a server bill, is the oldest and best documented. Every shard or index asked for page N must materialise pages one through N to find it; the Elasticsearch documentation warns that deep pages can degrade nodes outright, and GitLab.com users hit the 10,000-hit refusal as an HTTP 500 in 2018 and were still meeting it in a merge-request search path in 2026. Class two, the unbounded first page, is the one that kills servers rather than queries: both Kubernetes incidents below are page-one failures, not deep-page failures. Class three, the moving collection, is quieter: nothing crashes, rows are silently skipped or repeated, and the record of it is in cursor-design fixes rather than postmortems. Class four, the contract break, is organisational: the pagination scheme is load-bearing for clients you cannot see, and both the GitLab incident and its guideline's breaking-change rule exist because of it.
lower_relation_max_count_limit silently removed total-count and page information from billable-members API responses; a customer's tooling depended on those fields and broke.The moving-collection class deserves one more paragraph because it never files an
incident. The GitLab policy-store cursor bug is the cleanest recorded instance: the cursor
encoded a page number, so its meaning depended on the page size of the request that produced
it, and a client that changed first mid-walk skipped or repeated rows. The field
description papered over it by telling callers to keep first constant, until a
rework encoded a row offset in the cursor instead
(MR 256256, itself
closed unmerged in a stacked rework, so treat the design statement as the durable artifact).
Elasticsearch documents the same class at the engine level: without a point-in-time, a
refresh between search_after pages can reorder results across pages. Nothing alerts on a
skipped row. If the walk feeds a migration, reconciliation or billing job, this class is the
one that costs you months later, and it is invisible in every dashboard you have.
Defaults, caps and measured costs, each with its source and date. These are the lines other teams drew after getting it wrong; starting from them is cheaper than rediscovering them.
| Metric | Value | At | Context | As of | Source |
|---|---|---|---|---|---|
| Offset window ceiling | 10,000 | Elasticsearch | Default index.max_result_window; from + size beyond it is refused | 2026 | IndexSettings.java |
| Search result cap | 1,000 | GitHub | Maximum results any REST search query returns | 2026 | search docs |
| Count suppression line | 10,000 | GitLab | Past this many records, x-total, x-total-pages and the last-page link are withheld | 2026 | REST docs |
| Keyset enforcement line | 50,000 | GitLab | users endpoint requires keyset past this many requested records, since 17.0 | 2026 | REST docs |
| Continue token lifetime | ~5 min | Kubernetes | Default before a continue token answers 410 Gone | 2026 | API concepts |
| Default client page size | 500 | Kubernetes | kubectl get chunk size | 2026 | get.go |
| List memory multiplier | ~5× | Kubernetes | Temporary apiserver memory per LIST, as a multiple of the etcd response | 2023 | KEP-3157 |
| Streaming target | ~2 MB | Kubernetes | Per-watcher constant memory once lists stream over watch | 2023 | KEP-3157 |
| Incident memory spike | ~90 GB | Kubernetes operator | apiserver memory during a pending-pods list storm, ending in OOM; v1.18 | 2021 | kubernetes#98423 |
| Reproduced retention | 1.5 GiB | Kubernetes | apiserver memory held after 20 sequential ~50 MiB LIST responses | 2022 | kubernetes#114276 |
| Offset worked example | 999,980 | GitLab guideline | OFFSET at which the database reads a million rows to return twenty | 2026 | guidelines |
| Search page cap | 250 | npm registry | Maximum size per search page, default 20, offset via from; totals still served (6,364 on a sample query) | 2026 | registry API docs, live record |
| Page-number API, with counts | 3,923 | Docker Hub | Tag count for library/python; the API serves page and page_size and hands back a ready next URL | 2026-10-01 | live record |
The 90 GB and 1.5 GiB figures are operator-reported and maintainer-reproduced measurements from the Kubernetes tracker; the 5× multiplier and 2 MB constant are the KEP authors' own test figures, so treat them as measured by an interested party. The registry totals and the Docker Hub count are live values fetched on 2026-10-01 and will drift. Everything else is a shipped default, which can be changed per deployment and therefore tells you the vendor's judgement, not your workload's answer.
Every source behind this page, graded. This session's network reached code hosts, registries and vendor documentation only, so the blog, paper and talk tiers are absent by constraint rather than by judgement; the ledger shipped beside this page records what was located but unreachable.
A feature flag tightening an offset limit removed total-count and page fields from an API response; a customer's tooling broke and the flag was rolled back while they rewrote it.
An operator's production record: a 400-node upscale with thousands of pending pods spiked apiserver memory to roughly 90 GB and ended in an OOM kill, with ingress knock-on effects.
The continuation-token design: an opaque token carrying a snapshot version and a position, built on the store's native range read, returning consistent lists page by page.
The argument behind the design, recorded: tokens must be opaque to clients, and a reviewer flagged that embedding raw store keys could leak object names callers cannot see; filtered pages may legitimately return fewer items than the limit with a valid token.
Measures paginated LIST at roughly five times the etcd response in temporary memory, reports losing a test server at sixteen list-building clients, and replaces bulk LISTs with streamed watches at a ~2 MB per-watcher constant.
The internal rule set: offset reads page times page-size rows and is unsuitable for large tables, keyset reads only what it returns, removing a pagination type is a breaking change, and keyset cannot provide page numbers.
Reproduction with numbers: twenty sequential ~50 MiB LIST responses leave the apiserver holding 1.5 GiB, quantifying the serving path's allocation behaviour.
The project's own flagship client never issues an unpaginated list by default; the chunk size ships as 500 and is a flag, not a hope.
The shipped refusal point for offset paging, defined in code with the comment that the heap of hits is what it bounds.
GitLab.com search hit the 10,000 window in production (from + size of 36,920); the issue's judgement is that raising the window probably will not perform well and the real fix is a better pagination strategy.
Eight years after issue 8754, the same window surfaced as 500s in merge-request search; this capping fix was closed without merging, and the close reason is not publicly visible to this session.
Records a cursor-design bug in one sentence: a cursor that encoded a page number meant the rows it pointed at depended on the previous request's page size, so clients changing page size mid-walk skipped or repeated rows.
The registry's search API honours from and size (default 20, max 250 per its repository docs) and returns totals; a sample query returned a total of 6,364 at offset 3.
The v2 API pages by page and page_size, returns a full count (3,923 tags for library/python on the check date) and hands the client a ready-made next URL.
The full client contract: continue tokens expire after about five minutes, the server answers 410 Gone, and clients choose between a consistent restart and an inconsistent continuation the server may offer.
Why depth is memory: each shard must load the hits for all previous pages; past 10,000 use search_after, and open a point-in-time because a refresh between requests can reorder results.
Both schemes side by side: offset with page headers as the default, keyset with Link-header continuation for scale, offset deprecated on the users endpoint, keyset enforced there past 50,000 records, totals withheld past 10,000.
The Link-header contract: the server supplies next, prev, first and last URLs and the client follows them rather than computing page arithmetic; the worked example's last page is 515.
Search is framed as finding the few best matches, and the API provides up to 1,000 results per query, full stop; deep enumeration is pushed to other endpoints.
Pagination via max_id, min_id and since_id against snowflake-style IDs, with Link headers carrying next and prev; no offsets, no totals, live collection semantics.
The engine's one-sentence verdict on offset economics: the rows skipped by an OFFSET clause still have to be computed inside the server, so a large OFFSET might be inefficient.
Six rungs from an offset query to a lease-bearing cursor. The crossing from toy to production is rung four, where the collection starts moving.
Load ten million rows into PostgreSQL. Page through with LIMIT and OFFSET at page 1, page 1,000 and page 100,000, and plot latency against depth with EXPLAIN ANALYZE open.
Done when: your plot shows latency growing with offset while rows returned stays constant. Teaches: the skipped rows are computed, exactly as the manual says.
Rewrite the walk as keyset: WHERE id is greater than the last id seen, ORDER BY id, LIMIT. Re-plot the same depths.
Done when: latency is flat at any depth and the plan shows an index scan reading one page. Teaches: why GitLab calls keyset runtime independent of collection size.
Wrap the walk in an HTTP endpoint that returns a Link header with a next URL you mint, GitHub-style, rather than page arithmetic the client does.
Done when: a client can walk the whole collection knowing nothing but follow-the-link. Teaches: navigation as a server-owned contract you can re-point later.
Run inserts and deletes continuously while two clients walk: one by offset, one by keyset. Count rows each client missed or saw twice against a ground-truth snapshot.
Done when: you can state each scheme's duplicate and miss behaviour from your own numbers. Teaches: the silent failure class that never files an incident.
Mint an opaque token encoding last key, sort spec and a snapshot version. Keep a few versions of the collection; when the version is gone, answer 410 with a documented restart path. Reject nothing when the client changes page size mid-walk.
Done when: a client changing page size mid-walk neither skips nor repeats, and a stale token gets a clean 410. Teaches: the cursor as a lease, and the bug MR 256256 recorded.
Fire concurrent full-collection walks at the service while watching its memory. Then add totals to the response and measure what counting costs; cap or drop totals past a line you choose, GitLab-style, and publish the cap.
Done when: memory per request is bounded under concurrent walkers and the count line is in your API docs. Teaches: page one as the bomb, and the count as the luxury item.
The searches that found this material, copyable. The high-yield move was searching incident trackers and closed merge requests, not the technology's name.
https://gitlab.com/api/v4/projects/gitlab-com%2Fgl-infra%2Fproduction/issues?search=paginationsite:github.com kubernetes issues apiserver OOM LIST requests memory"Result window is too large" site:gitlab.comhttps://gitlab.com/api/v4/projects/gitlab-org%2Fgitlab/merge_requests?search=offset%20pagination&state=closedrepo:kubernetes/community "chunking" design pull requestgrep -rn "MAX_RESULT_WINDOW" server/src/main/java (elastic/elasticsearch)grep -rn "DefaultChunkSize" staging/src/k8s.io (kubernetes/kubernetes)curl "https://registry.npmjs.org/-/v1/search?text=anything&size=3&from=3"path:keps "pagination" OR "list" repo:kubernetes/enhancements"pagination_guidelines" OR "keyset_pagination" repo:gitlab-org/gitlab path:doc/development