Paginating the unbounded list  / field guide
Practitioner field guide · 2026-10-01

The cursor is a lease, not a bookmark

How production APIs page through collections too large for one response, reconstructed from the systems' own repositories, design records and public incident trackers: Kubernetes, Elasticsearch, GitLab, GitHub, Mastodon, PostgreSQL and two package registries. A reader finishes knowing what a continuation token must encode, where offset pagination stops scaling, and why the expensive thing to serve is not the page but the total.

25 primary sources 9 production systems 2 incident records Evidence through October 2026 Read: 18 min
01

The territory

Reading a collection that does not fit in one response, while it changes underneath you, at a server cost that must not depend on how deep the reader has walked.

Stated without any product name, the problem is this: a client needs to read a collection that is larger than any single response the server is willing to produce. The server must answer each page in bounded time and bounded memory, whatever the page's depth. The collection keeps changing while the client walks it, so the two of them need a shared understanding of what a page even means against a moving set. And the scheme they agree on becomes a public contract: clients encode it in retry loops, billing scripts and mirrors, so changing it later is a breaking change whether or not it is called one.

The surprise of this dig is where the cost actually lives. Going in, the expected villain was the deep page, and the deep page is real: Elasticsearch refuses offset reads past 10,000 hits by default, and PostgreSQL's manual says plainly that skipped rows are still computed. But the systems' own records say the two expensive things are elsewhere. The first is page one: an unpaginated or weakly bounded first request is the memory bomb, and it is how a Kubernetes API server reached roughly 90 GB of memory and an OOM kill during one operator's 400-node scale-up (kubernetes#98423, 2021). The second is the count: GitLab serves any page you ask for but stops telling you the total beyond 10,000 records, because computing the total is the part whose cost it cannot bound (REST docs). GitHub's search API simply caps every query at 1,000 results (docs). The page is cheap. Knowing how many pages there are is not.

~90GB
kube-apiserver memory spike, then OOM, during a list storm from thousands of pending pods
10,000
Elasticsearch's default ceiling on from + size, the refusal point for offset paging
5 min
Default lifetime of a Kubernetes continue token before the server answers 410 Gone
500
kubectl's default page size: the flagship client chunks every list by default

Figure 1 · Three shapes of pagination, and what each one costs the server

reads page × size rows,
discards all but one page

reads exactly one page
via the index

reads one page, but server
must hold and expire snapshots

Client loop:
fetch page, repeat

Offset / page number
LIMIT 20 OFFSET 999980

Keyset / seek
WHERE id > last ORDER BY id

Continuation token
opaque position + snapshot

Store

Store

Store

reads page × size rows,
discards all but one page

reads exactly one page
via the index

reads one page, but server
must hold and expire snapshots

Client loop:
fetch page, repeat

Offset / page number
LIMIT 20 OFFSET 999980

Keyset / seek
WHERE id > last ORDER BY id

Continuation token
opaque position + snapshot

Store

Store

Store

The same client loop lands on three very different server bills. Offset work grows with depth, keyset work does not, and a snapshot token moves the cost into state the server must hold and expire. Sources: PostgreSQL docs, GitLab guidelines, Kubernetes chunking design.
Diagram source

Scope. This guide covers list and search pagination in APIs and the storage reads behind them: offset, keyset, and token schemes, what the token encodes, the consistency contract against a changing collection, and the published failures. It does not cover GraphQL connection internals beyond where the sources touch them, client-side infinite-scroll design, or stream-processing checkpointing. One environment note, in the open: this session's network egress allowed code hosts, registries and a few documentation domains, so every source here is repository-native, a vendor document, or a public incident record; the engineering-blog, paper and talk tiers were unreachable and are absent, and the guide says below where that absence matters.

02

How the walk is actually built

The reference shape is a seek, not a skip: every system that survives scale turns the page request into an indexed range read anchored at a position, and the argument is only about where the anchor lives.

The common architecture across every system studied has three parts. First, an anchored range read: the page is produced by seeking to a position in an index and reading forward one page, never by counting past discarded rows. GitLab's engineering guideline states the offset alternative plainly: when a client requests a large page number, the database reads page times page-size rows to return one page, which it calls unsuitable for large tables (pagination guidelines). PostgreSQL's own manual backs the mechanism: the rows skipped by an OFFSET clause are still computed inside the server (queries.sgml). Kubernetes designed its chunking directly on the store's native range read, noting that etcd3 offers consistent chunking with minimal overhead and that SQL stores can implement the same (design proposal, 2017).

Second, a position the client carries. This is where implementations diverge, and the divergence is the decision table in the next section. Mastodon hands back entity IDs and lets the client pass max_id or min_id to anchor the next page (API guidelines). GitLab's keyset mode returns a ready-made next-page URL carrying id_after. Kubernetes wraps the position in an opaque continue token that also encodes the snapshot version the walk started from; the design review records both why it is opaque and a security objection, that a token embedding raw store keys could leak object names a caller is not allowed to see (community#896, 2017).

Third, an expiry story, because a position only means something relative to a version of the collection, and the server cannot keep old versions forever. Kubernetes is the system that says this out loud: a continue token expires after about five minutes by default, after which the server answers 410 Gone and the client must choose between starting over consistently or accepting an inconsistent continuation (API concepts). Elasticsearch makes the same state explicit as a point-in-time the client opens and must keep alive, because otherwise a refresh between pages can reorder results (paginate docs). This is the sentence the title of this guide compresses: a cursor is a lease on a snapshot, with an owner, a term and an eviction, not a bookmark in a book that holds still.

Figure 2 · Anatomy of a continuation token, and the machinery behind it

snapshot gone

Continuation token (opaque)

Position
last key seen, not a page number

Snapshot version
resourceVersion, PIT id

Indexed seek
range read from position

Snapshot store
MVCC history, PIT context

Compaction / GC
evicts old versions

410 Gone
client restarts or accepts
an inconsistent walk

One page, bounded cost
at any depth

snapshot gone

Continuation token (opaque)

Position
last key seen, not a page number

Snapshot version
resourceVersion, PIT id

Indexed seek
range read from position

Snapshot store
MVCC history, PIT context

Compaction / GC
evicts old versions

410 Gone
client restarts or accepts
an inconsistent walk

One page, bounded cost
at any depth

The token carries a position and a snapshot version; the server side must be able to seek to the position, serve the snapshot, and refuse cleanly once the snapshot is gone. Reconstructed from the Kubernetes chunking design and Elasticsearch pagination docs.
Diagram source

The count is the luxury item

Totals and last-page links require visiting everything the filter matches. GitLab withholds x-total, x-total-pages and the last-page link past 10,000 records; GitHub caps search at 1,000 results outright; Kubernetes never offers a count. The npm search endpoint still returns totals, and also caps a page at 250 items on a bounded result set.

Sources: GitLab REST docs, GitHub docs, npm registry docs

Page one needs the hardest bound

KEP-3157 measures a paginated LIST at roughly five times the store response in temporary memory, and reports losing a test server after sixteen concurrent list-building clients. Its remedy is radical: stop paging for cache-priming reads and stream objects over a watch instead, taking per-watcher cost to a roughly 2 MB constant.

Source: KEP-3157

The scheme is a public contract

GitLab's guideline calls removing a pagination type a breaking change, and its own 2021 incident proves the point: tightening an offset limit behind a feature flag broke a customer's integration, and the flag was rolled back while the customer rewrote their tooling.

Sources: guidelines, production#5843

03

The decisions that matter

Four forks, each with the condition that flips it. The recurring hinge is whether the client may anchor the next page to data it has already seen, or to a count the server must recompute.

Offset and page numbers, or keyset and seek?

Chosen
  • GitLab: keyset for large collections, with runtime independent of collection size; offset deprecated on the users endpoint in 16.5 and keyset enforced there past 50,000 records in 17.0 (REST docs).
Rejected
  • Offset everywhere, the original and easiest scheme. The guideline's worked example shows the page at offset 999,980 forcing the database to read a million rows to return twenty (guidelines).
Flips when
  • The product needs jump-to-page or a page selector and the collection is bounded. Keyset cannot express page numbers, the guideline says so explicitly, so admin tables and search UIs keep offset with a depth cap.

What does the cursor encode: page number, row offset, or key?

Chosen
  • A key (Mastodon's max_id, GitLab's id_after) or an opaque key-plus-snapshot token (Kubernetes), so the cursor's meaning cannot depend on any earlier request's parameters.
Rejected
  • A page number inside the cursor. A GitLab merge request records the resulting bug in one sentence: the rows a cursor pointed at depended on the page size of the previous request, so a client that changed first mid-walk skipped or repeated policies (MR 256256, closed unmerged; the close reason is not publicly visible).
Flips when
  • The store cannot seek by key (no usable index, or ranking is query-relative, as in search engines). Then the cursor carries an offset and the server must cap depth, which is exactly the Elasticsearch shape.

Does the walk see a frozen collection or a live one?

Chosen
  • Kubernetes: frozen. Chunking is intended to return consistent lists, so the token pins a resource version and the whole walk reads one snapshot (community#896).
  • Elasticsearch: frozen on request, via an explicit point-in-time (docs).
Rejected
  • A live walk for cache-priming reads. Mastodon pages a live timeline by ID, which is fine for feeds; the same choice for a migration or audit silently misses rows that move.
Flips when
  • Duplicates and misses are tolerable (human-paced feeds, infinite scroll). Snapshots cost server state and expire; feeds should not pay for consistency nobody observes.

Serve any depth, or refuse past a line?

Chosen
  • Refuse, visibly. Elasticsearch rejects from + size beyond 10,000 by default and points at search_after; GitHub caps search at 1,000 results; GitLab suppresses totals past 10,000 records.
Rejected
  • Raising the window. GitLab's own issue tracker, hitting the limit in production, judged that a higher index.max_result_window probably will not perform well and that the fix is a better pagination strategy (issue 8754, 2018).
Flips when
  • The collection is small and bounded by construction. npm's search serves offsets with totals, at a 250-item page cap, and that is a reasonable contract for result sets that top out in the thousands.

Figure 3 · Choosing a pagination scheme for a new API

whole collection

browse

yes

no

yes

no

collection outgrows the cap

Do clients read the whole
collection, or browse it?

Must the walk see
one consistent version?

Is jump-to-page or a total
a product requirement?

Opaque cursor: key + snapshot version,
with documented expiry and a 410-style
restart path

Keyset by unique indexed key,
server-minted next links,
no totals

Offset with page numbers,
hard depth cap and a
documented count line

whole collection

browse

yes

no

yes

no

collection outgrows the cap

Do clients read the whole
collection, or browse it?

Must the walk see
one consistent version?

Is jump-to-page or a total
a product requirement?

Opaque cursor: key + snapshot version,
with documented expiry and a 410-style
restart path

Keyset by unique indexed key,
server-minted next links,
no totals

Offset with page numbers,
hard depth cap and a
documented count line

Terminal nodes are actions. The tree compresses the four decisions above; note that two branches end in refusing depth rather than serving it, which is what the production systems in the evidence wall converged on.
Diagram source
DecisionChosenRejectedBecauseEvidence
Paging scheme at scaleKeyset seekOffset skipOffset reads and discards every skipped rowGitLab guidelines
Cursor contentsKey plus snapshot version, opaquePage number in the cursorA page-number cursor depends on the previous request's page sizeMR 256256
Consistency of the walkSnapshot (resourceVersion, PIT)Live walk for bulk readsA refresh between pages reorders results; chunking is meant to return consistent listscommunity#896
Depth policyVisible refusal and capped searchRaising the result windowDepth converts into per-shard memory; a bigger window moves the cliff, not the slopegitlab#8754
Client defaultChunked list, 500 per pageUnpaginated LISTUnbounded first pages are the documented memory bombkubectl get.go
Bulk cache primingStreamed watch instead of pagingPaginated LIST per informerEach LIST costs about five times the store response in temporary memoryKEP-3157
04

What broke in production

The public record sorts into four classes: the deep page the server pays for, the unbounded first page, the collection that moves under the walker, and the contract change that breaks the clients.

Class one, depth as a server bill, is the oldest and best documented. Every shard or index asked for page N must materialise pages one through N to find it; the Elasticsearch documentation warns that deep pages can degrade nodes outright, and GitLab.com users hit the 10,000-hit refusal as an HTTP 500 in 2018 and were still meeting it in a merge-request search path in 2026. Class two, the unbounded first page, is the one that kills servers rather than queries: both Kubernetes incidents below are page-one failures, not deep-page failures. Class three, the moving collection, is quieter: nothing crashes, rows are silently skipped or repeated, and the record of it is in cursor-design fixes rather than postmortems. Class four, the contract break, is organisational: the pagination scheme is load-bearing for clients you cannot see, and both the GitLab incident and its guideline's breaking-change rule exist because of it.

Figure 4 · The lease expires: a consistent walk meets compaction

"Store (MVCC)""API server""Client""Store (MVCC)""API server""Client"processes slowly,more than 5 minutescompaction evictshistory before vrestart from scratch for consistency,or continue against the latest versionand accept missed or repeated itemsLIST limit=500range read at version v500 items + continue token(position, v)LIST continue=tokenseek at version vversion no longer available410 Gone
"Store (MVCC)""API server""Client""Store (MVCC)""API server""Client"processes slowly,more than 5 minutescompaction evictshistory before vrestart from scratch for consistency,or continue against the latest versionand accept missed or repeated itemsLIST limit=500range read at version v500 items + continue token(position, v)LIST continue=tokenseek at version vversion no longer available410 Gone
A slow walker outlives its snapshot and must choose between restarting and accepting inconsistency. Reconstructed from the Kubernetes API concepts documentation.
Diagram source
Postmortem

Ninety gigabytes to list the backlog

AssumptionThe API server could absorb whatever list traffic a large scale-up produced; memory limits and fairness controls would hold.
What happenedA 400-node upscale with thousands of pending pods (a 2,000-replica deployment requesting 4 CPU and 4 GB each) drove list traffic that spiked kube-apiserver memory to roughly 90 GB, ending in an OOM kill, with knock-on ingress failures through cilium and contour with envoy.
Blast radiusControl-plane outage on the affected cluster within 30 to 60 minutes of the scale event; v1.18.14.
FixThe structural fixes landed upstream over years: bounded list work and, in KEP-3157, replacing bulk LISTs with streamed watches at a roughly 2 MB per-watcher constant.
Design ruleCapacity-plan the first page, not the deep page: every list consumer is a multiplier on response size, and the server pays about five times the store response per list in flight.
Sourcekubernetes#98423, 2021, mechanism in KEP-3157
Postmortem

The offset limit that broke a customer

AssumptionTightening an internal pagination limit behind a feature flag was a safe performance change, invisible to well-behaved API clients.
What happenedThe flag lower_relation_max_count_limit silently removed total-count and page information from billable-members API responses; a customer's tooling depended on those fields and broke.
Blast radiusOne known paying customer's integration; declared as a production incident on 2021-11-02 and the flag disabled while they rewrote their tooling.
FixRollback first; the lasting change is doctrinal, with the development guidelines naming any removal of a pagination type a breaking change to be announced, not flagged in quietly.
Design ruleCounts, page links and limits are part of the public contract. Clients build on whatever the response carries, so changing pagination follows the same deprecation path as removing an endpoint.
Source

Twenty lists, one and a half gigabytes

AssumptionServing a large response costs its size, briefly; memory returns to baseline between requests.
What happenedTwenty sequential LIST requests of about 50 MiB each left the API server holding 1.5 GiB, the serving path allocating far more than the response size and the runtime not returning it promptly.
Blast radiusA reproducible resource bug rather than an outage; it quantifies the mechanism behind the 90 GB incident.
FixBounded list processing upstream and the KEP-3157 streaming path; the issue is closed.
Design ruleMeasure list-path memory as a multiple of response size, not as the response size; five times is the documented Kubernetes figure.
Source

Page 1,846 returns a 500

AssumptionIf the UI offers a last-page button, the backend can serve it.
What happenedJumping to the last page of GitLab.com search results asked Elasticsearch for from + size of 36,920 against a 10,000 window; the indexer's 400 surfaced to users as a 500. A capping fix attempted in 2026 was closed without merging, and the same window still reaches users through newer search paths.
Blast radiusErrors for any user paging deep into search, from 2018 reports through a 2026 merge-request search path.
FixDirection recorded in the issue: an efficient pagination strategy rather than a bigger window, since raising the window probably will not perform well.
Design ruleNever render navigation the backend refuses to serve; cap the UI at the same line as the store, and prefer search_after-style continuation past it.

Figure 5 · The collection moves under an offset walker

"Collection""Server""Offset walker""Collection""Server""Offset walker"the row that moved intoposition 20 was never returned:one record silently skippedpage 1 (rows 1 to 20)rows 1 to 20row 7 is deleted, sorows 8 and later shift left by onepage 2 (OFFSET 20)rows that are now 21 to 40
"Collection""Server""Offset walker""Collection""Server""Offset walker"the row that moved intoposition 20 was never returned:one record silently skippedpage 1 (rows 1 to 20)rows 1 to 20row 7 is deleted, sorows 8 and later shift left by onepage 2 (OFFSET 20)rows that are now 21 to 40
A row deleted on an earlier page shifts every later row left by one, so the walker's next page silently skips a row it never saw. Keyset anchors to a key rather than a position, so the same deletion costs it nothing. Mechanism per the GitLab pagination guidelines and Elasticsearch pagination docs.
Diagram source

The moving-collection class deserves one more paragraph because it never files an incident. The GitLab policy-store cursor bug is the cleanest recorded instance: the cursor encoded a page number, so its meaning depended on the page size of the request that produced it, and a client that changed first mid-walk skipped or repeated rows. The field description papered over it by telling callers to keep first constant, until a rework encoded a row offset in the cursor instead (MR 256256, itself closed unmerged in a stacked rework, so treat the design statement as the durable artifact). Elasticsearch documents the same class at the engine level: without a point-in-time, a refresh between search_after pages can reorder results across pages. Nothing alerts on a skipped row. If the walk feeds a migration, reconciliation or billing job, this class is the one that costs you months later, and it is invisible in every dashboard you have.

05

Numbers you can plan against

Defaults, caps and measured costs, each with its source and date. These are the lines other teams drew after getting it wrong; starting from them is cheaper than rediscovering them.

MetricValueAtContextAs ofSource
Offset window ceiling10,000ElasticsearchDefault index.max_result_window; from + size beyond it is refused2026IndexSettings.java
Search result cap1,000GitHubMaximum results any REST search query returns2026search docs
Count suppression line10,000GitLabPast this many records, x-total, x-total-pages and the last-page link are withheld2026REST docs
Keyset enforcement line50,000GitLabusers endpoint requires keyset past this many requested records, since 17.02026REST docs
Continue token lifetime~5 minKubernetesDefault before a continue token answers 410 Gone2026API concepts
Default client page size500Kuberneteskubectl get chunk size2026get.go
List memory multiplier~5×KubernetesTemporary apiserver memory per LIST, as a multiple of the etcd response2023KEP-3157
Streaming target~2 MBKubernetesPer-watcher constant memory once lists stream over watch2023KEP-3157
Incident memory spike~90 GBKubernetes operatorapiserver memory during a pending-pods list storm, ending in OOM; v1.182021kubernetes#98423
Reproduced retention1.5 GiBKubernetesapiserver memory held after 20 sequential ~50 MiB LIST responses2022kubernetes#114276
Offset worked example999,980GitLab guidelineOFFSET at which the database reads a million rows to return twenty2026guidelines
Search page cap250npm registryMaximum size per search page, default 20, offset via from; totals still served (6,364 on a sample query)2026registry API docs, live record
Page-number API, with counts3,923Docker HubTag count for library/python; the API serves page and page_size and hands back a ready next URL2026-10-01live record
Read these carefully

The 90 GB and 1.5 GiB figures are operator-reported and maintainer-reproduced measurements from the Kubernetes tracker; the 5× multiplier and 2 MB constant are the KEP authors' own test figures, so treat them as measured by an interested party. The registry totals and the Docker Hub count are live values fetched on 2026-10-01 and will drift. Everything else is a shipped default, which can be changed per deployment and therefore tells you the vendor's judgement, not your workload's answer.

06

The evidence wall

Every source behind this page, graded. This session's network reached code hosts, registries and vendor documentation only, so the blog, paper and talk tiers are absent by constraint rather than by judgement; the ledger shipped beside this page records what was located but unreachable.

Postmortem GitLab2021-11

Silent change to offset pagination limit impacted a customer

A feature flag tightening an offset limit removed total-count and page fields from an API response; a customer's tooling broke and the flag was rolled back while they rewrote it.

Carry forwardWhatever your pagination responses carry today is the contract; tightening a limit is a deprecation, not a tweak.
gitlab.com/gitlab-com/gl-infra/production/-/issues/5843
Postmortem Kubernetes operator2021-01

kube-apiserver high memory usage on pending pods storm

An operator's production record: a 400-node upscale with thousands of pending pods spiked apiserver memory to roughly 90 GB and ended in an OOM kill, with ingress knock-on effects.

Carry forwardList traffic scales with cluster churn, not cluster size; size control-plane memory for the storm, not the steady state.
github.com/kubernetes/kubernetes/issues/98423
Design record Kubernetes2017-08

Design proposal: consistent API chunking

The continuation-token design: an opaque token carrying a snapshot version and a position, built on the store's native range read, returning consistent lists page by page.

Carry forwardBuild the token on what the store can seek; a token the store cannot serve cheaply is a promise the server will break.
github.com/kubernetes/design-proposals-archive/blob/main/api-machinery/api-chunking.md
Design record Kubernetes2017-08

community#896: the chunking design review

The argument behind the design, recorded: tokens must be opaque to clients, and a reviewer flagged that embedding raw store keys could leak object names callers cannot see; filtered pages may legitimately return fewer items than the limit with a valid token.

Carry forwardA cursor is attacker-visible state; audit what it discloses, and let pages run short rather than scanning to fill them.
github.com/kubernetes/community/pull/896
Design record Kubernetes2023, checked 2026

KEP-3157: streaming lists over watch

Measures paginated LIST at roughly five times the etcd response in temporary memory, reports losing a test server at sixteen list-building clients, and replaces bulk LISTs with streamed watches at a ~2 MB per-watcher constant.

Carry forwardWhen every client walks the whole collection, pagination is the wrong primitive; ship the collection as a stream and let clients build state incrementally.
github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/3157-watch-list/README.md
Design record GitLabchecked 2026-10

Database pagination guidelines

The internal rule set: offset reads page times page-size rows and is unsuitable for large tables, keyset reads only what it returns, removing a pagination type is a breaking change, and keyset cannot provide page numbers.

Carry forwardWrite the offset-to-keyset rules down before the first big table, because the migration later is a breaking change by your own definition.
gitlab.com/gitlab-org/gitlab/-/blob/master/doc/development/database/pagination_guidelines.md
Source Kubernetes2022-12

kubernetes#114276: memory held after large LISTs

Reproduction with numbers: twenty sequential ~50 MiB LIST responses leave the apiserver holding 1.5 GiB, quantifying the serving path's allocation behaviour.

Carry forwardBenchmark the list path by memory multiple, not response size; the gap is the capacity you actually need.
github.com/kubernetes/kubernetes/issues/114276
Source GitLab2018-12

gitlab#8754: jumping to the last page gives HTTP 500

GitLab.com search hit the 10,000 window in production (from + size of 36,920); the issue's judgement is that raising the window probably will not perform well and the real fix is a better pagination strategy.

Carry forwardWhen a backend refuses a depth, the frontend must stop offering it; the pair of them is the product.
gitlab.com/gitlab-org/gitlab/-/issues/8754
Source GitLab2026-09, closed unmerged

MR 251701: cap search pagination at the result window

Eight years after issue 8754, the same window surfaced as 500s in merge-request search; this capping fix was closed without merging, and the close reason is not publicly visible to this session.

Carry forwardA depth limit discovered in production will resurface on every new search path until the cap lives in one shared layer.
gitlab.com/gitlab-org/gitlab/-/merge_requests/251701
Source GitLab2026-09, closed unmerged

MR 256256: encode an offset in pagination cursors

Records a cursor-design bug in one sentence: a cursor that encoded a page number meant the rows it pointed at depended on the previous request's page size, so clients changing page size mid-walk skipped or repeated rows.

Carry forwardA cursor must be self-contained: position plus ordering, never an index into someone else's arithmetic.
gitlab.com/gitlab-org/gitlab/-/merge_requests/256256
Source npmfetched 2026-10-01

Live search endpoint: offset paging with totals, capped pages

The registry's search API honours from and size (default 20, max 250 per its repository docs) and returns totals; a sample query returned a total of 6,364 at offset 3.

Carry forwardOffset with counts is a fine contract when the result set is bounded by construction; the cap is what makes it safe.
registry.npmjs.org/-/v1/search?text=kubernetes&size=3&from=3
Source Docker Hubfetched 2026-10-01

Live tags endpoint: page numbers with a served next URL

The v2 API pages by page and page_size, returns a full count (3,923 tags for library/python on the check date) and hands the client a ready-made next URL.

Carry forwardIf you keep page numbers, serve the navigation yourself; a next URL you mint is a contract you can later re-point at a cursor.
hub.docker.com/v2/repositories/library/python/tags/?page_size=2
Vendor doc Kuberneteschecked 2026-10

API concepts: pagination semantics and 410 Gone

The full client contract: continue tokens expire after about five minutes, the server answers 410 Gone, and clients choose between a consistent restart and an inconsistent continuation the server may offer.

Carry forwardDesign the expiry answer before shipping the token: what a client does on 410 is part of your API, not their problem.
github.com/kubernetes/website/blob/main/content/en/docs/reference/using-api/api-concepts.md
Vendor doc GitLabchecked 2026-10

REST API pagination

Both schemes side by side: offset with page headers as the default, keyset with Link-header continuation for scale, offset deprecated on the users endpoint, keyset enforced there past 50,000 records, totals withheld past 10,000.

Carry forwardRun both schemes during any migration; the house rule that removal is a breaking change came from a real incident.
gitlab.com/gitlab-org/gitlab/-/blob/master/doc/api/rest/_index.md
Vendor doc GitHubchecked 2026-10

Using pagination in the REST API

The Link-header contract: the server supplies next, prev, first and last URLs and the client follows them rather than computing page arithmetic; the worked example's last page is 515.

Carry forwardHand clients URLs, not formulas; every piece of pagination arithmetic a client does is a compatibility promise you did not mean to make.
github.com/github/docs/blob/main/content/rest/using-the-rest-api/using-pagination-in-the-rest-api.md
Vendor doc GitHubchecked 2026-10

REST search API: the 1,000-result cap

Search is framed as finding the few best matches, and the API provides up to 1,000 results per query, full stop; deep enumeration is pushed to other endpoints.

Carry forwardSeparate search from enumeration in the API surface; ranking justifies a hard cap that listing cannot live with.
github.com/github/docs/blob/main/content/rest/search/search.md
Vendor doc Mastodonchecked 2026-10

API guidelines: cursoring by entity ID

Pagination via max_id, min_id and since_id against snowflake-style IDs, with Link headers carrying next and prev; no offsets, no totals, live collection semantics.

Carry forwardFor feeds, the entity ID is the cursor you already have; totals and page numbers are costs a timeline never needs to pay.
github.com/mastodon/documentation/blob/main/content/en/api/guidelines.md
Vendor doc PostgreSQLchecked 2026-10

Queries: LIMIT and OFFSET

The engine's one-sentence verdict on offset economics: the rows skipped by an OFFSET clause still have to be computed inside the server, so a large OFFSET might be inefficient.

Carry forwardOFFSET is a filter applied after the work is done; any scheme built on it inherits that arithmetic at every depth.
github.com/postgres/postgres/blob/master/doc/src/sgml/queries.sgml
07

Build a miniature, then productionise it

Six rungs from an offset query to a lease-bearing cursor. The crossing from toy to production is rung four, where the collection starts moving.

Feel the offset slope

Load ten million rows into PostgreSQL. Page through with LIMIT and OFFSET at page 1, page 1,000 and page 100,000, and plot latency against depth with EXPLAIN ANALYZE open.

Done when: your plot shows latency growing with offset while rows returned stays constant.  Teaches: the skipped rows are computed, exactly as the manual says.

Seek instead of skip

Rewrite the walk as keyset: WHERE id is greater than the last id seen, ORDER BY id, LIMIT. Re-plot the same depths.

Done when: latency is flat at any depth and the plan shows an index scan reading one page.  Teaches: why GitLab calls keyset runtime independent of collection size.

Expose it as an API with served navigation

Wrap the walk in an HTTP endpoint that returns a Link header with a next URL you mint, GitHub-style, rather than page arithmetic the client does.

Done when: a client can walk the whole collection knowing nothing but follow-the-link.  Teaches: navigation as a server-owned contract you can re-point later.

Make the collection move

Run inserts and deletes continuously while two clients walk: one by offset, one by keyset. Count rows each client missed or saw twice against a ground-truth snapshot.

Done when: you can state each scheme's duplicate and miss behaviour from your own numbers.  Teaches: the silent failure class that never files an incident.

Issue a real cursor, and let it expire

Mint an opaque token encoding last key, sort spec and a snapshot version. Keep a few versions of the collection; when the version is gone, answer 410 with a documented restart path. Reject nothing when the client changes page size mid-walk.

Done when: a client changing page size mid-walk neither skips nor repeats, and a stale token gets a clean 410.  Teaches: the cursor as a lease, and the bug MR 256256 recorded.

Load-test page one and ration the count

Fire concurrent full-collection walks at the service while watching its memory. Then add totals to the response and measure what counting costs; cap or drop totals past a line you choose, GitLab-style, and publish the cap.

Done when: memory per request is bounded under concurrent walkers and the count line is in your API docs.  Teaches: page one as the bomb, and the count as the luxury item.

08

Keep hunting

The searches that found this material, copyable. The high-yield move was searching incident trackers and closed merge requests, not the technology's name.

Incident records on public trackers

  • https://gitlab.com/api/v4/projects/gitlab-com%2Fgl-infra%2Fproduction/issues?search=pagination
  • site:github.com kubernetes issues apiserver OOM LIST requests memory
  • "Result window is too large" site:gitlab.com

Rejected and recorded arguments

  • https://gitlab.com/api/v4/projects/gitlab-org%2Fgitlab/merge_requests?search=offset%20pagination&state=closed
  • repo:kubernetes/community "chunking" design pull request

The mechanism in the source tree

  • grep -rn "MAX_RESULT_WINDOW" server/src/main/java (elastic/elasticsearch)
  • grep -rn "DefaultChunkSize" staging/src/k8s.io (kubernetes/kubernetes)
  • curl "https://registry.npmjs.org/-/v1/search?text=anything&size=3&from=3"

Design records before the code

  • path:keps "pagination" OR "list" repo:kubernetes/enhancements
  • "pagination_guidelines" OR "keyset_pagination" repo:gitlab-org/gitlab path:doc/development
09

References

  1. GitLab, production incident 5843: silent change to offset pagination limit impacted a customer gitlab.com incident tracker, 2021-11-02. Checked 2026-10-01.
  2. Kubernetes issue 98423: kube-apiserver high memory usage on pending pods storm github.com, 2021-01-26. Checked 2026-10-01.
  3. Kubernetes issue 114276: apiserver builds up high memory usage after serving a few large LIST requests github.com, 2022-12-04. Checked 2026-10-01.
  4. Kubernetes design proposal: consistent API chunking kubernetes/design-proposals-archive, merged 2017. Checked 2026-10-01.
  5. kubernetes/community pull request 896: design for consistent API chunking github.com, merged 2017-08-29. Checked 2026-10-01.
  6. KEP-3157: watch list, streaming lists over watch kubernetes/enhancements. Checked 2026-10-01.
  7. Kubernetes documentation source: API concepts, list pagination and 410 Gone semantics kubernetes/website. Checked 2026-10-01.
  8. kubectl get source: default chunk size 500 kubernetes/kubernetes, master. Checked 2026-10-01.
  9. Elasticsearch IndexSettings.java: index.max_result_window default 10,000 elastic/elasticsearch, main. Checked 2026-10-01.
  10. Elasticsearch documentation source: paginate search results elastic/elasticsearch, main. Checked 2026-10-01.
  11. GitLab issue 8754: Elasticsearch results jumping to last page gives HTTP 500 gitlab.com, 2018-12-07. Checked 2026-10-01.
  12. GitLab merge request 251701: cap search pagination at the Elasticsearch result window (closed unmerged) gitlab.com, closed 2026-09-18. Checked 2026-10-01.
  13. GitLab merge request 256256: encode an offset in policy store pagination cursors (closed unmerged) gitlab.com, closed 2026-09-18. Checked 2026-10-01.
  14. GitLab documentation source: REST API, offset and keyset pagination gitlab-org/gitlab, master. Checked 2026-10-01.
  15. GitLab development guidelines: pagination gitlab-org/gitlab, master. Checked 2026-10-01.
  16. GitHub documentation source: using pagination in the REST API github/docs, main. Checked 2026-10-01.
  17. GitHub documentation source: REST search API, 1,000-result cap github/docs, main. Checked 2026-10-01.
  18. Mastodon documentation source: API guidelines, pagination mastodon/documentation, main. Checked 2026-10-01.
  19. PostgreSQL documentation source: LIMIT and OFFSET (queries.sgml) postgres/postgres, master. Checked 2026-10-01.
  20. kubectl get.go, raw file as fetched raw.githubusercontent.com. Checked 2026-10-01.
  21. npm registry API documentation: search endpoint parameters npm/registry, main. Checked 2026-10-01.
  22. npm registry live search record (from and size honoured, total returned) registry.npmjs.org. Fetched 2026-10-01.
  23. Docker Hub live tags record (page_size, count and next URL) hub.docker.com. Fetched 2026-10-01.