Ten years of TiDB  / field guide
Practitioner field guide · 6 October 2026

Correctness, then cost, then neighbours

PingCAP has written down nearly every substantial change to TiDB and TiKV since 2018, in 136 dated design documents and 46 accepted RFCs, and has shipped a release note for every version since October 2017. Read in order, that record shows the binding constraint moving three times, in a sequence most shared-database teams will repeat, and it shows which of the grand redesigns actually reached users.

42 primary sources 3 repositories read in full 4 severity-critical incidents Evidence through August 2026 Read: 22 min
01

The territory

One transactional database, many applications, a decade of its own engineering paperwork. The question this guide answers is not how the system works but in what order its constraints arrived, and what that order predicts for anyone building the same thing.

182
dated design documents and accepted RFCs across the SQL layer and the storage layer, 2018 to 2026
96→256 MB
default region split size, raised in v8.3.0 against a design proposal of 10 GiB
148
release-note fixes for an inconsistency involving data or indexes; 36 of them in the schema-change path
0→12
design documents about resource governance: none before 2021, four of the ten filed in 2026

The problem, stated without any product in it: you operate one strongly consistent, horizontally scaled transactional store, several teams write to it, and you cannot stop it to rebuild it. Over ten years, what actually forces redesign? TiDB is an unusually good place to answer that, because PingCAP runs a design-document process that requires a dated file in the repository before substantial work starts, and states the bar in writing: a document "is accepted or rejected when at least two committers reach consensus". The result is a decade of decisions with dates on them, next to a release-note corpus of 208 files that records what shipped, when, and what broke afterwards.

Three answers come out of that record, and the first one is the surprise. The most ambitious storage redesign of the second half of the decade did not reach users as an architecture. RFC 0082 proposed replacing the 96 MiB region with regions up to "10GiB (before compression)" as "the first step that we try to support PiB scale cluster", and RFC 0093 followed it by giving every region its own RocksDB instance. That engine was introduced as experimental in v6.6.0 in February 2023, promising to "expand the storage capacity of the cluster from TB to PB", was last mentioned in v7.4.0 in October 2023 as "part of the near-term GA of the architecture", and appears in no release note after that, through v8.5.8 of 27 August 2026. What users did get is in the shipped source: the default split size moved from 96 MB to 256 MB "in version >= 8.3.0", August 2024. The redesign landed as a constant.

The second answer is a repeated move rather than a single decision. Once quorum replication is paying for durability, every local write that duplicates it becomes negotiable, and the record shows that cashed in four times: pessimistic locks no longer replicated, a key-value write-ahead log no longer written, a raft log no longer waited on, and flow control taken out of the storage engine. Each deletion is paired with a compensation, and the compensations are where the complexity now lives. The third answer is where the work went instead. Resource governance is absent from the design record before 2021 and dominates it afterwards, from request units and keyspaces to runaway-query kills, background-task demotion and, in 2026, a per-statement limit on in-flight requests to one store. The cluster stopped being a scaling problem and became a neighbourhood.

Scope

This guide covers what the public repositories of one company say about one system between October 2017 and August 2026: design documents, accepted and rejected RFCs, release notes, shipped configuration defaults and severity-critical bug reports. It does not cover TiDB Cloud internals, which are not in the open repositories; it does not compare TiDB with its competitors; and it contains no engineering-blog posts, conference talks or customer incident reports, because this session's network reached code hosts only. Where the public record stops, the guide says so rather than filling the gap.

Figure 1 · Three eras in one design record

2023-2026 · neighbours

2021-2023 · local cost

2017-2020 · compatibility

Raft KV + Percolator
v1.0 GA, Oct 2017

Pessimistic locks become
the default, Dec 2019

Locks kept in
leader memory

Raft Engine
default, v6.1

Bigger regions, shipped
as 256 MB in v8.3

Resource groups
GA, v7.1

Runaway query
KILL, 2023

Per-query per-store
limit, 2026

2023-2026 · neighbours

2021-2023 · local cost

2017-2020 · compatibility

Raft KV + Percolator
v1.0 GA, Oct 2017

Pessimistic locks become
the default, Dec 2019

Locks kept in
leader memory

Raft Engine
default, v6.1

Bigger regions, shipped
as 256 MB in v8.3

Resource groups
GA, v7.1

Runaway query
KILL, 2023

Per-query per-store
limit, 2026

The constraint moves down the page and never moves back: compatibility first, then the local cost of the guarantee, then isolation between tenants. Dates from the release notes and the design documents.
Diagram source
02

How it is actually built

Four planes, one of which did not exist for the first five years. Every box below is attributable to a document or a source file, and the late arrival of the fourth plane is the main architectural fact of the decade.

Figure 2 · Reference architecture, with the governance plane marked

Row store, TiKV

Control plane, PD

SQL tier

kill

learner

Parser, planner,
executor

RU token bucket
per resource group

Memory arbitrator

Timestamp oracle

Region scheduler

Keyspace metadata

Unified read pool,
dmclock priority

Raft group per region,
Percolator 2PC

RocksDB +
Raft Engine

Columnar learners,
TiFlash

Row store, TiKV

Control plane, PD

SQL tier

kill

learner

Parser, planner,
executor

RU token bucket
per resource group

Memory arbitrator

Timestamp oracle

Region scheduler

Keyspace metadata

Unified read pool,
dmclock priority

Raft group per region,
Percolator 2PC

RocksDB +
Raft Engine

Columnar learners,
TiFlash

The serving path has been stable since 2017; everything on the dotted path was added from 2021 onward and is what a shared cluster actually needs. Reconstructed from the TiDB and TiKV readmes, the resource control design and the keyspace design.
Diagram source

The serving shape is the one PingCAP has described since the beginning and it has not changed. A stateless SQL tier parses and plans; a control plane called PD hands out timestamps and moves data; a row store replicates each key range as its own Raft group and runs two-phase commit over it. The TiKV readme names the lineage directly: the design "is inspired by some great distributed systems from Google, such as BigTable, Spanner, and Percolator", the transaction model "is similar to Google's Percolator with some performance improvements", and "Similar to Google's Spanner, TiKV supports externally-consistent distributed transactions". Spanner's own paper describes that property as "externally consistent reads and writes, and globally-consistent reads across the database at a timestamp". Hold on to that sentence; the last decision in this guide walks away from it.

Two components are variation points rather than parts of the core. Columnar replicas are attached to the same Raft groups as learners, which is how one system serves analytical queries without a separate pipeline, and they are optional: a cluster with no analytical workload simply has none. The storage engine under the row store is also a variation point, and an unresolved one. The original arrangement is one RocksDB instance per store holding every region's data; RFC 0093 proposed one instance per region, calling each a tablet, and the release notes carried it as experimental for four versions before going quiet. Treat the per-region engine as something to ask a vendor about rather than something to plan on.

The fourth plane is the interesting one, because it is young. Before 2021 there is nothing in the design record that bounds what one application may consume. The global resource control document of November 2022 says so plainly: "TiDB lack of global admission control limits the request from the SQL layer to the storage layer, and different users or applications may influence each other in the shared environment". What was added is a token bucket at the SQL entry, denominated in request units, and a fair queue in the storage read pool using "the dmclock algorithm to ensure the fairness between different resource groups". The same document names the commercial reason in one line: the mechanism "can be used by Resource Manage in the future of TiDB Cloud. Users can pay for the actual usage."

Admission at the SQL entry

Resource groups carry a rate in request units, a priority and a burst flag, bound to a user, a session or a single statement by hint. The quota is counted across both the SQL and the storage layer, so one number covers work the planner cannot predict.

Specified in: resource control design, 2022-11-25; GA in v7.1.0

Fairness in the storage read pool

Priority queueing inside TiKV, with one rule worth copying: background work is pinned to the lowest priority whatever its group says, because "Background tasks (GC, compaction, statistics) use LOW group_priority regardless of their resource group's configured priority". A full queue rejects with a busy error rather than growing.

Specified in: queue fairness design doc

Tenancy as a key prefix

A keyspace is a prefix on every key plus a namespace in the metadata store, so many applications share one storage cluster while each SQL instance serves exactly one keyspace. The ceiling is explicit: "The max keyspace id is 16777216, no new keyspaces can be created after the maximum keyspace id has been reached."

Specified in: keyspace design, 2022-12-07

03

The decisions that matter

Four forks, each with the rejected option and the stated reason, taken from the documents that argued them. The last column is the one to steal: the condition under which the rejected option is the right one for your system.

Decision: should the database make applications handle write conflicts, or hold locks for them?

Chosen
  • Pessimistic locking, shipped as experimental in v3.0 GA of June 2019 and made the default for new clusters six months later: "Update the default value of the tidb_txn_mode variable from "" to "pessimistic"".
  • It won because the applications being migrated were written against MySQL and expect a statement that takes a lock to keep it.
Rejected
  • Percolator's optimistic model as the only mode, which was the founding design and is faster when conflicts are rare.
  • It lost on compatibility, not on performance: a conflict surfaces at commit time, and application code that never expected to retry a whole transaction cannot absorb that.
Flips when
  • Your clients are yours. If every writer is code you control and can make idempotent and retryable, optimistic concurrency keeps the lock traffic and the latency it costs.
  • Two and a half years is how long it took this team to conclude otherwise, which is a useful estimate of how long a compatibility argument takes to win.

Decision: once a quorum guarantees durability, what may a node stop writing locally?

Chosen
  • Stop replicating pessimistic locks and keep them in leader memory (RFC 0077), stop writing the key-value write-ahead log once each region owns its LSM tree (RFC 0093), and stop waiting for the local raft write before applying (RFC 0112).
  • The justification is one sentence of Raft: "committed entries are durable and will eventually be executed by all of the available state machines". If the quorum has it, the local disk is an optimisation.
Rejected
  • Paying for durability twice, which is the default behaviour of a replicated log on top of a durable engine.
  • It lost on measurable cost: RFC 0077 "expects to reduce disk write bandwidth by 20% and reduce the latency of pessimistic locking by 50%", and RFC 0112 was written because "in the cloud environment, we have encountered performance issues with disk IO latency jitters".
Flips when
  • You cannot enumerate the compensations. Each deletion has one: locks are shipped to peers before a voluntary leader transfer and moved on split and merge; every store state migrates into the raft log once the key-value log is gone; the unpersisted apply gap is bounded and switched off for merges.
  • Also when the lost thing is not retryable. Losing a pessimistic lock is survivable because the transaction fails and the client retries; losing an acknowledged commit is not.

Decision: where do you put the limit that stops one workload hurting another?

Chosen
  • At the tenant (request units per second, priority, burst) and at the statement: kill a runaway query by elapsed execution time, and from the July 2026 design cap one statement's in-flight coprocessor requests per store, "The default is 15".
  • It won because a per-request deadline does not bound a query: "a runaway query may not cost too much time on one single coprocessor request", so the 60-second coprocessor deadline never fires.
Rejected
  • Fairness per key range. RFC pull request 121, which proposed deprioritising hot regions, was closed unmerged on 14 April 2026 after three reviewers argued against it.
  • The reasons are recorded: one reviewer proposed "query level resource control" instead, a second warned it "might be counter-productive for complex queries", and a third that "the excessive pursuit of region-level fairness in business hotspot scenarios may harm the performance of critical business operations".
Flips when
  • The hotspot is data-shaped rather than query-shaped. If one key range is hot because of what it holds rather than because of who is reading it, no statement-level limit helps and splitting or scattering that range is the only lever.
  • In general: put the limit on the thing whose owner you can bill or page. A region has no owner.

Decision: what do you give up to let two regions both accept writes?

Chosen
  • Eventual consistency between clusters, with conflicts resolved by last write wins, in the active-active design of November 2025: "This design provides eventual consistency, and any write conflicts are resolved using the Last Write Wins (LWW) strategy."
  • The document states the loss without hedging: "the solution does't provide global transactional consistency", and each cluster only converges to the same final value, not to the right one.
Rejected
  • Stretching one cluster's Raft groups across regions, which preserves the externally consistent transactions the TiKV readme still claims.
  • It lost to physics and to the application: a cross-region quorum puts inter-region latency inside every commit, and a single cluster cannot accept local writes on both sides of a partition.
Flips when
  • Conflicting writes are impossible by construction, for example when each region owns a disjoint key space by tenancy or by sharding on a geography column. Then a stretched cluster keeps the stronger guarantee at no availability cost.
  • Note what had to be invented to make last write wins work at all: a soft-delete column, because comparing a delete against an insert requires the deleted row's commit timestamp to still exist. Conflict resolution needs tombstones that outlive garbage collection.

Figure 3 · Where the limit goes

another tenant

the whole cluster

one statement

background work

one key range

memory, not CPU

Who is harmed
by the overload?

Resource group:
RU per second,
priority, burst

Where does the
cost come from?

Runaway KILL by elapsed time,
per-store in-flight cap of 15

Pin GC, compaction and
statistics to LOW priority

Split and scatter the range;
region-level fairness was
rejected in PR 121

Memory arbitrator:
subscribe before allocating,
cancel, then kill

another tenant

the whole cluster

one statement

background work

one key range

memory, not CPU

Who is harmed
by the overload?

Resource group:
RU per second,
priority, burst

Where does the
cost come from?

Runaway KILL by elapsed time,
per-store in-flight cap of 15

Pin GC, compaction and
statistics to LOW priority

Split and scatter the range;
region-level fairness was
rejected in PR 121

Memory arbitrator:
subscribe before allocating,
cancel, then kill

The tree TiDB's own argument arrived at: bill the tenant, bound the statement, demote the background, and treat a hot key range as a placement problem rather than a fairness problem. Terminal boxes are the mechanisms that exist today.
Diagram source
DecisionChosenRejectedBecauseEvidence
Concurrency controlPessimistic locking by defaultOptimistic onlyMigrated MySQL applications cannot retry whole transactionsv3.0.8 release notes, 2019-12-31
Lock durabilityLocks in leader memory, not replicatedReplicated, persisted locks20% of disk write bandwidth and half the lock latency, against a "low probability that a pessimistic lock will be lost"RFC 0077, v6.0.0-DMR
Local persistence orderingApply at committed, not at persistedWait for the local raft writeCloud disk latency jitter on one node slows every transaction that touches itRFC 0112
Write flow controlThrottle in the scheduler, disable the engine's own stallKeep RocksDB write stallA default delayed write rate of 16 MB/s "doesn't take real disk ability into account" and produces latency spikesRFC 0067; contrast Pebble
Unit of scaleLarger regions, buckets for concurrency; shipped as a 256 MB defaultKeeping 96 MiB regionsRegion count, not region size, drove RPC fan-out, heartbeat CPU and slow toolingRFC 0082, split-check config
Isolation granularityTenant and statement limitsRegion-level fairnessDeprioritising a hot region stretches the tail of every query that touches itRFC PR 121, closed 2026-04-14
Cross-region writesTwo clusters, eventual consistency, last write winsOne cluster stretched across regionsA cross-region quorum puts inter-region latency in every commitactive-active design, 2025-11-05
04

What broke in production

PingCAP does not publish customer postmortems, so the failure record has to be read from the severity-critical bug reports and from the release notes that fixed them. Read that way it is remarkably concentrated: the serving path is not where the correctness incidents are.

Across all 208 release-note files there are 148 bullets fixing an inconsistency involving data or indexes. Classifying them by trigger puts the largest group, 36 bullets, in the online schema-change path, ahead of outbound replication (21) and the transaction layer itself (21). The storage and replication core, which is the part everyone worries about in review, accounts for four. That classification is mine, by keyword, and the boundaries are arguable; the imbalance is not.

Inside that group the trigger repeats with unusual consistency. An index build is a long-running background job that reads every row, sorts, and ingests the result, and it goes wrong when the ownership of the job changes underneath it: the owner node resigns, the owner is partitioned, the cluster is upgraded mid-build, the placement driver leader is killed, or a dependency is unavailable longer than a fixed retry budget. The serving path tolerates all five events by design. The maintenance path, which was distributed later and more quickly, did not.

Figure 4 · The failure path all four reports share

DDL owner BRow storeBackfill workersDDL owner AApplicationDDL owner BRow storeBackfill workersDDL owner AApplicationowner is partitioned, upgraded orresignsALTER TABLE ... ADD INDEXsplit the backfill into subtasksscan rows, write sorted filesconcurrent INSERT and REPLACE INTOresume subtasks from stored stateingest the remaining filesDDL reported as completeADMIN CHECK TABLEError 8223, data inconsistency
DDL owner BRow storeBackfill workersDDL owner AApplicationDDL owner BRow storeBackfill workersDDL owner AApplicationowner is partitioned, upgraded orresignsALTER TABLE ... ADD INDEXsplit the backfill into subtasksscan rows, write sorted filesconcurrent INSERT and REPLACE INTOresume subtasks from stored stateingest the remaining filesDDL reported as completeADMIN CHECK TABLEError 8223, data inconsistency
Ownership of the backfill changes while concurrent writes are still arriving, and the index is silently short until somebody runs a check. Reconstructed from issue 44619 and issue 54897.
Diagram source
Incident report

The index is half there after the owner changes twice

AssumptionA background job that persists its progress can be resumed by whichever node holds the DDL ownership lease.
What happenedWith an index being built in ingest mode, the owner was made to resign at least twice and the cluster was then upgraded. The index ended up with less than half the rows of the table.
Blast radiusReported check output "table count 10,008,00 != index(idx1) count 4,666,669". Labelled severity/critical against the 7.1 long-term support line. No timing is published, because nothing alerts: the gap is found by running a check.
FixFixed in v7.1.1, one of a long sequence of fixes to the same state machine rather than a redesign of it.
Design rulePersisting a job's progress is not enough to make it resumable. The resume path must be idempotent against the exact interleaving of progress record, ownership change and in-flight writes, and the only way to know it is to inject the ownership change in test.
Incident report

Concurrent writes slip between the backfill phases

AssumptionWrites arriving during an index build are captured either by the backfill scan or by the double-write path, with no window between them.
What happenedA REPLACE INTO executed between the import and merge stages of a unique-index build left index and row out of agreement. It was found by a schema-change fuzzer with failpoint injection, not by a user.
Blast radiusSeverity critical, reported against four long-term support lines at once: 6.5, 7.1, 7.5 and 8.1.
FixCorrected in the v7.5.2, v7.1.6 and v6.5.10 patch releases.
Design ruleA multi-phase data rewrite has one hazard per phase boundary. Enumerate the boundaries, then write a test that writes at each one; a phase boundary with no test is an outage waiting for traffic.
Incident report

Distributing the build distributed the failure modes

AssumptionMoving the backfill onto a distributed execution framework makes it faster without changing its correctness envelope.
What happenedWith distributed task mode enabled, a network partition injected at the DDL owner during an index build produced the user-visible "Error 8223 (HY000): data inconsistency in table".
Blast radiusImpacts 6.5, 7.1, 7.5 and 8.1. The fix was backported into four separate release lines (6.5.11, 7.1.6, 8.1.1, 8.3.0), which is the clearest available proxy for how much of the installed base was exposed.
FixPatch to the distributed backfill, with no change to the mode's default availability.
Design ruleWhen you distribute a maintenance job, it inherits every failure mode you spent years taming in the serving path, and it inherits them without the serving path's test suite. Budget the same partition and leader-change testing for background work as for foreground work.
Incident report

A dependency down for longer than the retry budget

AssumptionA rolling restart of the storage layer is transparent to an in-flight index ingest, because the ingest retries.
What happenedWhen a storage node stayed down longer than 810 seconds during a rolling restart, the ingest stopped retrying and the index was missing key-value pairs. The title states the threshold: "after tikv is down longer than 810s".
Blast radiusSeverity critical against v7.1, v7.5 and v8.1. The inconsistency is silent until a check runs.
FixFixed in the v7.5.4, v7.1.6 and v8.1.2 patch releases.
Design ruleEvery retry budget is an undeclared availability assumption about a dependency. Write the number down, compare it against your longest routine maintenance window, and make exhaustion fail the job loudly rather than finish it quietly.
Designed-in loss

What recovery from majority loss costs, written down in advance

AssumptionThree replicas across three failure domains means majority loss does not happen, so the recovery path is a tool rather than a feature.
What happenedIt happened often enough that the old offline tools, which required stopping nodes and were "not a documented routine", were replaced by an online recovery driven from the control plane.
Blast radiusThe RFC lists what the operator loses: "Committed writes can be lost, so linear consistency of Raft can be broken" and "For a given transaction, it can be partially committed, so Data Integrity can be broken".
FixA force-leader step that lets a surviving minority commit a configuration change, demote the dead voters and re-form a quorum, without stopping the cluster.
Design ruleWrite the recovery path's data-loss semantics into the design document, not the runbook. An operator deciding at three in the morning needs to know which invariants they are about to trade, and a recovery that needs a vendor engineer is a recovery you do not have.
Designed-in risk

The control plane as a self-sustaining overload

AssumptionClients that lose contact with the control plane will retry, and the retries will be absorbed.
What happenedIn clusters with hundreds of nodes, failure conditions produce a retry storm against the single control-plane leader, which the RFC describes as a state in which it "transitions into a metastable state and cannot recover autonomously".
Blast radiusThe whole cluster, because the leader hands out the timestamps every transaction needs. No incident report is published; the RFC is the evidence that it happens.
FixA circuit breaker in front of the control-plane client, configured by error-rate threshold. Note the shipped default in the RFC is zero, meaning off until an operator turns it on.
Design ruleAnything every transaction consults needs a breaker on the client side, and a breaker that ships disabled protects nobody. Decide at design time whether the safe default is protection or compatibility, and write down which you chose.
The detection gap

None of these failures pages anyone. The symptom is error 8223 from ADMIN CHECK, a command a user has to decide to run, and the 2021 design document that added permanent in-path assertions says why that matters: inconsistency causes "Data Loss: Data written cannot be read", and "It is very difficult for support engineers to locate data index inconsistencies when they occur". Two of the four reports above were produced by PingCAP's own fuzzing and fault injection rather than by a customer, which is a credit to the practice and a warning about the record: the public corpus counts the bugs that were caught, not the ones that ran.

05

Numbers you can plan against

Every figure below comes from a design document, a release note or a shipped default, with the date it was published. Nothing here is a benchmark run for this guide.

MetricValueAtContextAs ofSource
Default region split size256 MBTiKV96 MB before v8.3.0; region max size is 1.5 times the split size2024-08split-check config
Region size proposed by the redesign10 GiBTiKVWith 128 MiB logical buckets for concurrency; shipped bucket default is 50 MB and buckets are off by default2021RFC 0082
Keys per region at the old default960,000TiKVSplit triggers on size or key count, whichever comes first2021RFC 0082
Update-index latency, async commit on7.01 msTiDB v5.0Down from 12.04 ms, a 41.7% reduction, Sysbench at 64 threads2021-04v5.0.0 release notes
Clustered index effect+39%TiDB v5.0TPC-C tpmC, vendor's own test2021-04v5.0.0 release notes
Analytical engine against Greenplum 6.15 and Spark 3.1.12 to 3×TiDB v5.0Vendor benchmark, same cluster resources, some queries 8 times faster2021-04v5.0.0 release notes
Raft log store replacement-40% I/OTiKV v6.1Also 10% less CPU, about 5% more foreground throughput, 20% lower tail latency "under certain loads"2022-06v6.1.0 release notes
In-memory pessimistic locks, projected-20% write bandwidthTiKVAnd 50% lower lock latency, "according to preliminary tests on the TPC-C workload"2021RFC 0077
In-memory pessimistic locks, as shipped-10% latencyTiDB v6.0And 10% more queries per second, under the bottleneck the feature targets2022-04v6.0.0-DMR release notes
RocksDB default delayed write rate16 MB/sTiKVThe number that made engine-level write stalls unacceptable2021RFC 0067
Coprocessor request deadline60 sTiKVPer request, which is why it never bounds a long query2023-06runaway queries design
In-flight coprocessor requests per statement per store15TiDBNew system variable; zero disables the statement-level limit2026-07per-store limiter design
Buffered transaction size limit100 MBTiDBDefault txn-total-size-limit, the constraint pipelined writes were designed to remove2024-01pipelined DML design
Maximum keyspaces per cluster16,777,216PDIdentifiers are never reused, so this is a lifetime budget rather than a concurrent one2022-12keyspace design
Leader lease9 sTiKVShipped raft_store_max_leader_lease; bounds how long a stale leader can serve a local read2026-10raftstore config
Unpersisted apply gap1,024 logsTiKVShipped max_apply_unpersisted_log_limit; the RFC's own default for the feature switch was zero, so whether it is on by default cannot be read from this file alone2026-10raftstore config
Dependency outage that truncates an index build810 sTiDBObserved during a rolling restart, before the fix2024-09issue 55808
Stated horizontal scale of the row store100+ TBTiKVReadme claim; the per-region engine promised "from TB to PB" and has not reached general availability in any release note2026-10TiKV readme
Read these carefully

Measured by the vendor, not independently: every percentage in the table comes from PingCAP's own Sysbench, TPC-C or TPC-H runs, published without the harness. Treat them as the direction of an effect, not its size on your workload. Projected versus shipped: the in-memory lock rows are the same feature described twice, once as a design estimate of 20% write bandwidth and 50% lock latency, once as a shipped claim of 10% and 10%. When a vendor publishes both, the second number is the one to plan with. Derived by this guide: the 148 and 36 bullet counts, the 0 to 12 governance-document count, and the comparison between a 10 GiB proposal and a 256 MB default; the method is in the ledger and the classification is keyword-based. Nobody has published: the price of a request unit, how many clusters run the per-region engine, how often these inconsistencies reached customers, or any latency figure for the active-active design.

06

The lesson, and the order it arrives in

One pattern explains most of the middle of the decade, and one sequence explains the whole of it. Both transfer to systems that are nothing like this one.

Figure 5 · Spend the quorum once

Quorum commit already
guarantees durability

Stop replicating
pessimistic locks

Ship locks before a
leader transfer; move
them on split and merge

Stop writing the
key-value WAL

Move every store state
into the raft log

Stop waiting for the
local raft write

Bound the apply gap;
exclude merges and
log compaction

Price: a lost lock
fails one transaction

Quorum commit already
guarantees durability

Stop replicating
pessimistic locks

Ship locks before a
leader transfer; move
them on split and merge

Stop writing the
key-value WAL

Move every store state
into the raft log

Stop waiting for the
local raft write

Bound the apply gap;
exclude merges and
log compaction

Price: a lost lock
fails one transaction

Each deletion of local work is justified by the same guarantee and paid for with a different compensation; the compensations, not the deletions, are where the remaining complexity sits. Sources: RFCs 0077, 0093 and 0112.
Diagram source

The pattern first. A replicated system pays once for durability at the quorum, and then usually pays again on every node, because the components it is built from were designed to be durable on their own. Three of the decade's measured wins come from noticing that and deleting the second payment: locks that are not replicated, a write-ahead log that is not written, a local disk write that is not waited for. The reported gains are not marginal, with the log-store replacement alone claiming 40% less write traffic and 20% lower tail latency. The discipline that makes this safe is visible in every one of those documents: each states the event that would expose the deletion, and adds a compensation for it. RFC 0077 ships locks to peers before a voluntary leader transfer; RFC 0093 migrates every piece of store state into the raft log because without the log "every state can be lost"; RFC 0112 refuses the optimisation for merge preparation and bounds how far ahead applying may run. An architect copying the move should copy the enumeration, not just the deletion. If you cannot list the events that make the duplicate load-bearing, you are not removing redundancy, you are removing a safety net you have not found yet.

Now the sequence, which is the part worth taking to a review. For five years the binding constraint was correctness and compatibility, and the design record is full of SQL surface and transaction semantics. For the next three it was the local cost of the guarantee, and the record fills with storage layout and durability. Since 2021 it has been the neighbours: request units, keyspaces, runaway kills, background demotion, a memory arbitrator, and per-statement limits on fan-out. Nothing moves back. Each era's work remains, and each new constraint arrives while the previous one is still being paid for.

The last era is the one everybody builds last and needs first. Admission control was retrofitted into a serving path that was already mature: first design document late 2022, tenant-level half generally available in May 2023, statement-level half still arriving in 2026. Meanwhile the change that would have spared the second era, a genuinely bigger unit of scale, is the one that did not ship: absent from release notes after October 2023, with the workaround it was meant to retire still on by default.

If you are building a shared transactional system, the order of your own constraints is probably correctness, then the cost of the guarantee, then isolation between tenants. The first two you will be forced into. The third you can either design in while the system is young, or retrofit later into a serving path that has no idea who is asking. The claim this guide is prepared to defend, from the record above
07

The evidence wall

Twenty-six sources, graded. Every one was read in this session from a repository clone or a fetched page. The full ledger, with the quote supporting each claim and the derived counts, is in sources.md beside this file.

Postmortem PingCAP2023-06

Issue 44619: data index inconsistent after DDL owner change

An index built in ingest mode while the owner resigned twice, followed by an upgrade, finished with under half the rows indexed. Severity critical against the 7.1 support line.

Carry forwardA resumable job needs idempotence against ownership change, not just against process restart.
pingcap/tidb issue 44619
Postmortem PingCAP2024-04

Issue 52914: inconsistency from a write between backfill phases

A REPLACE INTO landing between the import and merge stages of a unique-index build corrupted the index. Found by a schema-change fuzzer with failpoint injection, reported against four support lines at once.

Carry forwardOne hazard per phase boundary; test a write at each boundary or you have not tested the migration.
pingcap/tidb issue 52914
Postmortem PingCAP2024-07

Issue 54897: error 8223 after an injected owner partition

The distributed backfill framework plus an injected network partition at the owner produces the user-visible data inconsistency error. Backported into four release lines, which is the best available proxy for exposure.

Carry forwardDistributing a maintenance job imports every failure mode of the serving path without importing its test suite.
pingcap/tidb issue 54897
Postmortem PingCAP2024-09

Issue 55808: ingest truncated after 810 seconds of downtime

A storage node down longer than the ingest retry budget during a rolling restart left key-value pairs unwritten, and the index silently short.

Carry forwardA retry budget is an availability assumption about a dependency; compare it with your longest maintenance window.
pingcap/tidb issue 55808
Decision record PingCAP2021

TiKV RFC 0082: dynamic size region

Proposes raising the region ceiling from 96 MiB to 10 GiB with logical buckets for concurrency, naming the three costs of region count: extra RPCs per transaction, per region driving cost on each node, and tooling that loops over every region.

Carry forwardWhen the metadata unit and the concurrency unit are the same object, scaling forces you to split them.
tikv/rfcs text/0082
Decision record PingCAP2022

TiKV RFC 0093: one LSM tree per region

Gives every region its own RocksDB instance to stop cross-region compaction and lock contention, and deletes the key-value write-ahead log as a consequence, moving every store state into the raft log.

Carry forwardPhysical isolation between units of replication buys predictable maintenance, and its price is that every piece of local state has to move somewhere else.
tikv/rfcs text/0093
Decision record PingCAP2021

TiKV RFC 0077: in-memory pessimistic locks

Stops replicating pessimistic locks, on the stated ground that losing one is safe because the transaction fails and retries. Lists the compensations: ship locks before a voluntary leader transfer, move them on split and merge, bound the memory per region.

Carry forwardYou can drop a guarantee whose loss is retryable; write down the events that make it lossy before you do.
tikv/rfcs text/0077
Decision record PingCAP2024

TiKV RFC 0112: apply raft log before persistence

Raises the apply bound from the minimum of committed and persisted to committed, because cloud disk latency jitter on one node slows every distributed transaction that touches it. Excludes merge preparation and log compaction, and bounds the gap.

Carry forwardQuorum commit already gives durability; local persistence on the critical path is latency you may be able to reclaim, with named exceptions.
tikv/rfcs text/0112
Decision record PingCAP2021

TiKV RFC 0067: substitute the engine's write stall

Turns off RocksDB's own write stall and throttles at the scheduler with a token bucket, because the engine's default delayed write rate of 16 MB/s ignores the real disk and converts pressure into latency spikes.

Carry forwardBackpressure belongs where you can shape it and report it, which is above the storage engine, not inside it.
tikv/rfcs text/0067
Decision record Cockroach Labsundated

Pebble: differences from RocksDB

The same question answered the other way by another storage team: Pebble removes artificial write delays entirely, arguing they "increase write latencies without a clear benefit" in an open-loop system. TiKV's RFC cites this document and keeps throttling anyway, one layer up.

Carry forwardTwo credible teams disagree here. If your clients are open loop and cannot be slowed, deleting the delay is defensible; if they are, shaping it is.
cockroachdb/pebble docs/rocksdb.md
Decision record PingCAP2021

TiKV RFC 0091: online unsafe recovery

Replaces offline recovery tooling with a control-plane driven procedure, and states exactly which invariants an operator trades: committed writes, transaction atomicity, and the agreement between a table and its indexes.

Carry forwardPublish the data-loss semantics of your recovery path in the design, not in a runbook written during the incident.
tikv/rfcs text/0091
Decision record PingCAP2022-11

Global resource control in TiDB

Admits the absence of global admission control, then introduces request units as the single unit of account across the SQL and storage layers, with a fair queue in the storage read pool. Names the billing motive in the same document.

Carry forwardThe unit you meter becomes the unit you can limit, schedule and charge for; pick it before the cluster is shared, not after.
pingcap/tidb design, 2022-11-25
Decision record PingCAP2022-12

Proposal: keyspace

Tenancy as a key prefix plus a metadata namespace, with each SQL node bound to exactly one keyspace until restarted, and a hard ceiling of 16,777,216 never-reused identifiers.

Carry forwardPrefix-based tenancy is cheap to add late and constrains you twice: identifiers are finite and a node serves one tenant at a time.
pingcap/tidb design, 2022-12-07
Decision record PingCAP2025-11

Active-active deployment

Two or more clusters, each committing locally, replicated by change data capture, with conflicts resolved by last write wins and a soft-delete column so a delete keeps a comparable timestamp. States that global transactional consistency is not provided.

Carry forwardMulti-region write availability costs the consistency model, and last write wins needs tombstones that outlive garbage collection.
pingcap/tidb design, 2025-11-05
Decision record PingCAP2023-06

Runaway queries management

Explains why per-request deadlines do not bound a query, then adds rules on elapsed execution time with three escalating actions, the strongest being to kill the statement and watch for others like it.

Carry forwardBound the statement, not the request. Users define the rule because no absolute threshold distinguishes a runaway from a legitimate full scan.
pingcap/tidb design, 2023-06-16
Decision record PingCAP2025-04

TiDB global memory arbitrator

Replaces after-the-fact memory accounting with a subscribe-then-allocate model, ranks the three problems it is solving (out of memory, heavy garbage collection, killed sessions) and reserves killing for genuine exhaustion risk.

Carry forwardMemory is the one resource you cannot throttle after the fact, so it needs admission rather than accounting.
pingcap/tidb design, 2025-04-15
Source PingCAP2026-04

RFC pull request 121: region level isolation, closed unmerged

A proposal to deprioritise hot regions, argued over from December 2025 and closed on 14 April 2026. Three reviewers reject it in favour of query-level control, on the ground that deprioritising a region stretches the tail of every query touching it.

Carry forwardPut the limit on an object that has an owner. The rejected alternative is the clearest statement of why region-level fairness is the wrong granularity.
tikv/rfcs pull request 121
Source PingCAP2026-10

Split-check configuration defaults

The shipped answer to the region-size redesign, in a comment next to the constant: the default split size was 96 MB before v8.3.0 and 256 MB after. Buckets default to 50 MB and are off unless the newer raftstore is in use.

Carry forwardRead the defaults, not the design document. The gap between them is the part of the redesign that did not survive contact with support obligations.
tikv coprocessor config.rs
Source PingCAP2026-10

Raftstore configuration defaults

Hibernation of idle regions is still on by default in 2026, five years after the RFC that expected larger regions to make it unnecessary. The unpersisted apply gap is capped at 1,024 logs and the leader lease is nine seconds.

Carry forwardA workaround kept as a default is evidence that the architecture meant to replace it never arrived.
tikv raftstore config.rs
Source PingCAP / CNCF2026-10

TiKV readme

States the lineage (BigTable, Spanner, Percolator, Raft), the scale claim of more than 100 TB, externally consistent distributed transactions, and graduated status in the Cloud Native Computing Foundation.

Carry forwardThe readme is the standing promise. Compare it with the newest design document to find where a product's guarantees have quietly been scoped.
tikv/tikv readme
Vendor PingCAP2019-12

TiDB 3.0.8 release notes

One line in a patch release changes the concurrency model for every new cluster: the default transaction mode becomes pessimistic, six months after the feature shipped as an experiment.

Carry forwardThe most consequential architectural changes often appear as a default change in a patch note. Diff the defaults across versions before upgrading.
pingcap/docs release-3.0.8
Vendor PingCAP2021-04

TiDB 5.0 release notes

The clearest set of published numbers in the corpus: async commit cutting update-index latency from 12.04 ms to 7.01 ms, clustered index worth 39% on TPC-C, and the analytical engine benchmarked against Greenplum and Spark.

Carry forwardLatency wins of this size come from removing a round trip, not from tuning. Look for the round trip first.
pingcap/docs release-5.0.0
Vendor PingCAP2023-02

TiDB 6.6.0 and 7.4.0 release notes

The per-region storage engine enters as experimental with a promise to take clusters "from TB to PB", is described eight months later as approaching general availability, and is then absent from every subsequent release note through August 2026.

Carry forwardTrack a vendor's experimental features by their silence. A feature that stops appearing in release notes has not shipped, whatever the roadmap says.
pingcap/docs release-6.6.0
Vendor PingCAP2023-05

TiDB 7.1.0 release notes

Resource control reaches general availability and is framed as consolidation: combine applications from different systems into one cluster, "reduce the number of clusters", and lay "the foundation for multi-tenancy".

Carry forwardIsolation features are consolidation economics. The business case for them is fewer clusters, which is also the risk they create.
pingcap/docs release-7.1.0
Paper Ongaro, Ousterhout2013

In Search of an Understandable Consensus Algorithm

Read from a mirrored copy in a public repository. The sentence that underwrites a decade of local optimisation: "Raft guarantees that committed entries are durable and will eventually be executed by all of the available state machines."

Carry forwardKnow precisely which guarantee your replication protocol gives, because that is the list of local work you are entitled to delete.
papers-we-love mirror
Paper Corbett et al., Google2012

Spanner: Google's Globally-Distributed Database

The system TiKV names as its model, offering "externally consistent reads and writes, and globally-consistent reads across the database at a timestamp". The 2025 active-active design declines to provide that property between clusters.

Carry forwardExternal consistency is bought with coordination. When a design drops the coordination, read the guarantee section twice.
papers-we-love mirror
What this evidence base lacks

No engineering-blog posts, no conference talks, no customer case studies and no independent benchmarks: this session's network reached code hosts only, so anything published on a company blog, a documentation site, a conference archive or a paper repository was unreachable and has been left out rather than cited from memory. The practical effect is that every performance figure here is the vendor's own, and that the failure section is built from severity-critical bug reports instead of incident reviews with timelines and customer impact. Where a claim needed one of those, this guide says that nobody has published it.

08

Build a miniature, then productionise it

Six rungs on a three-node cluster you can run on one machine. The interesting ones are the middle two, where the exercise stops being a tutorial and starts producing evidence about your own system.

Stand up the three planes and watch the regions appear

Run one SQL node, one control-plane node and three storage nodes locally, load a few gigabytes, then list the regions and their leaders. Change the split size from the default and reload the same data.

Done when: you can state the region count for a given data size at two different split sizes.  Teaches: the metadata unit and the replication unit are the same object, which is the constraint behind RFC 0082.

Build an index under concurrent writes, then check it

Start a steady insert and update workload, add an index on a column the workload touches, and run the table check afterwards. Time the check.

Done when: you can quote how long verification takes relative to the build.  Teaches: detection of the most common correctness failure in this corpus is a command somebody has to decide to run.

Take the owner away mid-build

Repeat the previous rung, and while the backfill runs, restart or isolate the node holding the schema-change ownership. Then check the table again, and repeat with the distributed execution framework enabled and disabled.

Done when: you have run both configurations at least three times and recorded any check failure.  Teaches: why four separate severity-critical reports in this guide share one trigger, and whether your version still shares it.

Make one workload starve another, then stop it

Run a well-behaved transactional workload and a full-table scan from a second account. Measure the first workload's latency. Now put the scan in a resource group with a request-unit cap and a low priority, and add a runaway rule that kills statements over an elapsed-time threshold. Measure again.

Done when: you can show the protected workload's tail latency with and without the group.  Teaches: what admission control is worth, and that a kill action is part of the design rather than an operator's improvisation.

Price the local guarantee you are paying for

With a write-heavy workload, measure disk write bandwidth and commit latency with in-memory pessimistic locks on and off. Then kill the leader of a hot region during the run and count the transactions that fail.

Done when: you have both numbers, the saving and the failure count.  Teaches: the trade in RFC 0077 as a measurement rather than a claim, and whether your application can absorb the retries.

Lose a majority on purpose

On a throwaway cluster, stop two of three replicas of some regions, then run the online unsafe recovery procedure. Afterwards, check the tables and compare against what you wrote.

Done when: the cluster serves again and you can name which rows and indexes are wrong.  Teaches: that recovery from majority loss is a data-loss event with a documented shape, which is a different conversation with your risk owner than "we have three replicas".

09

Keep hunting

The commands and filters that produced this page. They work on any company that keeps its design record in public, which is the point: the technique outlives this particular database.

Read a decade of decisions in an hour

  • git clone --depth 1 --filter=blob:none --sparse https://github.com/pingcap/tidb.git && git sparse-checkout set docs
  • ls docs/design | sort | awk -F- '{print $1}' | uniq -c
  • grep -l "Motivation" docs/design/202[4-6]-*.md | xargs grep -A4 -h "^## Motivation"

Find the arguments that were lost

  • is:pr is:closed is:unmerged
  • repo:tikv/rfcs is:pr is:closed is:unmerged rfc
  • repo:pingcap/tidb is:issue label:severity/critical "data inconsistency"

Mine release notes as an incident record

  • grep -ri "inconsisten" releases/*.md | grep -iE "data|index" | wc -l
  • grep -ho "deprecated starting from[^.]*" releases/*.md | sort -u
  • grep -l "Partitioned Raft KV" releases/*.md | sort -V

Check what actually shipped

  • grep -rn "impl Default for Config" -A60 components/raftstore/src
  • grep -rn "ReadableSize::mb\|Default value: 0" components/ | head -40
  • pdftotext paper.pdf - | tr '\n' ' ' | grep -o "guarantees that[^.]*\."
10

References

  1. PingCAP, TiDB design documents: the process pingcap/tidb, master. Checked 2026-10-06.
  2. PingCAP, TiDB design document directory, 136 files dated 2018-07-01 to 2026-09-09 pingcap/tidb, master. Checked 2026-10-06.
  3. PingCAP, TiKV RFC directory, 46 accepted RFCs tikv/rfcs, master. Checked 2026-10-06.
  4. TiKV RFC pull request 121, region level isolation, closed unmerged 14 April 2026 tikv/rfcs. Checked 2026-10-06.
  5. TiKV RFC 0023, hibernate region tikv/rfcs, undated. Checked 2026-10-06.
  6. TiKV RFC 0067, substitute RocksDB write stall tikv/rfcs, tracking issue 10137. Checked 2026-10-06.
  7. TiKV RFC 0077, in-memory pessimistic locks tikv/rfcs, tracking issue 11452. Checked 2026-10-06.
  8. TiKV RFC 0082, dynamic size region tikv/rfcs, tracking issue 11515. Checked 2026-10-06.
  9. TiKV RFC 0091, online unsafe recovery tikv/rfcs, tracking issue 10483. Checked 2026-10-06.
  10. TiKV RFC 0093, physical isolation between regions tikv/rfcs, tracking issue 12842. Checked 2026-10-06.
  11. TiKV RFC 0112, apply raft log before persistence tikv/rfcs, tracking issue 16717. Checked 2026-10-06.
  12. TiKV RFC 0115, circuit breaker for the control plane tikv/rfcs, tracking issue pd/8678. Checked 2026-10-06.
  13. PingCAP, queue fairness in the TiKV unified read pool tikv/rfcs, undated. Checked 2026-10-06.
  14. PingCAP, proposal: reduce data inconsistencies pingcap/tidb, 22 September 2021. Checked 2026-10-06.
  15. PingCAP, global resource control in TiDB pingcap/tidb, 25 November 2022. Checked 2026-10-06.
  16. PingCAP, proposal: keyspace pingcap/tidb, 7 December 2022. Checked 2026-10-06.
  17. PingCAP, runaway queries management pingcap/tidb, 16 June 2023. Checked 2026-10-06.
  18. PingCAP, pipelined DML pingcap/tidb, 9 January 2024. Checked 2026-10-06.
  19. PingCAP, TiDB global memory arbitrator pingcap/tidb, 15 April 2025. Checked 2026-10-06.
  20. PingCAP, active-active deployment design pingcap/tidb, 5 November 2025. Checked 2026-10-06.
  21. PingCAP, TiDB TopRU pingcap/tidb, 19 January 2026. Checked 2026-10-06.
  22. PingCAP, query-level per-store coprocessor request limiter pingcap/tidb, 29 July 2026. Checked 2026-10-06.
  23. pingcap/tidb issue 44619, data index inconsistent after DDL owner change Opened 13 June 2023. Checked 2026-10-06.
  24. pingcap/tidb issue 52914, data inconsistency during adding index with replace into Opened 26 April 2024. Checked 2026-10-06.
  25. pingcap/tidb issue 54897, admin check failed after injected DDL owner network partition Opened 25 July 2024. Checked 2026-10-06.
  26. pingcap/tidb issue 55808, ingest incomplete after a storage node is down longer than 810 seconds Opened 3 September 2024. Checked 2026-10-06.
  27. TiKV split-check configuration defaults tikv/tikv, master. Checked 2026-10-06.
  28. TiKV raftstore configuration defaults tikv/tikv, master. Checked 2026-10-06.
  29. TiKV readme tikv/tikv, master. Checked 2026-10-06.
  30. TiDB readme pingcap/tidb, master. Checked 2026-10-06.
  31. Cockroach Labs, Pebble: differences from RocksDB cockroachdb/pebble, master, undated. Checked 2026-10-06.
  32. PingCAP, TiDB release notes, 208 files from v1.0 to v8.5.8 pingcap/docs, master. Checked 2026-10-06.
  33. TiDB 1.0 release notes PingCAP, 16 October 2017. Checked 2026-10-06.
  34. TiDB 3.0 GA release notes PingCAP, 28 June 2019. Checked 2026-10-06.
  35. TiDB 3.0.8 release notes PingCAP, 31 December 2019. Checked 2026-10-06.
  36. TiDB 5.0 release notes PingCAP, 7 April 2021. Checked 2026-10-06.
  37. TiDB 6.0.0-DMR release notes PingCAP, 7 April 2022. Checked 2026-10-06.
  38. TiDB 6.1.0 release notes PingCAP, 13 June 2022. Checked 2026-10-06.
  39. TiDB 6.6.0 release notes PingCAP, 20 February 2023. Checked 2026-10-06.
  40. TiDB 7.1.0 release notes PingCAP, 31 May 2023. Checked 2026-10-06.
  41. TiDB 7.4.0 release notes PingCAP, 12 October 2023. Checked 2026-10-06.
  42. TiDB 8.3.0 release notes PingCAP, 22 August 2024. Checked 2026-10-06.
  43. TiDB 8.5.8 release notes, the evidence cutoff PingCAP, 27 August 2026. Checked 2026-10-06.
  44. Diego Ongaro and John Ousterhout, In Search of an Understandable Consensus Algorithm Stanford University, draft of 7 October 2013, read from a public mirror. Checked 2026-10-06.
  45. James C. Corbett and others, Spanner: Google's Globally-Distributed Database Google, OSDI 2012, read from a public mirror. Checked 2026-10-06.