Search Indexing Service · View 14 of 21 · Runtime
Decisions
- Progress is checkpointed per source partition, so a dead worker costs minutes rather than restarting a four-hour rebuild.
- The swap is refused while the new index's lag exceeds the fast-lane budget: a correct index that is 40 seconds behind is not a safe index to serve.
- The gate is four independent detectors, because each catches a different failure — count catches a truncated rebuild, sampled diff catches a mapping change, judgement score catches a relevance regression, and a canary catches what none of them predicted.
- The gate refuses rather than escalating. Waking a human at 2am to approve a swap they cannot evaluate is not a control.
Assumptions
- Largest index 38,000,000 documents, rebuilt in ≤ 4 h at ≥ 8,000 docs/second, while query p99 rises by no more than 15%.
- Dry run reports duration, storage and cost before promotion; a rebuild above a declared cost threshold needs an explicit acknowledgement.
Risks
- The 72-hour rollback window is the real limit on how long a bad relevance change can hide. A regression discovered on day four needs a rebuild, not an alias move.