Embedding Pipeline Service · View 17 of 22 · Operations
Decisions
- One write region. Active-active indexing would make the contract and alias state a consensus problem across regions for no retrieval benefit the latency budget needs.
- Three GPU postures: on-demand for the interactive and standard lanes, spot for bulk only, and spill from spot to on-demand on preemption.
- PostgreSQL primary and synchronous standby in different zones, because RPO 0 on the chunk ledger is the one hard durability promise in the set.
Failure domains
- Zone loss: retrieval survives on two zones; a Qdrant replica is rebuilt from snapshot rather than from the ledger.
- Spot capacity withdrawal: bulk work stops and the migration takes longer. Nothing else notices, which is the point of putting bulk there.
- Region loss: Phase 3 read standby serves retrieval at a published staleness; ingest stops until the region returns.
Assumptions
- ≈ 12 TB of index data including structures, before quantisation, for 4.8 billion 768-dimensional vectors at 2 bytes per component.
- Index snapshot restore RTO 4 hours; chunk ledger RTO 15 minutes.
- Corpus sources reach Kafka through the ingest API shown on view 9; that edge is omitted here.