Embedding Pipeline Service  ·  View 17 of 22  ·  Operations

Deployment Architecture

One write region across three zones, the node pools that separate CPU from GPU from spot, and a read standby that is Phase 3.

Editable source SVG draw.io All views
Region eu-west — primary write region Zone A Retrieval API HPA on p99 Extract + chunk CPU pool Qdrant shard replica 1 Kafka broker Zone B Retrieval API GPU node pool KubeRay, on-demand Qdrant shard replica 2 PostgreSQL primary Patroni Zone C Retrieval API GPU spot pool bulk lane only OpenSearch node PostgreSQL standby synchronous Shared regional services Object storage MinIO erasure-coded Index snapshots RTO 4 h Control plane etcd aliases Argo Workflows rebuilds OpenBao SPIRE server Telemetry Prometheus + Mimir ClickHouse quality + drift Region us-east — read standby (Phase 3) Read serving Retrieval API read only Qdrant replica snapshot shipped PostgreSQL replica async Global load balancer latency routing retrieval snapshots spills on preemption Deployment — Kubernetes, Node Pools and Failure Domains Application we own Data store Queue / topic Security / platform Interface / broker synchronous batch failure / alternate Corpus sources reach Kafka through the ingest API on view 9. The bulk lane runs only on spot GPU capacity: a cheaper rebuild is a longer dual-write window. v 1.0 · owner Data & AI Platform Architecture · date 2026-10

Decisions

  • One write region. Active-active indexing would make the contract and alias state a consensus problem across regions for no retrieval benefit the latency budget needs.
  • Three GPU postures: on-demand for the interactive and standard lanes, spot for bulk only, and spill from spot to on-demand on preemption.
  • PostgreSQL primary and synchronous standby in different zones, because RPO 0 on the chunk ledger is the one hard durability promise in the set.

Failure domains

  • Zone loss: retrieval survives on two zones; a Qdrant replica is rebuilt from snapshot rather than from the ledger.
  • Spot capacity withdrawal: bulk work stops and the migration takes longer. Nothing else notices, which is the point of putting bulk there.
  • Region loss: Phase 3 read standby serves retrieval at a published staleness; ingest stops until the region returns.

Assumptions

  • ≈ 12 TB of index data including structures, before quantisation, for 4.8 billion 768-dimensional vectors at 2 bytes per component.
  • Index snapshot restore RTO 4 hours; chunk ledger RTO 15 minutes.
  • Corpus sources reach Kafka through the ingest API shown on view 9; that edge is omitted here.