Embedding Pipeline Service  ·  View 14 of 22  ·  Runtime

Embedding Pipeline

Three lanes, one GPU fleet, and the explicit trade between batch efficiency and freshness.

Editable source SVG draw.io All views
Work arrives Interactive lane p50 30 s Standard lane p95 30 min Bulk lane no SLO Deduplicate Vector cache hash x contract Hit: no GPU reverted edits, templates Batch Batcher size + max wait Wait vs cost declared per lane Infer Ray Serve autoscaled replicas Model replicas pinned by digest GPU node pool preemptible for bulk Guard Reference probe frozen chunk set Over-length rule split, recorded Quarantine typed reason Persist Cache write Index upsert contract-checked priority preempted first repeat failure gates rollout Embedding Pipeline — Lanes, Batching and GPU Economics Queue / topic Data store Opportunity Application we own Decision point Security / platform Risk / gap synchronous failure / alternate Background re-embedding runs at 4,000 chunks/s for a 14-day rebuild; surge runs at 18,500 for 72 hours at about 4.5x the hourly cost. v 1.0 · owner Data & AI Platform Architecture · date 2026-10

The trade this view makes visible

  • GPU inference is dramatically cheaper per vector in large batches, and a large batch is bought with waiting. The batcher's max-wait is therefore a per-lane published number, not a tuning constant.
  • The interactive lane buys freshness with a short wait and poor batch efficiency; the bulk lane buys cost with no freshness promise at all.
  • Bulk runs on preemptible capacity and is preempted rather than interactive being queued behind it. That is what makes a 14-day rebuild affordable.

Guards on the fleet

  • Replicas are addressed by model digest. A provider or registry serving different weights behind the same name would silently invalidate every stored vector.
  • A frozen reference set is embedded on every rollout and compared against stored expected vectors. Divergence fails the rollout, not the pipeline.
  • Over-length chunks are split by a declared rule and the fact is recorded on the chunk, so a truncation is never invisible.

Rebuild rates (assumptions)

  • Live edit stream: 3,000 chunks/s sustained, 5× burst for 15 minutes.
  • Background rebuild 4,000 chunks/s → ≈ 14 days for 4.8 billion chunks.
  • Surge rebuild 18,500 chunks/s → ≈ 72 hours at about 4.5× the hourly cost. The fleet is sized on these, not on 280/s.