Embedding Pipeline Service · View 14 of 22 · Runtime
The trade this view makes visible
- GPU inference is dramatically cheaper per vector in large batches, and a large batch is bought with waiting. The batcher's max-wait is therefore a per-lane published number, not a tuning constant.
- The interactive lane buys freshness with a short wait and poor batch efficiency; the bulk lane buys cost with no freshness promise at all.
- Bulk runs on preemptible capacity and is preempted rather than interactive being queued behind it. That is what makes a 14-day rebuild affordable.
Guards on the fleet
- Replicas are addressed by model digest. A provider or registry serving different weights behind the same name would silently invalidate every stored vector.
- A frozen reference set is embedded on every rollout and compared against stored expected vectors. Divergence fails the rollout, not the pipeline.
- Over-length chunks are split by a declared rule and the fact is recorded on the chunk, so a truncation is never invisible.
Rebuild rates (assumptions)
- Live edit stream: 3,000 chunks/s sustained, 5× burst for 15 minutes.
- Background rebuild 4,000 chunks/s → ≈ 14 days for 4.8 billion chunks.
- Surge rebuild 18,500 chunks/s → ≈ 72 hours at about 4.5× the hourly cost. The fleet is sized on these, not on 280/s.