Embedding Pipeline Service  ·  View 01 of 22  ·  Context and scope

System Context

What is inside the boundary, who uses it, and which systems stay authoritative over content and permissions.

Editable source SVG draw.io All views
People Knowledge worker Product engineer Platform engineer Tenant admin Consuming product surfaces Workspace search 12k queries/s Ask-your-docs assistant answer layer Similar documents Duplicate detection Corpora (systems of record) Document service 400 M documents File store uploads, PDFs Comment threads Connected SaaS tenant-authorised Embedding Pipeline Service corpus in, retrieval out Platform dependencies Permission authority query-time ACLs Workload identity SPIRE Observability platform Prometheus retrieval top-k neighbours scores changes bytes changes webhooks searches registers migrates erases checks mTLS OTLP Embedding Pipeline Service — System Context Person or role External / third party Security / platform synchronous event / async Out of scope: the generative answer layer, the document stores themselves, and the product user interfaces. The nightly warehouse export is omitted here and appears on view 9. v 1.0 · owner Data & AI Platform Architecture · date 2026-10

Decisions

  • The corpus stays with its owner. The platform holds a pointer, a version and a derived representation — never an authoritative copy of content or of permissions.
  • Permissions are resolved against the source authority at query time rather than copied into the index, which keeps the platform out of the business of mirroring an ACL model it does not own.
  • Four consumer surfaces, one retrieval contract. A product team integrates against the retrieval API, not against the index.

Out of scope

  • The generative answer layer: prompting, synthesis and citation rendering belong to the consuming surface.
  • The document stores themselves, and the product user interfaces.
  • Omitted from this view only: the nightly warehouse export, which appears on view 9.

Assumptions

  • 25,000 customer organisations, 3 million monthly active users, 400 million documents, 4.8 billion live chunks.
  • 12,000 retrieval queries per second at steady state; 8 million document versions per day.
  • Every figure in this set is a stated assumption, sized for a company of that shape. None is drawn from a production system.