Enterprise Metadata Management System  ·  View 02 of 22  ·  Context and scope

High-Level Architecture

The six stages a piece of metadata passes through, from a source system to a person who trusts it.

Editable source SVG draw.io All views
Data & tool estate
Data & tool estate
Warehouses & lakes
Warehouses & lakes
BI & reporting
BI & reporting
Pipelines & jobs
Pipelines & jobs
SaaS & apps
SaaS & apps
Harvest
Harvest
Connector Fleet
40 types
Connector Fleet...
Ingestion Agents
run in-VPC
Ingestion Agents...
Event Listeners
webhook · CDC
Event Listeners...
Harvest Orchestrator
watermarks
Harvest Orchestrator...
Process & enrich
Process & enrich
Canonical Mapper
Canonical Mapper
Classify & Profile
rules + ML
Classify & Profile...
Lineage Parser
column level
Lineage Parser...
Merge & Precedence
curated wins
Merge & Precedence...
Metadata repository
Metadata repository
Aspect Store
PostgreSQL
Aspect Store...
Knowledge Graph
Neo4j
Knowledge Graph...
Search Index
OpenSearch
Search Index...
Payload Archive
S3 · replay
Payload Archive...
Access
Access
GraphQL API
primary read
GraphQL API...
REST API
write · bulk
REST API...
Event Egress
Kafka · webhook
Event Egress...
Policy Sync
outbound tags
Policy Sync...
Experience
Experience
Catalog & Search
Catalog & Search
Glossary Workbench
Glossary Workbench
Lineage Explorer
Lineage Explorer
Governance Console
Governance Console
High-Level Architecture — Source to Consumption
High-Level Architecture — Source to Consumption
External / third party
External / third party
Interface / broker
Interface / broker
Queue / topic
Queue / topic
Application we own
Application we own
Data store
Data store
Identity, workflow, audit, notification and observability are cross-cutting; see views 03 and 04.
Identity, workflow, audit, notification and observability are cross-cutting; see views 03 and 04.
v 1.0 · owner Data & AI Architecture · date 2026-08
v 1.0 · owner Data & AI Architecture · date 2026-08
Text is not SVG - cannot display

Why this shape

  • Harvest, process, store, serve and present are separated so a connector failure degrades freshness for one source rather than availability for everyone.
  • The repository is deliberately four stores, not one: no single engine does versioned writes, graph traversal, ranked search and cheap archival well (view 08).
  • GraphQL is the primary read interface because the model is a graph; REST carries writes and bulk operations where GraphQL adds nothing.

Targets

  • 5 million assets, 50 million relationships at design scale.
  • Search p95 under 300 ms; asset profile p95 under 500 ms; 3-hop lineage p95 under 1.5 s.
  • 99.9% availability on the read path; ingest is allowed to lag at 99.5%.

Left out here

  • Identity, workflow, audit, notification and observability cut across every stage; drawing them here would obscure the flow. They are in views 03, 04, 13, 18 and 21.