AI Executive Office — CXO Assistant Platform (On-Premises)

Architecture Views

30 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

Thirty views of a multi-tenant AI Executive Office that runs entirely inside the customer’s own data centre — no cloud plane, no vendor API call, no weight or prompt that leaves the building: a governed conversational layer over enterprise data, specialist agents, decision intelligence and controlled execution, architected as a decision intelligence platform rather than a chatbot with company data. Every component is open source or an enterprise product licensed to run on the customer’s own iron. The set reads in seven acts: the boundary, the people, the structure, the data, the runtime, the operations and the assurance. Below the index sits the architecture one-pager and the full decision record: thirty-four high-level decisions covering every component and technology on these views, each with the alternatives that lost, what the choice costs, and when you should choose differently.

1 · Context and scope

What sits inside the boundary, who and what touches it, and the shape of the whole platform in one picture.

2 · People and journeys

Who the platform is for, what each of them gets to do, and the three journeys whose worst moments the rest of the set has to answer.
03 The executive office — the people the platform is for CXO 3–8 per tenant Goal — Tell me what changed overnight and settle the decisions that need me, before my first meeting. Core journeys Morning brief to approval daily Ask across every system Weigh an option set Chief of staff 1–3 per CXO Goal — Have the investigation already done, with the numbers reconciled, before my principal asks. Core journeys Delegated investigation Prepare the option set Chase an executed action Business unit head 20–200 per tenant Goal — See the exception in my area before it reaches the CEO, and be the one who fixes it. Core journeys Own an escalation Respond to a variance Assurance and operation Governance officer risk and compliance Goal — Show a regulator exactly how the platform reached a decision that moved money. Core journeys Reconstruct a decision on demand Review the abstention log Implementation lead delivery partner Goal — Stand a new organisation up on the platform in six weeks without writing product code. Core journeys Onboard a tenant 6 weeks Map a source to canon Publish a KPI pack Tenant administrator customer side Goal — Change a threshold or an approval limit myself, on a Tuesday, without raising a ticket. Core journeys Set approval thresholds Grant and revoke access Platform SRE platform run team Goal — Know which tenant is degraded and why, without opening tenant data to find out. Core journeys Triage a degraded tenant Drain a site Machines that act without being asked Detection sweep hourly + on event Goal — Find the exception worth an executive's attention, and nothing else. Core journeys Score signals against baselines Raise a situation Ingestion pipelines per source contract Goal — Land every source inside its freshness contract, or say loudly that I did not. Core journeys Land and reconcile a source Publish a staleness flag Outcome monitor per executed action Goal — Close the loop: prove the intervention worked, or reopen the decision. Core journeys Track a KPI after an action Reopen a failed intervention Source systems 12 classes Goal — Be read on a schedule I can sustain, and written to only through my own front door. Core journeys Serve a governed read Accept an approved write Actors and Their Core Journeys Person or role Journey / task Security / platform External / third party v 1.0 · owner Data & AI Global Practice · date 2026-09 Actors and Their Core Journeys Who the platform is for, in their own words, and the named things each of them gets to do with it. HTML page SVG draw.io

3 · Structure

The layering rule, the deployable units in one tenant, every interface in and out, and how one codebase serves many organisations.
07 Experience Mattermost app brief · approve Executive web React on Kubernetes Mobile MDM managed Tenant console config, not code Edge and API HAProxy + WAF TLS 1.3 · rate limit Kong Gateway per-tenant quota Experience API one contract per surface Orchestration Executive Orchestrator intent · plan · compose Specialist agents finance · risk · projects Decision engine situation to options Investigation runner Temporal Conversation state Redis · 24 h Grounding Tool plane MCP servers · typed Retrieval service hybrid + trimming Metric service the only KPI path Prediction service forecast · anomaly Guardrail service ground · abstain AI platform vLLM serving reasoning · routing · embed LangGraph + Langfuse agents · evals · tracing OpenSearch vector + BM25 Llama Guard + NeMo shields · groundedness Prompt registry versioned · pinned Data and analytics MinIO object store medallion Trino + Iceberg canonical model Semantic model Cube measures Entity graph Neo4j MLflow + KServe registered endpoints Kafka Streams signal streams Integration Airflow + dbt batch + CDC Kafka events in Action queue Temporal sagas Camunda 8 execution plane Connector identities one per system Cross-cutting Keycloak token exchange · JIT Vault + HSM per-tenant key OpenMetadata catalogue · lineage Prometheus + Grafana traces · cost Falco + OpenSearch posture · SIEM OPA Gatekeeper admission guardrails one call per turn MCP tool call Cube only filtered query outcome callback Layered Architecture — What Depends on What Application we own Interface / broker Data store Security / platform Queue / topic synchronous event / async Dependencies point down. The one upward arrow is the execution plane reporting an outcome, and it is asynchronous by design. v 1.0 · owner Data & AI Global Practice · date 2026-09 Layered Architecture What depends on what, and the one dependency that is deliberately allowed to point the other way. HTML page SVG draw.io
08 Primary data centre — one tenant runtime, in country Edge HAProxy + WAF active / active Kong Gateway internal only Mattermost bot webhook service Application plane — Kubernetes namespace per tenant Experience API FastAPI · 3–30 pods Orchestrator host LangGraph MCP tool servers one per domain Decision service option sets Investigation runner Temporal worker Config service tenant packs Governance API audit export AI plane — GPU node pool, no egress vLLM inference Llama 3.3 70B Agent runtime LangGraph + MCP OpenSearch 3 data nodes Guardrail service Llama Guard 3 KServe endpoints forecast · anomaly Data plane Trino + Spark MinIO + Iceberg Semantic model Cube · row policies Neo4j causal cluster PostgreSQL decision store Redis state · semantic cache Evidence store MinIO · object lock Integration and execution Airflow pipelines Kafka ingest Temporal action sagas Camunda 8 execution plane Edge collector DMZ reach Platform services Keycloak token exchange Vault + HSM tenant key OpenMetadata lineage Prometheus Grafana · Loki · Tempo OpenSearch SIEM SIEM ERP and finance Procurement Projects Document estate Mail and calendar mTLS in-cluster mTLS approved write CDC pull as the caller Container Architecture — Deployable Units in One Tenant Interface / broker Application we own Data store Security / platform Queue / topic External / third party synchronous batch Pooled-tier tenants share the AI and data planes with a per-tenant index, schema and row policy. The siloed tier deploys this whole picture on its own cluster — view 10. v 1.0 · owner Data & AI Global Practice · date 2026-09 Container Architecture The deployable units inside one tenant, the technology behind each, and which of them holds a credential. HTML page SVG draw.io

4 · Data

Which store owns what, what can be rebuilt and what cannot, the entities every agent shares, and how documents become citable evidence.

5 · Runtime

What actually happens on a question, who computes what, how an exception is found before anyone asks, and how an approval becomes a verified outcome.

6 · Operations

Where it runs, how a change reaches an executive, how anyone knows the answers are still good, and what is allowed to degrade.
21 Primary data centre — in country Application tier — spread over 3 fault domains HAProxy cluster VRRP · active pair Kong Gateway 3 nodes Kubernetes 3–30 pods Temporal 3 nodes · 3 domains AI tier vLLM on GPU nodes 8 × H100 · MIG OpenSearch 3 data · 3 shards KServe endpoints 2 nodes minimum Data tier PostgreSQL Patroni · 3 nodes Trino + Spark MinIO erasure coded Neo4j 3 nodes Redis cluster 3 nodes Evidence store object lock Hub network Egress firewall egress control Internal DNS split horizon Bastion no public admin MPLS link to the enterprise Secondary site — only where the tenant's tier permits it PostgreSQL standby RPO 5 min Evidence replica second site Redeploy from Git RTO 4 h · cold Sovereign tenants: no second site RTO is a restore, in country Enterprise network MPLS link Mail and calendar on-site servers GPU supply is the binding constraint size and order before contract private log shipping Deployment Topology — Sites and Failure Domains Interface / broker Application we own Data store Security / platform Risk / gap External / third party synchronous event / async Availability is bought inside one site, across racks and power feeds. A second site is a tier option, because for a sovereign tenant a second building may be a contract question rather than a resilience feature. v 1.0 · owner Data & AI Global Practice · date 2026-09 Deployment Topology Where it runs, what the failure domains are, and why a second data centre is a tier option rather than a default. HTML page SVG draw.io

7 · Assurance

Where the trust boundaries are, how the caller's authority reaches the data, what never leaves the country, and how a decision is reconstructed months later.
26 Internet — untrusted Executive device MDM compliant Attacker credential · injection Chat and mail Mattermost · SMTP Perimeter — public ingress ends here HAProxy + WAF OWASP CRS · limits Perimeter firewall Keycloak MFA · device policy Application — cluster-internal, no route from outside Kong internal no public route Experience API workload identity Orchestrator no data credential Tool plane MCP · the only door AI and data — mesh-internal, keys in the HSM vLLM serving no egress OpenSearch ACL fields Decision store encrypted · Vault key Trino + MinIO catalogue RBAC Vault + HSM per-tenant key Execution and egress — the only outbound path Execution plane write scopes only Egress firewall FQDN allow-list Enterprise systems over the MPLS link Management — separate identities, no standing access JIT admin access approval + time box Bastion OpenSearch SIEM SIEM · UEBA Falco + Trivy runtime · posture TLS 1.3 blocked at the edge mTLS mesh typed call only caller identity allow-listed FQDN private Security Zones — Where an Attacker Arrives, and What Stops Them Person or role Risk / gap External / third party Interface / broker Security / platform Application we own Data store synchronous failure / alternate The orchestrator holds no credential for any store. Compromising it yields the ability to ask questions as the caller, and nothing more. v 1.0 · owner Data & AI Global Practice · date 2026-09 Security Zones Where an attacker arrives, what stops them, and what they would actually get if they took the orchestrator. HTML page SVG draw.io
28 In country — the enterprise data centre, customer-held keys Data at rest Object store encrypted · Vault key Decision store encrypted · Vault key Search index encrypted · Vault key Evidence immutable + CMK Logs in-site cluster Processing Model inference on-site GPUs Embedding on-site ML training on-site compute Application compute on-site Key and identity control Vault + HSM customer holds the key Break-glass approval vendor access logged JIT access no standing admin Allowed out — explicitly, and only these Vendor support bundles no customer content Threat intelligence signatures in Mirrored package feeds build time only Never leaves the boundary Prompts and completions Retrieved passages Business data and KPIs Decision and audit records Control framework mapping evidence per control GPU supply is the open risk size and order before signing revoke = unreadable declared Sovereignty — What Stays In Country, and What Is Allowed Out Data store Application we own Security / platform Interface / broker Risk / gap External / third party synchronous batch Residency is a physical property here — the estate has no cloud plane at all. Admission control refuses any workload without a residency label, and the firewall denies every egress that is not on the list. v 1.0 · owner Data & AI Global Practice · date 2026-09 Sovereignty and Residency What stays in country, the short list of what is allowed out, and the constraint that could invalidate the design. HTML page SVG draw.io

Architecture One-Pager

The whole argument on one page: the problem, the shape, the decisions that carry it, the numbers, and what is deliberately not being built.

⬇ Download this one-pager as a Word documentEverything in this section, plus the technology table, the risk register and the index of decisions — in the format a reviewer can mark up.

A governed conversational layer over enterprise data, specialist AI agents, decision intelligence and controlled execution — delivered as one product that runs entirely inside the customer's own data centre, on open source, and that many organisations can buy.

An executive's view of their organisation is assembled by people. A question such as "which suppliers are putting projects that are already financially at risk into further trouble?" crosses an ERP, a procurement system and a portfolio tool, and today it is answered by a chief of staff spending two days reconciling three exports. The information exists; the joins, the trust and the audit trail do not. The failure is not a missing dashboard — it is that no single component owns the relationship between a supplier, a contract, a project and a budget line, and nothing records why a decision was taken once it is.

The platform puts a single conversational front door over that estate and makes four things structural. Numbers come from one governed semantic layer and never from a language model. Every retrieval carries the caller's own permissions, so the AI can never see more than the person asking. Every claim in an answer is bound to an evidence identifier that is snapshotted, not re-queried. And the system may propose an action but never perform one — a separate execution plane, holding the only write credentials in the estate, acts after a human approves, and then watches whether the intervention worked. A fifth property is structural to this variant in particular: the model weights sit on the customer's own accelerators, so no prompt, no retrieved passage and no figure ever crosses the building's perimeter. What the executive experiences is a morning brief and a conversation; what the architecture actually is, is a decision record with a loop closed around it.

What it is, and what it is not

A decision intelligence platforma chatbot with company data attached
A governed read layer plus one narrow write pathan agent with credentials to enterprise systems
One product configured per tenanta bespoke build repeated per customer
A semantic layer that owns every KPIa model that calculates figures from retrieved text
An audit trail that reconstructs what the approver sawa chat log with timestamps
Sovereign because the hardware is in the buildingsovereign by contractual assurance
Open-source components the customer can forkan open-source badge on a hosted dependency

The decisions that are the architecture

01The decision record is the system of record

Situation, evidence, options, approval, action and outcome are one durable entity. The conversation is a rendering of it. Everything about auditability, closed-loop monitoring and reopening a failed intervention follows from this one modelling choice.

ADR-01

02The model never touches data

All access is through a tool plane of first-party Model Context Protocol servers, each call carrying the caller’s exchanged token. There is no text-to-SQL against production, and the orchestrator holds no credential for any store — so compromising it yields the ability to ask questions as the caller, and nothing more.

ADR-02

03Deterministic before generative

KPIs come from measures in a governed semantic layer; forecasts and anomaly scores from models registered in MLflow and served on KServe. The language model classifies, plans and narrates. It never produces a figure. This is both the trust control and, on a fixed GPU budget, the largest capacity control.

ADR-03

04Abstention is a first-class outcome

Every claim is bound to an evidence identifier and mechanically checked for groundedness. Unsupported claims are removed and the gap is stated. A platform that never says "I do not have sufficient evidence" is not more accurate, only less honest.

ADR-04

05Isolation is a purchased tier

Pooled tenants share a cluster with a per-tenant index, row-level security and per-tenant keys. Siloed tenants get their own cluster, their own GPUs and their own database hosts, from the same Terraform. One codebase, two deployment shapes, and a tier change is a migration rather than a fork.

ADR-25

06Execution is a proposal, never a write

Approved actions go to a single execution plane holding a distinct connector identity per target system, with an idempotency key and an authority re-check at execution time. The AI identity holds no write scope anywhere in the estate.

ADR-21

07Sovereignty is physical, then enforced

There is no cloud plane to leak into: every process runs on hardware the customer can point at. Admission control refuses any workload without a residency label, the perimeter firewall denies unlisted egress, and the customer holds the encryption key in its own HSM. Revoking the key makes the data unreadable — which is a guarantee rather than a clause.

ADR-27

Non-functional targets

Targets are stated so they can be tested and argued with. Where a number is an assumption rather than a measurement it says so, because a target invented to fill a table is worse than an admitted gap.

QualityTargetHow it is metView
Simple KPI question p95 under 5 s One small-model classification, one semantic-layer measure, one composition call on a resident model. Semantic cache keyed on question, tenant and permission fingerprint. 15
Cross-system investigation 30–180 s, asynchronous Temporal orchestration with progressive disclosure, resumable across a pod restart. Anything needing more than four tool calls is promoted to this path rather than made to wait. 16
Morning brief Generated 06:30, delivered 07:10 local Hourly deterministic detection sweep; narration only for signals a detector already raised. 18
Platform availability 99.9% monthly Measured on the ability to answer a KPI question, not on resource health. Each tier spread across three racks on independent power and network paths inside one hall. 21
Recovery — decision store RPO 5 min, RTO 4 h PostgreSQL under Patroni with synchronous replicas across racks and continuous WAL archive to the object store. Where the tenant permits no second site, RTO is a same-building restore onto spares, targeted at 8 h and tested quarterly. 11
Approval to executed write Under 30 s Temporal saga with a lease and an idempotency key, Camunda for the human step, and authority re-checked at execution time. 20
Answer groundedness No unsupported claim released Claim-level binding verified against retrieved evidence before release; failures dropped and recorded. Rate monitored on a rolling 24 hours as a release-blocking signal. 29
Data freshness Declared per source, shown per answer Finance CDC 15 min, projects hourly, documents 4 h, HR and CRM nightly. Staleness past contract is surfaced in the answer, never hidden. 12
Tenant scale Hundreds of concurrent users per tenant Sized for 3–8 executives and 20–200 business-unit heads per tenant. The load is briefing-time spiky rather than sustained, which is why resident GPU capacity and an admission queue matter more than replica count. 10
Audit retention 10 years, immutable Object-lock retention on the store with legal hold, per-tenant daily hash chain, and a writer identity that cannot delete. 30

Scope

In scope

  • Conversational front door in the enterprise chat client, web and mobile, with a role-aware executive Today view
  • Executive orchestrator plus ten specialist capability packs sharing one runtime
  • Governed tool plane over ERP, finance, procurement, projects, HR, CRM, supply chain and the document estate
  • Canonical semantic layer, entity graph and enterprise retrieval with per-caller security trimming
  • Decision engine: situation, evidence, root cause, forecast, options, recommendation, confidence
  • Governed execution with human approval, plus closed-loop outcome monitoring
  • Multi-tenant control plane, industry packs, and a tenant configuration studio
  • Sovereignty, audit, lineage and AI governance as product capabilities rather than a later phase

Explicitly out of scope

  • Replacing any system of record. The platform reads them and, in two places, writes to them
  • Replacing the enterprise BI estate. The customer's BI tool reads the same semantic layer rather than a parallel one
  • Autonomous action above a tenant-configured threshold. Low-risk automation is a later, opt-in capability
  • Master data management as a product. Entity resolution is performed for the platform's own model, not offered as an MDM service
  • Model training on tenant data. Grounding is retrieval and tools; no tenant corpus is used to fine-tune a shared model
  • Voice, and email and calendar where the tenant does not permit it — both are per-tenant switches, and the architecture must be correct with them off

The four-week prototype

Not a slice of the platform, and not twenty disconnected features. The prototype proves one thing: that the loop closes. A single executive scenario carried end to end — see, understand, predict, decide, act, monitor — is more convincing to a sponsor than any breadth demonstration, and it is the only way to discover early that entity resolution, not the AI, is the hard part. On-premise adds one constraint to the four weeks: the accelerators have to already be in the building. Where they are not, the prototype runs on a single borrowed GPU host with a smaller model and states that latency figures are provisional.

  1. Three sources: ERP/finance, procurement, projects — chosen because they share the supplier and project keys the flagship question needs
  2. Three capability packs: finance, procurement risk, projects
  3. One orchestrator, one tool plane with roughly eight typed tools
  4. One semantic layer with a dozen governed measures, reconciled against the customer's own board pack
  5. One entity graph joining supplier, contract, purchase order and project
  6. One executive Today view and one conversational surface, in the enterprise chat client
  7. Decision record, approval and one governed write to procurement
  8. Golden-question suite written by the customer, run as the go-live gate
  • "What should I know this morning?" — the detection sweep surfaces a project delay
  • "Why is this happening?" — the investigation crosses ERP, procurement and project data
  • "What happens if we do nothing?" — a registered forecast endpoint, with an interval
  • "What are my options?" — three interventions with cost, risk and delay
  • "Proceed with option 2" — approval, governed write, external reference recorded
  • "Is the situation improving?" — the outcome monitor answers, or reopens the decision

Open risks, carried rather than hidden

RiskIf it landsResponse
GPU supply, and power and cooling in the target hall The central premise — inference on the customer's own hardware — fails if accelerators cannot be obtained in time or the hall cannot power and cool them. Lead times run to months and a dense GPU rack materially changes a facility's power draw Size and order the accelerators, and confirm rack power and cooling with the customer's facilities team, before contract rather than after. Three pre-agreed fallbacks: run a smaller open-weight model on the accelerators actually available and declare the capability gap; place inference on a customer-approved hosted endpoint under a written data-boundary commitment; or ship the data platform first and defer the AI layer. The choice belongs in the contract (ADR-29)
Entity resolution across ERP, procurement and projects The flagship cross-domain question is unanswerable until supplier and project identity is resolved. This is the most commonly underestimated item in the programme Treated as a first-class workstream with a human review queue, not a pipeline step. Reversible merges via retained source keys. Proven in the prototype on three sources before scope grows (ADR-16)
Executive adoption The real incumbent is a chief of staff. If the platform is slower or less trusted than a person, it becomes shelfware regardless of accuracy Design for the delegate as well as the principal. Recommendation acceptance is tracked as a platform metric with an alert, so disengagement is visible as an outage rather than discovered at renewal (view 23)
One number, two definitions A figure that differs from the board pack destroys trust faster than a wrong answer, because it is not obviously wrong A single semantic layer owned by finance, read by both the platform and the customer's BI tool. Reconciliation against the customer's existing reporting is an explicit onboarding gate (ADR-17)
Cross-tenant leakage in the pooled tier A single incident ends the multi-tenant commercial model for public-sector customers Server-side mandatory tenant filter, per-tenant keys, and automated cross-tenant probes on every release. Sovereign tenants are siloed by default (ADR-25)
Integration breadth Fourteen interfaces is a programme risk that dwarfs the AI work; legacy public-sector ERPs often have no change-data capability Cadence is declared per source and a nightly full extract is an accepted pattern. The prototype takes three. What a slow source means for "what changed since yesterday?" is stated in the answer rather than papered over (ADR-20)
Guardrail latency against the 5-second target The first thing questioned when the target is missed will be the controls that protect trust Groundedness runs on the extracted claim set rather than full text, and the latency budget for guardrails is stated up front rather than discovered under pressure (ADR-04)
Cost per interaction A multi-agent platform with unbounded fan-out has unpredictable unit economics, which breaks per-tenant pricing Per-turn budgets for tokens, tool calls and wall-clock; routing that defaults to the cheapest sufficient path; cost per interaction tracked per tenant with a budget alert (ADR-10)
Operating roughly twenty pieces of infrastructure software A managed cloud hides the patching, upgrade and capacity work behind a service. On-premise it is real work with a real headcount, and it is the most commonly underestimated cost of an on-premise decision — more so than the hardware The estate is deliberately narrowed: one runtime, one search engine, one queue, one workflow engine, one relational store, one object store. Everything is deployed by Argo CD from Git, upgraded on a published cadence, and exercised in a quarterly game day. The run team's size is stated in the commercial model rather than discovered in year two (ADR-11)

Architecture Decision Record

Why every component and every technology on these 30 views is what it is, and what each choice costs.

Thirty-four decisions that make up this architecture. Everything else on these thirty views is convention, and convention needs no defending. Each record is written to teach as well as to record: it opens with the forcing question and the context that makes it hard, states the decision so it can be checked, then shows the concrete on-premise mechanism that realises it — which package, configured how, on whose hardware. It weighs the credible alternatives — some rejected, some genuinely right for an organisation with different constraints — and states what the choice buys, what it costs, the conditions that would flip it, and the principle that transfers to systems that are not this one. Decisions are grouped into nine areas; use the filter to read one area at a time.

Status of this document. This is a design, not a post-mortem of a running system. No figure in it is a measured production number: targets are engineering commitments to be tested, and volumes are stated assumptions drawn from the requirement. Every component named here is open source, or an enterprise product licensed to run on hardware the customer owns; nothing in the design depends on a public cloud service, an external API or a vendor-hosted endpoint. That is the point of the exercise, and it moves the risk rather than removing it: instead of a provider's regional service list, the binding constraints become GPU supply, rack power and cooling, and the customer's own capacity to operate roughly twenty pieces of infrastructure software — all three of which are assessed in ADR-29 and ADR-11 rather than assumed away. Open-source version currency is a standing obligation, not a one-off: the same team that owns the platform owns its patching. Where a decision rests on a customer fact not yet confirmed — hall capacity, spares contract, directory topology, an existing lake's table format — the record says so rather than assuming in the platform's favour.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it works on-premiseThe concrete mechanism: which package, configured how, on whose hardware.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
LessonThe principle that transfers beyond this platform.

Decision map

Foundations 4

The four commitments everything else is built on, and the ones a reviewer should attack first.

ADR-01The decision record is the system of record, not the conversation ADR-02The model never touches data; a typed MCP tool plane does ADR-03Deterministic before generative: the model never produces a number ADR-04Every claim is bound to evidence, and abstention is a first-class outcome

Experience and access 3

Where an executive meets the platform, and how a request gets in.

ADR-05The enterprise chat client is the primary surface; the web app is the deep surface ADR-06One front door: an edge proxy for the perimeter, Kong for the contract ADR-07Simple questions are synchronous; investigations are durable orchestrations

Orchestration and agents 5

What plans a turn, what runs it, and where it is hosted.

ADR-08LangGraph for the agent runtime, with the plan owned in our code ADR-09Specialist agents are capability packs in one runtime, not separate applications ADR-10Two model classes, routed — and a budget on every turn ADR-11Kubernetes for the stateless tier, dedicated hosts for the stateful one ADR-33One MCP server per domain, with tools resolved per role at start-up

Grounding and knowledge 3

How documents become citable evidence without leaking.

ADR-12Hybrid retrieval with the security filter applied before scoring ADR-13Permissions and sensitivity are captured at ingestion and carried on the chunk ADR-14The semantic cache key includes the caller's permission fingerprint

Data and semantics 6

The analytical foundation, the shared model, and who is allowed to compute a number.

ADR-15An Iceberg lakehouse on the customer's own object store, queried in place rather than copied ADR-16A canonical model with resolved entities, not query federation ADR-17One measure definition, owned by finance, shared with the BI tool ADR-18A property graph for cross-domain relationships ADR-19Forecasts and anomaly scores come from registered models, not prompts ADR-20Freshness is declared per source and published with every answer

Decision and execution 4

Turning a recommendation into a governed action, and knowing whether it worked.

ADR-21One execution plane holds every write credential in the estate ADR-22Options with cost, risk and delay — not a single recommendation ADR-23The loop closes on the record it opened ADR-24Proactive detection is deterministic and scheduled

Multi-tenancy 2

How one product serves many organisations without becoming many products.

ADR-25Isolation is a purchased tier, not an engineering compromise ADR-26Configuration, not custom code — but the configuration is engineered

Sovereignty and security 4

Residency, key custody, and the authority chain from person to row.

ADR-27Residency is physical first, then enforced by network, admission and key custody ADR-28The caller's authority reaches the row; no read-everything identity exists ADR-29GPU supply and facility capacity are carried as a contracted risk, not an assumption ADR-34First-party MCP servers only, and tool descriptions are untrusted input

Operations and assurance 3

Releasing safely, degrading honestly, and proving it afterwards.

ADR-30Evaluation is a release gate, and a prompt is a release artefact ADR-31Evidence is snapshotted, and the writer cannot delete ADR-32A written degradation contract, and one place the platform fails closed

Technology by capability

What each capability is built from, the credible alternative, and why this one. Every row links to the record that argues it. Origin labels: Open source is a project the customer runs from its own repository or registry mirror; Enterprise on-prem is a commercially licensed product installed on the customer's hardware; This design means the requirement is met by a pattern the team builds rather than a package it installs. Every row runs inside the customer's data centre; there is no hosted control plane anywhere in the table.

Open source Enterprise on-prem This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
Executive surface Self-hosted chat client (Mattermost) + React web app on Kubernetes Open source Web portal only; a hosted collaboration suite Executives approve where they already talk; the web app carries the depth a chat attachment cannot. ADR-05
Perimeter and WAF HAProxy pair with VRRP, Coraza WAF on the OWASP core rule set Open source NGINX + ModSecurity; an enterprise load balancer appliance TLS termination, rule-set WAF and failover from two small components the team can read end to end. ADR-06
API gateway Kong Gateway, cluster-internal only Open source Apache APISIX; ingress plus custom middleware Per-tenant quota, token budgets and one enforced contract per surface, without writing a gateway. ADR-06
Application hosting Kubernetes on bare metal, one cluster per site Open source A supported distribution such as OpenShift; plain hosts under configuration management A tenant becomes a namespace, and scaling down between briefings returns capacity to a fixed pool. ADR-11
GPU platform NVIDIA GPU Operator with MIG partitioning on H100-class nodes Enterprise on-prem Time-sliced whole GPUs; CPU-only inference on a smaller model The router, embedding and reranking models share one card while the reasoning model keeps its own. ADR-10
Service identity and east-west security Istio mesh with SPIFFE workload identity and mutual TLS Open source Cilium mTLS; TLS terminated per service Every hop is authenticated and encrypted without a credential in any configuration file. ADR-27
Long-running investigations Temporal Open source A queue plus a hand-written state machine An investigation is a resumable orchestration with checkpoints, not a long HTTP request. ADR-07
Agent runtime LangGraph on a FastAPI host, checkpointing to PostgreSQL Open source AutoGen; a bare tool-calling loop Durable graph execution and thread state, with the routing policy kept in code the team owns. ADR-08
Language models vLLM serving an open-weight 70B-class reasoning model and a 7B-class router Open source One large model for everything; a hosted API under a data-boundary commitment Most turns need routing, not reasoning, and on fixed accelerators that is the largest capacity lever there is. ADR-10
Enterprise retrieval OpenSearch — BM25 and k-NN vectors fused, filters in query context Open source Elasticsearch; pgvector on PostgreSQL; a vector-only store Executive questions mix concepts with exact identifiers, and the security filter must run before scoring. ADR-12
Embeddings and reranking BGE-M3 and a cross-encoder reranker on Text Embeddings Inference Open source A smaller embedding model; a domain-tuned one Served on the same accelerators as everything else, so there is no separate residency argument to make. ADR-12
Guardrails and grounding Llama Guard 3 for input shields, an NLI entailment verifier for groundedness, Presidio for PII Open source A second large model as judge; a commercial moderation service Versioned separately from the generating model, which is what makes the check independent. ADR-04
Document processing Docling and Apache Tika with an OCR pass Open source A commercial document-understanding appliance Layout and table extraction from PDFs and board papers without a document leaving the building. ADR-13
Object storage MinIO, erasure coded, with object lock Open source Ceph object gateway; an enterprise NAS with S3 emulation S3 semantics on the customer’s own disks, and immutability enforced by the store rather than the application. ADR-15
Analytical foundation Apache Iceberg tables on MinIO, queried by Trino, built by Spark Open source A traditional relational warehouse; a Spark-centric platform One storage layer and one catalogue, and the customer’s existing lake is read in place rather than copied. ADR-15
Transformation and orchestration Apache Airflow with dbt Open source Dagster; scheduled scripts Declared dependencies, per-source watermarks and a lineage event the catalogue can consume. ADR-20
KPI definitions Cube semantic layer over the gold tables, defined as version-controlled files Open source dbt metrics; measures defined in the BI tool alone The finance team owns one definition, and the platform and the BI tool cannot disagree about a number. ADR-17
Business intelligence Apache Superset reading the same semantic layer Open source Metabase; the customer’s existing BI estate The platform does not build a parallel definition of a KPI the customer already owns. ADR-17
Cross-domain relationships Neo4j causal cluster, rebuilt from the gold layer Open source Apache AGE on PostgreSQL; recursive SQL; no graph at all The flagship question is a multi-hop traversal over supplier, contract, purchase order and project. ADR-18
Forecasting and anomaly detection MLflow registry with KServe endpoints on CPU nodes Open source Asking the language model; notebooks on a schedule A forecast needs a registered, versioned, testable artefact with an interval — three things a prompt cannot provide. ADR-19
Decision and audit store PostgreSQL under Patroni, synchronous replicas across racks Open source A document store; a commercial RDBMS the customer already licenses The decision record is relational and transactional, and needs row-level security and point-in-time restore. ADR-01
Evidence custody MinIO object lock in compliance mode, with legal hold Open source Keeping evidence in the database; a WORM tape archive Immutability enforced by the storage layer, so an application compromise cannot rewrite history. ADR-31
Conversation state and semantic cache Redis Open source State in the decision store; no cache Turn state is ephemeral and hot; the cache key includes the caller’s permission fingerprint, which is what makes it safe. ADR-14
Change-data capture Debezium into Kafka Open source Nightly full extracts everywhere; trigger-based capture Low-latency finance data without polling a production ERP, where the source supports a log reader. ADR-20
Event backbone Apache Kafka (Strimzi on Kubernetes) Open source RabbitMQ; NATS High-throughput, replayable, and it lands in the same lake as everything else. ADR-20
Action execution Temporal sagas with a lease and an idempotency key Open source A queue with dead-lettering; a database-backed outbox An approved action must survive a crash exactly once, and be re-authorised on resume. ADR-21
Approval workflow Camunda 8 for approval, escalation, delegation and SLA Enterprise on-prem Approval logic in the decision service; the customer’s existing BPM engine Thresholds, escalation paths and delegation are business rules that change without a release. ADR-21
Execution connectors Apache Camel routes, one connector identity per target system Open source Direct API clients per system; calls from the agent itself Hundreds of enterprise integration patterns, and a blast radius bounded to one system per identity. ADR-21
Identity Keycloak federated to Active Directory, RFC 8693 token exchange Open source A privileged service account with application-side filtering The AI must inherit the caller’s authority; a read-everything identity makes every bug a data breach. ADR-28
Keys and secrets HashiCorp Vault fronting the customer’s HSM over PKCS#11, per-tenant keys Open source Secrets in the cluster’s own store; keys held by the platform team The customer holding the key makes revocation a real control rather than a contractual promise. ADR-27
Catalogue, classification and lineage OpenMetadata with OpenLineage events from the pipelines Open source DataHub; lineage inferred from pipeline metadata Sensitivity labels captured at ingestion are what the retrieval filter and the egress controls both act on. ADR-13
Network segmentation and egress Calico default-deny policy, one egress firewall with an FQDN allow-list Open source Flat network with host firewalls; segmentation at the switch only One controlled egress path is what makes an exfiltration claim defensible — and makes air-gapping a configuration. ADR-27
Admission and policy as code OPA Gatekeeper for admission, OPA for authorisation decisions Open source Kyverno; policy asserted in review rather than enforced Residency and isolation are enforced at the cluster boundary, so a developer cannot deploy around them. ADR-27
Observability Prometheus, Grafana, Loki, Tempo and OpenTelemetry, with Langfuse for turn-level AI traces Open source A commercial APM installed on-premise; logs only A turn must be readable end to end — plan, tool calls, model version, grounding verdict — without leaving the boundary. ADR-32
Runtime security and SIEM Falco and Trivy for posture, an OpenSearch cluster for security analytics Open source Wazuh; Elastic Security; the customer’s existing SIEM with log forwarding Native signal from the AI and data services, and reading a decision record is itself an audited event. ADR-31
Supply chain and registry Harbor registry with Cosign signatures, mirrored package feeds Open source Pulling images and packages from the internet at build time Nothing is fetched from outside the boundary, so the air-gapped deployment is the same deployment. ADR-34
Delivery Self-hosted GitLab pipelines with Argo CD reconciling from Git Open source Jenkins; deployment by hand at a change window A cluster rebuild is a reconciliation rather than a runbook, which is what makes the restore RTO credible. ADR-30
Infrastructure as code Terraform, Helm and Ansible, with the isolation tier as a parameter Open source Manual provisioning; a per-tenant repository fork A tier is a deployment parameter. Onboarding tenant twelve must not cost what tenant two cost. ADR-25
Backup and restore Velero for cluster state, WAL archive for PostgreSQL, MinIO replication for objects Open source Storage-array snapshots alone; an enterprise backup product Three stores with three different recovery shapes; only the operational estate carries a real objective. ADR-32
Tool plane First-party MCP servers behind the gateway, tools resolved per role This design Bespoke REST tool contracts; direct database access from the agent The one door between the model and the data, and therefore the one place authorisation can be audited. ADR-02
Model-to-data contract One MCP server per domain, tool set pinned at start-up This design Run-time tool discovery; one server for everything A capability set known only at run time cannot be reviewed by security or asserted in an audit. ADR-33
Decision engine Situation → evidence → root cause → forecast → options → recommendation, as a service This design Prompting the model to produce a recommendation directly Options with cost, risk and delay are what let an executive decide. A single recommendation asks them to trust instead. ADR-22
Proactive detection Scheduled deterministic detectors over gold-layer measures and the graph This design An agent that continuously monitors the business Repeatable, affordable and explainable. A model that watches is a model you cannot test. ADR-24
Tenant configuration Versioned, schema-checked configuration artefacts in Git, promoted through rings This design Configuration rows in a database edited through an admin UI "Configuration, not custom code" only holds if the configuration is engineered, reviewable and reversible. ADR-26
Evaluation Golden-question suites per tenant scored by Ragas, run as a release gate This design Manual spot checks; published benchmark scores Retrieval and generation are scored separately so a regression has a stage rather than a shrug. ADR-30

The decisions, and the alternatives that lost

FoundationsThe four commitments everything else is built on, and the ones a reviewer should attack first.

ADR-01

The decision record is the system of record, not the conversation

Accepted

When an executive approves an intervention through a conversational interface, what exactly has been created?

Context
The brief asks for auditability, closed-loop monitoring, and the ability to reconstruct why a decision was taken. A chat transcript satisfies none of these: it has no state machine, no outcome, no relationship to the KPI the intervention was meant to move, and no way to be reopened. The requirement also asks for a decision intelligence platform rather than a chatbot, and the difference has to be visible somewhere in the data model or it is only a slogan.
Decision
Situation, evidence, options, approval, action and outcome form one durable, append-only entity with an explicit state machine. The conversation is a projection of that entity, not the other way round. Every surface — chat card, web thread, governance export — renders the same record.
How it works on-premise
PostgreSQL holds the decision, evidence, action and outcome tables with row-level security by tenant. State transitions are appended to a separate history table rather than updated in place, which is what makes the timeline reconstructable four months later. Evidence payloads are snapshotted to the object store under an object-lock retention policy and referenced by identifier. The cluster runs under Patroni with synchronous replicas on separate racks, and continuous WAL archive gives point-in-time restore without a cloud backup service.
Options weighed
  • ChosenDecision record as the durable entity: Makes audit, reopening and outcome tracking properties of the model rather than features bolted on later.
  • RejectedChat thread with structured annotations: Cheaper to build and the default for conversational products, but there is no state machine, so "which decisions are awaiting me" becomes a search rather than a query.
  • RejectedRecords only for actions above a value threshold: The threshold is the argument. A decision not to act is exactly the one an auditor asks about.
  • Right elsewherePush the record into the customer's existing case system: Right where the customer already runs governance on a platform they trust. It costs the closed loop, because the outcome monitor then has no home.
Consequences
What it buys
  • Audit, closed-loop monitoring and reopening all fall out of one model rather than three subsystems
  • The Today view is a query, not an aggregation over chat history
  • A decision survives the conversation, the session and the user's departure
What it costs
  • A relational store with a real backup and retention obligation, which the analytical estate does not need
  • Every surface must write through the decision service rather than talking to the model directly
  • More upfront modelling before anything demonstrable exists
Choose differently when
Choose a lighter model when the assistant genuinely only answers questions and never proposes an action — an internal knowledge assistant with no execution path does not need this and should not pay for it.
LessonIn any assistant that leads to action, the durable artefact is the decision, not the dialogue. Get that wrong and auditability becomes log archaeology.
Shown on views06 13 19 30
ADR-02

The model never touches data; a typed MCP tool plane does

Accepted

How does a language model get access to enterprise data without being given access to enterprise data?

Context
The requirement is explicit that AI must never have unrestricted access and must inherit the user's authorisation context. The tempting shortcut — give the model a database connection and let it write SQL — is fast to demonstrate and impossible to govern: the generated query is unbounded, the identity is the service's rather than the caller's, and nothing about the access is reviewable in advance.
Decision
All data access is through a catalogue of typed tools, exposed to the model as first-party Model Context Protocol servers with declared inputs, outputs and required scopes. The model proposes which tool to call; the tool plane decides whether the caller may, executes the read, and mints an evidence identifier. The orchestrator holds no credential for any store.
How it works on-premise
MCP servers run as their own containers on the Kubernetes cluster, one per domain, reached through Kong with a plugin that requires a token exchanged for the caller. Each server calls the semantic layer, the search cluster or the graph using that token, so row-level security and index filters evaluate against the person rather than the service. The orchestrator binds to the servers a capability pack declares. A mesh-issued workload identity is used only for the platform's own operational stores, never for tenant business data.
Options weighed
  • ChosenTyped tools as first-party MCP servers, caller identity carried through: Authorisation is reviewable in advance and every read is replayable for audit, and the contract is an ecosystem standard rather than a bespoke schema only this platform understands.
  • RejectedBespoke REST tool contracts behind the gateway: Functionally equivalent and entirely defensible; MCP wins only on portability — the same servers can be driven by a different runtime, which matters when the model layer is open-weight and expected to change.
  • RejectedText-to-SQL against a read replica: Demonstrates in a day and fails review in an hour. Unbounded queries, no per-caller authority, and no way to know what a future prompt will produce.
  • Right elsewherePre-computed answer sets only: Right for a fixed executive dashboard with no drill-down, and considerably cheaper. It cannot answer a question nobody anticipated.
Consequences
What it buys
  • Compromising the orchestrator yields the ability to ask questions as the caller, and nothing more
  • Every read is typed, logged and replayable, which is what makes the audit trail meaningful
  • A standard contract: the same server backs the production orchestrator, a developer client and whatever runtime comes next
  • Tool latency and failure rate become measurable per source
What it costs
  • Every new question class may need a new tool, which is slower than a general query interface
  • The tool catalogue becomes a governed artefact with its own change process
  • A dependency on a young protocol whose authorisation and transport details are still settling, which is why identity is enforced at the gateway rather than by the protocol
  • Per-caller queries cache poorly, addressed in ADR-14
Choose differently when
A single-tenant internal assistant over one warehouse can let a read-only role query directly and skip the plane entirely. The cost of this decision is real — a tool catalogue is work, and every new question may need a new tool — and it is only worth paying where the data is genuinely sensitive or the estate genuinely plural.
LessonGive a model a capability, never a credential. Standardising how that capability is described costs nothing and buys portability; it does not, by itself, buy you authorisation.
Shown on views02 15 16 26
ADR-03

Deterministic before generative: the model never produces a number

Accepted

When an executive asks for the cash position, what actually computes the figure?

Context
The requirement warns against asking the LLM to calculate financial KPIs when a deterministic engine can. The deeper problem is asymmetric consequence: a clumsy narrative is embarrassing and recoverable, whereas a wrong figure presented to a board is not. Executives also calibrate on the answers that turn out wrong, so a single fabricated number costs more trust than fifty good answers earn.
Decision
KPI values come from measures in the governed semantic layer. Forecasts, anomaly scores and risk scores come from registered models served behind endpoints, with intervals. The language model classifies the question, plans the retrieval, frames the scenario and narrates the result. It never produces a figure, a forecast or a score.
How it works on-premise
The metric service is the only component permitted to query the semantic layer, and it does so with the caller's identity so row policies apply. KServe endpoints serve forecasting and anomaly models registered in MLflow, as versioned artefacts with intervals. The routing table in view 17 is implemented as a classifier plus a dispatch map, and a test asserts that no other code path returns a numeric measure.
Options weighed
  • ChosenDeterministic computation, generative narration: Trust control and cost control in one decision: most questions route to a cheap path with a single small-model call.
  • RejectedLet the model compute from retrieved tables: Works impressively in a demo with small tables and fails silently on aggregation, currency, period boundaries and null handling.
  • RejectedModel computes, then a checker verifies: Two model calls to reach a number that a measure already defines, with a checker that shares the first model's blind spots.
  • DeferredCode interpreter in a sandbox: Genuinely useful for ad-hoc what-if analysis on data already retrieved, and a candidate for a later scenario-modelling capability under the same evidence rules.
Consequences
What it buys
  • Numbers are reproducible, testable against expected values and explainable without reference to a model version
  • The dominant cost path is avoided for most questions
  • A model upgrade cannot change a reported figure
What it costs
  • A question with no corresponding measure cannot be answered until the measure is built
  • Measure coverage becomes an onboarding workstream with its own backlog
  • The platform will sometimes abstain where a competitor confidently guesses
Choose differently when
Relax this for exploratory analysis clearly labelled as such, on data the user already has, where no figure leaves the session. Never relax it for anything that reaches a decision record.
LessonDecide which of your outputs are opinions and which are facts, and route them to different machinery. A system that blurs the two will eventually be wrong in the expensive direction.
Shown on views02 17 19
ADR-04

Every claim is bound to evidence, and abstention is a first-class outcome

Accepted

What does the platform do when it cannot support part of an answer?

Context
The brief asks the system to say "I don't have sufficient evidence to answer this reliably" rather than fabricate. That is easy to write and hard to build, because it requires knowing which parts of a generated answer are supported — which means the binding between claim and evidence must be mechanical rather than requested politely in a prompt.
Decision
The composition step returns a claim-to-evidence map alongside the prose. Each claim is verified against the retrieved evidence by an independent check. Unsupported claims are removed and the gap is stated. Where the whole answer fails, the platform abstains in those words. Abstention rate is a monitored metric and is deliberately not driven to zero.
How it works on-premise
The check is a natural-language-inference verifier — a small entailment model served on the inference host, separate from the generating model and versioned independently, so the checker is not marking its own work. Llama Guard 3 handles the input-side shields and Presidio the PII pass. Evidence identifiers are minted by the tool plane at read time and carried through composition. The verdict, the dropped claims and the abstention are all written to the decision record and surface in view 30's evidence pack.
Options weighed
  • ChosenMechanical claim binding with an independent groundedness check: The only version of this that survives contact with a regulator, because the check does not depend on the model's own confidence.
  • RejectedPrompt the model to cite and to refuse when unsure: Improves the average case and fails exactly when it matters, since a confidently wrong model is also confidently cited.
  • RejectedShow the sources and let the reader judge: Transfers the verification burden to the person least able to carry it at 07:10.
  • Right elsewhereHuman review before every executive answer: Right in a small, very-high-stakes setting such as regulatory filings. It does not scale to a conversational front door and removes the reason to have one.
Consequences
What it buys
  • Trust is earned by the answers that are declined as much as by the ones given
  • The evidence pack in view 30 becomes producible, because binding already exists
  • Grounding rate is a release-blocking signal rather than a sentiment
What it costs
  • Added latency on every turn, which is the first thing challenged when the 5-second target is missed
  • Answers are shorter and occasionally less satisfying than an ungrounded competitor's
  • An additional service in the critical path and in the cost model
Choose differently when
A lower bar is reasonable for an internal brainstorming assistant where the reader is expected to verify. It is not reasonable anywhere an answer can become an approval.
Lesson"Do not hallucinate" is not a control. A control is a component that can fail the output, run by something other than the thing that produced it.
Shown on views04 15 29

Experience and accessWhere an executive meets the platform, and how a request gets in.

ADR-05

The enterprise chat client is the primary surface; the web app is the deep surface

Accepted

Where does an executive actually meet this platform?

Context
Executive adoption is the largest product risk, and the real incumbent is a chief of staff rather than a competing tool. Anything requiring an executive to open a new application, remember a URL and learn a navigation model starts at a disadvantage that accuracy will not recover. Approvals in most organisations already happen in whatever chat client the enterprise has standardised on — which, in a fully on-premise estate, is a self-hosted one.
Decision
The morning brief, notifications and approvals are delivered as interactive messages in the enterprise chat client, self-hosted alongside everything else. Investigation, option comparison, drill-down and the governance centre live in a web application that a message can deep-link into. Mobile shares the web contract rather than having its own.
How it works on-premise
A bot service subscribes to the chat server's outgoing webhooks and posts message attachments carrying the brief and the approval buttons; the button posts to the same Experience API the web app uses. The web app runs on the cluster behind the edge proxy. Single sign-on flows from the same Keycloak session the chat client uses, so an approval requires no second authentication.
Options weighed
  • ChosenChat for brief and approval, web for depth: Meets the executive where approvals already happen without forcing a rich comparison view into a message attachment.
  • RejectedStandalone portal only: Cleanest to build and the surface most likely to go unopened after week three.
  • RejectedEverything in chat: An option set with cost, risk and delay, and a drill-down into evidence, do not fit a chat attachment without becoming unreadable.
  • Right elsewhereA hosted collaboration suite's own assistant framework: A strong choice for an organisation standardising on a vendor suite with lighter governance needs. It is not available here — the estate is on-premise by requirement — and it would cede the orchestration and evidence contract, which is the product.
Consequences
What it buys
  • Approvals happen where the executive already is, which is the difference between daily use and abandonment
  • One identity, one session, no second sign-in for an approval
  • Push and presence come free rather than being built
What it costs
  • Two front-end surfaces to maintain and to keep consistent
  • A dependency on the customer's self-hosted chat deployment, its upgrade cadence and its own availability
  • Self-hosted chat attachments are less capable than a hosted suite's card framework, so the brief is plainer than its cloud equivalent
Choose differently when
Lead with the web app where the organisation has no chat client it trusts with decision content, or where a regulator objects to that content transiting a collaboration platform — the brief then becomes an email deep-link.
LessonFor executive software, distribution beats features. The surface that is already open wins arguments that the better interface loses.
Shown on views01 02 04
ADR-06

One front door: an edge proxy for the perimeter, Kong for the contract

Accepted

What enforces tenant quota, contract and public exposure, and where?

Context
A multi-tenant platform needs per-tenant rate limits and token budgets, one enforced API contract per surface, and no public endpoint on anything behind the perimeter. Building that into the application means every service re-implements it and the sovereign tenants get a different answer from the pooled ones.
Decision
A highly available HAProxy pair terminates TLS and runs the WAF, and is the only thing in the estate with a route from the user network. Kong, reachable only inside the cluster, is the single ingress to the application and is where tenant quota, token budgets and contract validation are enforced.
How it works on-premise
Two HAProxy nodes share a virtual address by VRRP, terminate TLS 1.3, and run Coraza with the OWASP core rule set in front of every request. Behind them Kong validates the Keycloak token, extracts the tenant claim, applies per-tenant rate and token quotas, and forwards with claims to the application services. Nothing else is published: every other service is cluster-internal, reached over mutual TLS in the mesh, and the perimeter firewall has no inbound rule for any of them. Rate limiting replaces a cloud scrubbing service, and volumetric protection is the enterprise's existing perimeter appliance — which is an honest limitation, not an equivalent.
Options weighed
  • ChosenHAProxy at the edge + internal Kong: Edge protection and contract enforcement as two small, well-understood open-source components, with quota policy in one place for every tenant.
  • RejectedNGINX Ingress + custom middleware: Fewer moving parts and re-implements quota, versioning and developer contract in application code.
  • RejectedCluster ingress directly: Adequate for a prototype and leaves per-tenant throttling — the control that protects one tenant from another — in the application.
  • Right elsewhereThird-party API gateway: Right where the customer already standardises on one and wants a single policy surface across their estate.
Consequences
What it buys
  • A noisy tenant is throttled before it reaches shared compute or shared model quota
  • One contract per surface, versioned, with the policy visible to security review
  • No public endpoint anywhere behind the perimeter
What it costs
  • Two more components the platform team patches and upgrades itself, with no vendor carrying the on-call
  • Another hop in the latency budget
  • Two products to configure before the first request is served
  • No anycast scrubbing tier: volumetric denial of service is the enterprise perimeter's problem, and the platform can only state that honestly
Choose differently when
Drop Kong for a single-tenant deployment with no quota requirement and one client; the policy layer is then genuinely unearned cost.
LessonPut multi-tenant fairness controls in infrastructure, not in application code. The tenant you need to throttle is the one whose code path you are least keen to touch.
Shown on views02 08 21 26
ADR-07

Simple questions are synchronous; investigations are durable orchestrations

Accepted

What happens when a question genuinely takes ninety seconds to answer?

Context
The requirement distinguishes a simple KPI question, targeted under five seconds, from a complex investigation where an asynchronous experience is acceptable. Holding an HTTP request open for ninety seconds fails at every proxy, loses all work on a restart, and gives the user a spinner instead of progress.
Decision
Turns needing at most four tool calls run synchronously. Anything larger is promoted to a durable orchestration with checkpoints, progressive disclosure of intermediate findings, and a resumable identifier the executive can leave and return to.
How it works on-premise
Temporal runs the investigation as a workflow with activity checkpoints in its own PostgreSQL store, so a pod restart resumes rather than restarts. Findings stream to the client over Server-Sent Events as each stage completes, and the same workflow identifier is addressable later from the chat message or the web thread.
Options weighed
  • ChosenThreshold-based promotion to a durable orchestration: The user waits only when waiting is short, and long work is crash-safe and resumable.
  • RejectedEverything synchronous: Simple and correct until the first cross-system question, which is the question the product exists for.
  • RejectedEverything asynchronous: Uniform and honest, but a five-second question behind a job-status poll feels broken.
  • RejectedA queue plus a hand-written state machine: Entirely buildable, and it is the checkpointing, retry and replay semantics — not the queue — that take a year to get right.
Consequences
What it buys
  • Long investigations survive a deployment or a host failure
  • Progressive disclosure makes ninety seconds feel like work rather than a hang
  • The promotion threshold is a tuning knob rather than a rewrite
What it costs
  • Two execution paths to build, test and observe
  • The client must handle streaming and resumption
  • Orchestration state is another store to secure and to reason about for residency
Choose differently when
Stay synchronous throughout if every question in your domain is bounded by a single system call. Go fully asynchronous if none of them are.
LessonChoose the interaction model from the work's distribution, not its average. The tail is what the user remembers.
Shown on views02 15 16

Orchestration and agentsWhat plans a turn, what runs it, and where it is hosted.

ADR-08

LangGraph for the agent runtime, with the plan owned in our code

Accepted

Build the agent runtime, or adopt one?

Context
Agent frameworks are abundant and change quickly. The parts that matter here — tool calling with the caller's identity, durable threads, tracing and evaluation — are infrastructure. The part that is genuinely ours is the planning policy: which capability answers which question, under what budget, with what fallback. In a fully on-premise estate there is no managed agent service to adopt, so the choice is not build-versus-buy but which self-hosted framework to stand on — and how much of the product's judgement to let it own.
Decision
Adopt LangGraph for graph-structured agent execution, thread state and checkpointing. Keep the planning policy, the budget enforcement and the tool authorisation in our own orchestrator service, so the platform's distinctive behaviour is not a framework's default. Tracing and evaluation are separate, swappable components rather than the framework's own, so the framework can be replaced without losing the governance surface.
How it works on-premise
The orchestrator runs on the cluster as a FastAPI service using LangGraph for plan composition and execution, with the checkpointer writing thread state to PostgreSQL so a restart resumes mid-investigation. Tracing follows OpenTelemetry GenAI conventions into Langfuse for prompt-level detail and Tempo for the distributed trace; Ragas provides the offline scoring used by the release gate in ADR-30. Every part of this is a container from the platform's own registry mirror, and none of it calls out of the building.
Options weighed
  • ChosenLangGraph + our planner: Buys durable graph execution and checkpointing; keeps the routing policy, which is where the product's judgement lives.
  • RejectedFully hand-rolled orchestration loop: Maximum control, and a year of building threads, tracing and evaluation that a managed service already provides in-boundary.
  • Right elsewhereAutoGen, or a bare tool-calling loop: Both are reasonable and the right answer for a team already fluent in them. AutoGen's conversational multi-agent model is a poorer fit for a plan that must be budgeted and audited step by step.
  • RejectedA vendor's low-code agent builder: Fast to a demonstration, and unavailable in an air-gapped estate in any case; it would cede the tool authorisation and evidence contract that this platform is fundamentally about.
Consequences
What it buys
  • Tracing and evaluation exist from day one rather than being retrofitted
  • Everything runs inside the building, so there is no residency argument to make about the agent runtime at all
  • The routing policy is testable code we own
What it costs
  • A dependency on a fast-moving open-source project, whose breaking changes land on the platform team rather than on a vendor's compatibility promise
  • Framework upgrades become a release concern
  • Some capability duplication between the framework and our planner
Choose differently when
Build your own when you need portability across clouds, or when your orchestration is genuinely unusual. Adopt more of the platform when your team is small and your governance needs are ordinary.
LessonBuy the plumbing, build the judgement. The framework is not your product; the routing policy might be.
Shown on views02 07 16
ADR-09

Specialist agents are capability packs in one runtime, not separate applications

Accepted

Is a finance agent a deployment, or a configuration?

Context
The requirement names ten specialist domains and describes them as modular capabilities rather than independent LLM applications. It also asks for an agent marketplace as a product module. Those two only reconcile if adding a capability is cheap, and it is only cheap if it is not a deployment.
Decision
A specialist agent is a prompt pack plus a tool catalogue plus an evaluation set, versioned with the tenant configuration and loaded into one shared runtime. Agents return structured results to the orchestrator and never call each other.
How it works on-premise
Packs are configuration artefacts in the tenant configuration repository, schema-checked in the pipeline and loaded by the orchestrator at start-up and on change. The tool catalogue per pack resolves to the MCP servers and tools the caller's role permits, and Kong enforces the same allow-list independently. Thread state keeps per-agent context without a per-agent deployment.
Options weighed
  • ChosenCapability packs in a shared runtime: Adding a domain is configuration, which is what makes a marketplace a feature rather than a programme.
  • RejectedOne microservice per agent: Attractive for team autonomy at large scale; here it multiplies deployments, network policy and cost by ten for no isolation benefit.
  • RejectedOne monolithic prompt covering every domain: Cheapest, and it degrades as domains are added because every question pays for every domain's instructions.
  • RejectedFree agent-to-agent negotiation: Where latency and cost hide. Structured returns to a single orchestrator keep a turn bounded and traceable.
Consequences
What it buys
  • A new domain is a configuration change with an evaluation set, not a release
  • A tenant can be behind on a pack without being behind on the platform
  • One runtime to secure, observe and scale
What it costs
  • Packs share a blast radius: a bad shared runtime release affects every domain
  • Per-domain resource limits must be enforced in the budget rather than by isolation
  • Pack versioning becomes its own governance surface
Choose differently when
Split into services when a domain needs genuinely different scaling, a different compliance boundary, or an independent team release cadence.
LessonModularity is a property of the configuration model, not of the deployment topology. Ten deployments are not ten modules; they are ten things to patch.
Shown on views02 07 16
ADR-10

Two model classes, routed — and a budget on every turn

Accepted

Which model answers, and what stops a single question costing a fortune?

Context
Most executive turns are routing and narration over a computed figure. A minority are genuine multi-step investigations. Sending everything to a reasoning model makes the common case slow and expensive, and an agentic system with unbounded fan-out has unpredictable unit economics — which breaks per-tenant pricing before it breaks the latency target.
Decision
A small, fast model performs intent classification, entity resolution and routing. A reasoning model is used only for planning an investigation and for final composition. Every turn carries an explicit budget for tokens, tool calls and wall-clock, set by the orchestrator before any agent runs.
How it works on-premise
Both models are served by vLLM on the tenant's GPU nodes: a 70B-class open-weight reasoning model held resident with paged attention and continuous batching, and a 7B-class router on a MIG slice of the same hardware. Kong enforces the per-tenant token quota and an admission queue in front of the GPU pool; the orchestrator enforces the per-turn budget and degrades to a narrower answer rather than exceeding it. GPU-seconds and tokens per interaction are exported to the lake and alerted on per tenant — the scarce resource here is the accelerator, which cannot be bought in an afternoon, so the budget is a capacity control before it is a cost control.
Options weighed
  • ChosenRouted two-class models with per-turn budgets: The largest single lever on both latency and cost, and it makes unit economics predictable enough to price.
  • RejectedOne capable model for everything: Simplest to operate and reason about, and it pays reasoning-model prices for "what is our cash position?".
  • DeferredOpen-weight models on managed compute: Worth revisiting for the classification path where volume is high and the task is narrow; it adds an operational burden that is hard to justify at prototype scale.
  • RejectedNo budget, monitor and react: The first runaway investigation is discovered on the invoice, and by then a tenant has been affected.
Consequences
What it buys
  • Predictable cost per interaction, which a commercial model can be built on
  • The common path is fast because it never touches the expensive model
  • A runaway turn degrades gracefully instead of consuming the tenant's quota
What it costs
  • Two deployments to manage, evaluate and keep pinned
  • The router itself can be wrong, sending a hard question down a cheap path
  • Budget exhaustion is a user-visible state that must be explained well
Choose differently when
Use a single model when volume is low enough that operational simplicity beats unit cost, or when routing errors are more damaging than latency.
LessonIn an agentic system, cost is an architectural property, not an operational surprise. Decide the budget before the first token, or discover it on the invoice.
Shown on views16 17 23
ADR-11

Kubernetes for the stateless tier, dedicated hosts for the stateful one

Accepted

What runs the application, given a mid-size tenant and a small platform team?

Context
The workload is a handful of stateless HTTP services with a pronounced daily spike around the morning brief and long quiet periods, plus a set of stateful systems — PostgreSQL, the object store, the graph, the search cluster — that behave badly under a scheduler. The platform team runs many tenants. On-premise there is no managed container service to hide cluster operations behind, so the cluster is a thing the team genuinely owns, and the honest question is how much to put on it.
Decision
Kubernetes runs every stateless service: the experience API, the orchestrator, the MCP servers, the decision and governance services, and the GPU inference pool. The stateful systems run outside it on dedicated hosts, managed by their own operators or by Ansible. One cluster per site, a namespace per tenant in the pooled tier, a dedicated cluster in the siloed tier.
How it works on-premise
A three-master cluster per site on bare metal, with Istio for mesh identity and mTLS, Calico network policy, and the NVIDIA GPU Operator managing the accelerator nodes and their MIG partitions. Application services scale from three to thirty pods on concurrency. PostgreSQL under Patroni, the object store, the graph and the search cluster sit on dedicated hardware, because a stateful rebuild during an incident is not the moment to discover a storage driver's limits. Everything is deployed by Argo CD from Git, so a cluster rebuild is a reconciliation rather than a runbook.
Options weighed
  • ChosenKubernetes for stateless, dedicated hosts for stateful: Right-sized for the workload, and it keeps the databases away from a scheduler that will move them at the worst possible moment.
  • Right elsewhereEverything on Kubernetes, databases included: Tidy, increasingly viable with mature operators, and it puts the recovery of the one store with a real RPO behind an extra layer the team must also debug under pressure.
  • RejectedVirtual machines and configuration management only: Familiar to most enterprise operations teams and the lowest-risk choice for a small estate; it costs the per-tenant packaging and the scale-down that make many tenants affordable.
  • RejectedA serverless-style function runtime for everything: Self-hosted function runtimes exist and suit the event paths; they are a poor fit for the long-lived streaming connections the investigation experience depends on.
Consequences
What it buys
  • One deployment model for every stateless service, and a tenant is a namespace rather than a fleet of virtual machines
  • Scaling down between briefings returns accelerators and cores to the pool, which is what makes the pooled tier affordable on fixed hardware
  • Deployment is an image reference, which keeps the release pipeline simple
What it costs
  • The platform team owns cluster upgrades, certificate rotation and CNI behaviour — real work that a managed service absorbs, and the single largest operational cost of going on-premise
  • A cluster is a large thing to operate for a modest service count, and the exit — back to plain hosts — is recorded rather than assumed away
  • Two operating models — the cluster and the dedicated database hosts — to run and to staff
Choose differently when
Use a commercially supported Kubernetes distribution instead of vanilla where the customer's operations team wants a vendor to call at 3 a.m., and accept the licence cost as buying an on-call rota. Skip Kubernetes entirely, and run containers under systemd on a dozen hosts, if the estate is genuinely one tenant and one site — a cluster is a large thing to operate for six services.
LessonChoose the hosting platform for the team that will operate it at 3 a.m., not for the architecture diagram. Record the exit so the choice stays reversible.
Shown on views08 21
ADR-33

One MCP server per domain, with tools resolved per role at start-up

Accepted

How many MCP servers, and how does an agent learn which tools it may call?

Context
MCP makes it cheap to expose tools, which is the danger. One server holding every tool means a finance change redeploys procurement and every agent sees every tool. One server per tool means dozens of containers for a handful of operations. Separately, a protocol that supports run-time tool discovery invites a design where the set of things a model can do is not knowable until the model asks — which cannot be reviewed, tested or audited in advance.
Decision
One MCP server per business domain — finance, procurement, projects, risk and so on — matching the capability packs. A pack declares the servers it binds to and the subset of tools it may use; that binding is resolved per role at start-up and pinned for the turn. Nothing is discovered at run time.
How it works on-premise
Each server is a container in the cluster with its own workload identity and scale rule. The orchestrator loads the pack's server list and role-filtered tool set at start-up and on configuration change, and Kong enforces the same allow-list independently, so a bug in the orchestrator cannot widen the set. Server version and tool-set hash are written into every decision record.
Options weighed
  • ChosenOne server per domain, tools resolved per role at start-up: Matches the capability packs, so a domain is one server, one prompt pack and one evaluation set. The callable set is knowable before a turn begins.
  • RejectedA single monolithic MCP server: Fewer containers and one deployment, and every domain change becomes a platform release while every agent sees every tool.
  • RejectedDynamic run-time tool discovery: The protocol supports it and it is the wrong default here: a capability set known only at run time cannot be reviewed by security or asserted in an audit.
  • Right elsewhereOne server per tool: Right where tools have genuinely independent scaling or compliance boundaries. Here it multiplies operational surface for no isolation gain.
Consequences
What it buys
  • A domain ships independently: new tool, new server version, no orchestrator release
  • The callable tool set is a reviewable artefact, enforced in two places
  • A misbehaving server is restarted or rolled back alone
  • Per-domain latency, failure rate and cost fall out of the topology
What it costs
  • Ten or so more containers to deploy, observe and patch
  • A cross-domain question fans out across servers, adding a round trip
  • Tool-set versioning is another thing that must be pinned into the record for replay
Choose differently when
A single server is right below roughly a dozen tools with one team owning all of them. Move to run-time discovery only if you build the review and attestation machinery that makes a changing capability set auditable.
LessonLet the protocol’s cheapest feature be the one you refuse. Run-time discovery is convenient for a demo and unreviewable in an enterprise; pin the capability set and you can still answer "what could it have done?" a year later.
Shown on views07 08 16

Grounding and knowledgeHow documents become citable evidence without leaking.

ADR-12

Hybrid retrieval with the security filter applied before scoring

Accepted

How is the document estate searched without returning something the caller may not read?

Context
Executive questions mix concepts with exact identifiers — a contract number, a project code, a supplier name. Vector search alone handles identifiers poorly. Separately, the obvious implementation of permissions — retrieve, then filter the results — leaks nothing but silently returns a worse top-k, because the documents the caller may read were crowded out before filtering and never scored.
Decision
Retrieval is hybrid: vector and keyword, fused, then reranked. The caller's group membership is applied as a filter inside the search query, evaluated before scoring, so the top-k is drawn only from documents the caller may read.
How it works on-premise
OpenSearch with a k-NN vector field from BGE-M3 alongside a BM25 text field, combined by reciprocal rank fusion and reranked by a cross-encoder served on the inference host. Group identifiers captured at ingestion are stored as a keyword field, and every query carries a terms filter built from the caller's group claims, placed in the query's filter context so it is evaluated before scoring. Retrieve fifty, rerank to eight.
Options weighed
  • ChosenHybrid retrieval, filter inside the query: Correct top-k and correct permissions, which are two problems that look like one.
  • RejectedVector-only retrieval: Fine for conceptual questions and weak exactly where executives are precise — contract numbers and project codes.
  • RejectedRetrieve then filter in the application: The failure nobody notices: no leak, and a quietly worse answer that looks fine.
  • RejectedAn index per security group: Correct and unmanageable; group membership changes daily and index count explodes.
Consequences
What it buys
  • Identifier-shaped queries work, which is most of what an executive types
  • Permissions cannot be bypassed by an application bug in the result-handling code
  • Top-k is honest, so citation quality is measurable
What it costs
  • Filter cardinality affects query performance on large indexes
  • The security field must be maintained accurately at ingestion, which ADR-13 addresses
  • Two retrieval modes and a reranker to tune and evaluate
Choose differently when
Vector-only is fine for a corpus with no access differentiation and no identifiers. Post-filtering is acceptable only when every user can read everything, which is rarely true and never true in government.
LessonAuthorisation belongs inside the query, not after it. A filter applied to results protects the data and corrupts the ranking.
Shown on views14 26 27
ADR-13

Permissions and sensitivity are captured at ingestion and carried on the chunk

Accepted

Where does the index learn who may read a document?

Context
Re-deriving permissions at query time means calling the source system for every candidate document, which is slow and creates a dependency on a system that may be down. Deriving them once and forgetting them means the index becomes wrong the moment a permission changes. Sensitivity labels have the same problem and additionally drive the egress controls.
Decision
Source access-control identifiers and sensitivity labels are captured during ingestion and written as fields on every chunk. They are never re-derived at query time. Documents are superseded rather than deleted, so a citation made four months ago still resolves to the version that was cited.
How it works on-premise
The crawl reads permissions from each source's own authorisation model — file-share ACLs resolved against the directory, or the source API's own grants — and writes the resulting group identifiers into a filterable field on each chunk. A classification pass runs at landing and its label is stored alongside, registered in the catalogue. Re-crawl every four hours; high-sensitivity libraries are crawled more often, and content labelled restricted is re-checked against the source at query time, at a stated latency cost.
Options weighed
  • ChosenCapture at ingestion, carry on the chunk: Fast queries and a bounded, stated exposure window rather than an unbounded unstated one.
  • RejectedRe-derive at query time for everything: Always correct, always slow, and it makes retrieval unavailable whenever the source is.
  • RejectedTrust a single-level document classification: Coarse enough to be either over-restrictive or wrong; executive material is precisely where per-document permissions matter.
  • DeferredPush-based permission change notifications: Ideal where a source emits permission-change events, and most on-premise document stores do not, so the design cannot depend on it.
Consequences
What it buys
  • Query latency stays flat regardless of corpus size
  • Retrieval keeps working when the source system does not
  • Sensitivity is available to both the retrieval filter and the egress DLP
What it costs
  • A permission drift window of up to one crawl interval, which must be stated to the tenant rather than glossed
  • Ingestion becomes more complex and more failure-prone
  • Superseding rather than deleting increases storage
Choose differently when
Re-derive at query time when the corpus is small, permissions change constantly, or the exposure window is legally unacceptable — and accept the latency openly.
LessonCopying an authorisation decision creates a staleness window. You cannot remove it; you can only measure it, shorten it for the material that matters, and tell people it exists.
Shown on views14 26 30
ADR-14

The semantic cache key includes the caller's permission fingerprint

Accepted

How does a per-caller authorisation model coexist with a cache?

Context
ADR-27's rule — every read carries the caller's identity — makes results caller-specific, and caller-specific results cache badly. Caching on the question text alone is the obvious optimisation and is a data breach with a hit rate: two executives asking the same question are entitled to different answers.
Decision
Cached entries are keyed on the normalised question, the tenant, the model and prompt version, the data-freshness stamp, and a hash of the caller's effective permission set. Two callers share a cache entry only when they would have been entitled to identical results.
How it works on-premise
Redis holds the entries. The permission fingerprint is a stable hash of the caller's directory group claims relevant to the sources involved, recomputed per session. Entries are invalidated on model or prompt version change and on the freshness stamp advancing, so a cached answer can never be stale in a way the answer does not admit.
Options weighed
  • ChosenPermission-fingerprint cache key: Keeps most of the benefit — executives in the same role do share entries — with no possibility of a cross-authority hit.
  • RejectedCache on question text only: The highest hit rate and an authorisation bypass, which is not a trade-off worth having.
  • RejectedNo cache at all: Safe and needlessly expensive during the morning briefing spike, when many executives ask overlapping questions.
  • DeferredCache only the retrieval, not the answer: A useful middle ground worth measuring once real query distributions exist.
Consequences
What it buys
  • Real hit rates during the briefing spike, where identical questions cluster
  • A permission change naturally invalidates by changing the fingerprint
  • Model and prompt version in the key means an upgrade cannot serve stale reasoning
What it costs
  • Lower hit rate than a naive cache, and highly variable by tenant
  • The fingerprint must be computed correctly or the whole guarantee is void
  • Another store inside the residency boundary
Choose differently when
Cache on content alone only where every user has identical entitlements — a public corpus, or a single-role deployment.
LessonA cache key is an authorisation statement. Anything the answer depended on belongs in the key, and permissions are the thing most often left out.
Shown on views15 27

Data and semanticsThe analytical foundation, the shared model, and who is allowed to compute a number.

ADR-15

An Iceberg lakehouse on the customer's own object store, queried in place rather than copied

Accepted

What holds the analytical data, given that most customers already have a lake?

Context
Every target customer has an existing warehouse, lake or BI estate. Copying it into the platform doubles the storage bill, creates a second version of the truth, and makes the platform responsible for data it did not produce. The requirement asks for integration with existing lakes and BI, not their replacement.
Decision
Apache Iceberg tables on a self-hosted S3-compatible object store are the analytical foundation, with Trino for interactive query, Spark for batch, and one catalogue governing both. The customer's existing lake is registered as an external catalogue and queried in place wherever its table format allows, rather than copied.
How it works on-premise
MinIO provides erasure-coded object storage on the customer's own disks; Iceberg gives table semantics, snapshot isolation and time travel over it; Trino serves interactive and federated query and Spark the heavy batch work, both against one catalogue so there is a single place where a table's permissions are defined. A medallion layout — bronze, silver, gold — sits in the object store. Airflow orchestrates batch and dbt builds the transformations. Where the customer's existing lake is already Iceberg, Delta or Hive, Trino reads it as an external catalogue; where it is a proprietary warehouse, a nightly extract is accepted and the copy is declared.
Options weighed
  • ChosenIceberg on object storage, queried by Trino: One storage layer and one catalogue to size and secure, and the customer's lake stays theirs instead of being copied.
  • RejectedA traditional relational warehouse for everything: Simpler to operate and familiar to the customer's DBAs; it neither holds the document estate nor scales to the event volume without becoming very expensive hardware.
  • Right elsewhereA self-hosted Spark-centric platform: The stronger choice where the workload is engineering-heavy or the team already lives in Spark notebooks. Here the semantic layer and interactive query are the centre of gravity, which is Trino's ground.
  • RejectedRead the source systems directly at query time: No historical baseline, no cross-system joins, and every executive question becomes load on production ERP.
Consequences
What it buys
  • One storage layer, one catalogue, one security model, and no per-query cloud bill to model
  • The customer's existing investment is read rather than duplicated
  • The semantic model sits natively beside the warehouse that feeds it
What it costs
  • Capacity is bought in advance as disks and cores, so sizing errors are corrected by a purchase order rather than a slider
  • Four projects — object store, table format, query engine, orchestrator — to version, patch and keep compatible with each other
  • Shortcut support varies by source storage, so some copies remain unavoidable
Choose differently when
Choose a commercially supported lakehouse distribution where the customer's team wants vendor support on the query engine, or drop Trino entirely and use PostgreSQL alone where the analytical volume is genuinely small — a lakehouse for forty tables is architecture theatre.
LessonPrefer mounting to copying. A second copy of the truth is a second thing to reconcile, and reconciliation is where executive trust is lost.
Shown on views08 11 12
ADR-16

A canonical model with resolved entities, not query federation

Accepted

How does the platform answer a question that spans three systems that share no key?

Context
"Which suppliers are affecting projects that are already financially at risk?" requires supplier identity to be the same thing in the ERP, the procurement system and the portfolio tool. It is not: the same supplier is three records with three keys and often three spellings. Query federation joins tables; it does not resolve identity, and no amount of virtualisation supplies a key that does not exist.
Decision
A canonical model with explicit entity resolution for supplier, project, cost centre and contract. Conformed entities retain their source keys so a merge is reversible, and resolution runs as a pipeline stage with a human review queue for low-confidence matches.
How it works on-premise
Spark jobs in the silver layer perform deterministic matching on identifiers and probabilistic matching on names and addresses, writing a crosswalk table with confidence scores. Matches below threshold enter a review queue surfaced in the tenant console. The crosswalk is the one table that is never cached downstream, because a merge is correct on approval and wrong for every consumer until they agree.
Options weighed
  • ChosenCanonical model with entity resolution: The only approach that makes the flagship cross-domain question answerable at all.
  • RejectedQuery federation over source systems: Avoids copying data and cannot join on a key that does not exist. It also puts executive queries onto production ERP.
  • RejectedAsk the language model to match entities: Plausible-looking matches with no confidence score, no review path and no reversibility — the worst possible property for a supplier merge.
  • Right elsewhereBuy a master data management platform: Correct for an organisation that needs mastered data across many consumers, and a programme of its own — twelve to eighteen months before this platform could ship. Self-hosted MDM products exist and are not cheap.
Consequences
What it buys
  • Cross-domain questions become answerable, which is the product's central claim
  • Reversible merges, because entity resolution is a judgement and judgements get revised
  • Match confidence is visible, so an answer can state that identity is uncertain
What it costs
  • The hardest and most underestimated workstream in the programme, and it is not an AI problem
  • A human review queue means an operational process, not just a pipeline
  • A bad merge propagates into every answer until it is corrected
Choose differently when
Federate when the systems already share a reliable enterprise key. Buy MDM when several programmes need resolution and this one would be building it for everyone.
LessonCross-system intelligence is an identity problem wearing an integration costume. Solve identity first, or the joins will be confidently wrong.
Shown on views11 12 13
ADR-17

One measure definition, owned by finance, shared with the BI tool

Accepted

Who owns the definition of a KPI the platform reports?

Context
The fastest way to destroy executive trust is a figure that differs from the board pack, because it is not obviously wrong — it just makes both numbers suspect. Building a metric layer inside the platform guarantees this happens, since two definitions of gross margin will diverge within a quarter no matter how carefully they start.
Decision
KPI definitions live once, as measures in the governed semantic model, owned by the customer's finance function. The platform reads them; it does not restate them. Where the customer already owns a definition in their BI estate, the platform adopts it rather than writing its own.
How it works on-premise
A Cube semantic layer over the gold Iceberg tables holds the measures and the row-level policies, defined as version-controlled files rather than clicked into a tool. The platform's metric service is the only component that queries it, using the caller's identity. Superset reads the same layer, so a report and an answer cannot disagree. Measure changes go through the same versioned promotion as any other artefact.
Options weighed
  • ChosenOne semantic model, shared with the customer's BI: Makes disagreement structurally impossible rather than a thing to police.
  • RejectedA metric service over SQL, owned by the platform: More portable and it recreates the exact divergence the platform exists to eliminate.
  • Right elsewheredbt metrics, or measures defined in the BI tool alone: dbt metrics are a fine choice and closer to where the transformations already live; they lack a query-time row-policy story, which is the half of this decision that carries the security argument.
  • RejectedLet each agent compute its domain's measures: Ten definitions of headcount within a year, and no way to say which is right.
Consequences
What it buys
  • A platform answer and a board pack cannot disagree
  • Finance owns the definition, which is both correct and politically necessary
  • Row-level security is defined once and inherited by every consumer
What it costs
  • A hard dependency on the semantic model's availability and refresh schedule
  • Measure coverage becomes an onboarding gate, and gaps are visible as abstentions
  • Change control on measures is slower than the platform team would like
Choose differently when
Build your own metric layer when there is no incumbent BI estate to align with, or when the incumbent's definitions are known to be wrong and replacing them is in scope.
LessonNever own a second definition of a number someone else already owns. The reconciliation meeting costs more than the integration.
Shown on views11 12 17
ADR-18

A property graph for cross-domain relationships

Accepted

What answers a question that is three relationship hops deep?

Context
The flagship question traverses supplier to contract to purchase order to project to budget line, then asks which of those projects are already at financial risk. In a star schema that is a chain of joins that must be written in advance; in a graph it is a traversal that can be composed at query time. Detection also needs to propagate a supplier signal outward to the projects it affects, which is the same traversal in reverse.
Decision
A property graph holds the canonical entities and the relationships between them, rebuilt from the gold layer. It holds relationships and identifiers, never measures — figures always come from the semantic model.
How it works on-premise
Neo4j in a causal cluster, rebuilt from the gold layer on each refresh. Vertices carry canonical identifiers and a small set of attributes used for traversal; every numeric answer is fetched from the semantic layer by identifier after the traversal returns. The graph is therefore fully rebuildable and carries no independent recovery objective — which is also why the community edition is defensible here if the licence matters.
Options weighed
  • ChosenProperty graph alongside the warehouse: Multi-hop traversal and signal propagation become natural rather than a generated join chain.
  • RejectedRecursive SQL over the warehouse: Avoids another store and becomes unreadable and slow past two or three hops, which is where the interesting questions start.
  • Right elsewhereRecursive SQL in the warehouse, or a graph extension on PostgreSQL: Apache AGE or recursive CTEs remove a component from the estate and are the right answer at two hops; they become unreadable and slow at the four-hop questions this platform is built to answer.
  • RejectedNo graph; pre-compute the known paths: Workable for a fixed question set and precisely wrong for a platform whose promise is questions nobody anticipated.
Consequences
What it buys
  • Multi-hop questions are composable at query time rather than pre-written
  • Detection can propagate a signal along real relationships
  • Fully rebuildable, so it needs no backup story of its own
What it costs
  • Another store to secure, size and keep in step with the warehouse
  • Another datastore to operate, patch and monitor — justified only because the flagship question is a multi-hop traversal
  • A rebuild lag means the graph can briefly disagree with the warehouse, which must be stated
Choose differently when
Skip the graph when relationships are shallow and stable — two hops over a well-modelled star schema does not need one.
LessonAdd a graph when the questions are about relationships rather than aggregates. Keep the measures out of it, or you now have two places a number can come from.
Shown on views11 13 17 18
ADR-19

Forecasts and anomaly scores come from registered models, not prompts

Accepted

What produces the forecast behind "what happens if we do nothing?"

Context
The requirement asks for forecasting, anomaly detection, risk scoring and scenario analysis, and explicitly separates LLM reasoning from statistical calculation. A forecast presented to an executive needs an interval, a method that can be described, and reproducibility — a language model asked to project a trend supplies a number with the appearance of all three and none of the substance.
Decision
Forecasting, anomaly detection and risk scoring are registered, versioned models served behind managed endpoints. Every prediction returns an interval and the model version. The language model frames the scenario and explains the result; it never generates the projection.
How it works on-premise
MLflow holds the registry and KServe serves the models as endpoints with at least two replicas, on CPU nodes so they never contend with the inference pool for accelerators. Features come from the gold layer. Model version accompanies every prediction into the decision record, so a forecast shown four months ago can be attributed to the model that made it.
Options weighed
  • ChosenRegistered models behind managed endpoints: Versioned, testable, reproducible, and able to express uncertainty honestly.
  • RejectedAsk the language model to project the trend: Fast and plausible, with no interval, no method and no reproducibility — three properties an executive forecast cannot do without.
  • RejectedForecasting inside the semantic model: Fine for simple time intelligence and unable to carry anomaly detection or multivariate risk scoring.
  • Right elsewhereNotebooks scheduled in the batch layer: Fine for a first forecast and it produces an artefact nobody can version, roll back or attribute a four-month-old number to.
Consequences
What it buys
  • Every projection carries an interval, which is what makes "confidence" a number rather than an adjective
  • Model version is auditable alongside the decision that used it
  • Statistical quality can be improved without touching the conversational layer
What it costs
  • An ML lifecycle to run — training, registration, drift monitoring — alongside the AI lifecycle
  • Endpoint cost even at low utilisation
  • Cold-start models for a new tenant with no history, which must be declared rather than hidden
Choose differently when
Use simple statistical baselines in the semantic model when the forecast horizon is short and the series is well behaved. Do not use a language model either way.
LessonA forecast without an interval is an opinion with a decimal point. If the component cannot express uncertainty, it should not be producing the number.
Shown on views17 18
ADR-20

Freshness is declared per source and published with every answer

Accepted

What does "current" mean when twelve source systems refresh at twelve different rates?

Context
The requirement notes that not every source needs real-time integration, which is true and incomplete: the consequence is that "what changed since yesterday?" means something different for a source on fifteen-minute change capture than for one on a nightly extract. Silence about that difference is the failure mode that destroys trust fastest, because the answer looks authoritative either way.
Decision
Every source declares a freshness contract at registration. The platform records the actual as-of timestamp per source, publishes it with every answer that used the source, and raises an alert when a source falls behind its contract. Zones are organised by rebuildability, so only the operational estate carries a real recovery objective.
How it works on-premise
Airflow tasks write an as-of watermark per source into a freshness register in the gold layer. The tool plane reads it and attaches it to every result, so the composition step can state it. Prometheus alerts when a watermark exceeds its contract, and view 25's degradation contract says what the answer does in that case: serve the last good value with the date shown, up to 24 hours, then withhold.
Options weighed
  • ChosenDeclared contracts, published as-of, alert on breach: Makes staleness a visible property of an answer rather than an invisible property of the platform.
  • RejectedPresent everything as current: The default in most BI, and the reason executives quietly stop believing dashboards.
  • RejectedReal-time integration everywhere: Unaffordable and unnecessary; several source systems cannot support it at all.
  • RejectedRefuse to answer whenever anything is stale: Honest and useless — one late nightly feed would silence the whole morning brief.
Consequences
What it buys
  • An executive can tell how much weight to put on a figure
  • Stale-source incidents are detected by the platform rather than reported by the user
  • Ingestion cost is spent where freshness is actually needed
What it costs
  • Every answer carries more qualification, which takes design work to keep readable
  • A freshness register to maintain per source and per tenant
  • Some questions become unanswerable rather than answered approximately
Choose differently when
A single global freshness statement is enough when all sources genuinely refresh together. That is rare outside a single-system deployment.
LessonPublish the age of your data. An answer that hides how old it is has substituted confidence for accuracy.
Shown on views09 11 12 25

Decision and executionTurning a recommendation into a governed action, and knowing whether it worked.

ADR-21

One execution plane holds every write credential in the estate

Accepted

What is actually permitted to change a record in an enterprise system?

Context
The requirement states that AI must never have unrestricted write access. The weaker reading — give the agent a write tool and guard it with a prompt — fails to the first prompt injection through a retrieved document. The strong reading is that the AI identity should hold no write scope at all, so that there is nothing to be talked into using.
Decision
Agents may only create action proposals. A separate execution plane, holding a distinct connector identity per target system, performs the write after a human approval, with an idempotency key and an authority re-check at execution time.
How it works on-premise
Approved proposals become Temporal workflows with a lease, a retry policy and a dead-letter path. Each workflow calls the target system through a connector service — a Camel route or a thin API client — running under its own identity, scoped to one system and one operation set, with its credential leased from Vault at execution time and never held on disk. Camunda carries the human approval, escalation and delegation steps that the requirement asks for. The external reference returned by the target system is written back to the decision record. The inference and orchestrator identities have no write credential anywhere in the estate.
Options weighed
  • ChosenSeparate execution plane, connector identity per system: Prompt injection cannot reach a write path that the model has no credential for.
  • RejectedWrite tools in the agent's catalogue, guarded by policy: One identity away from a very bad afternoon, and the guard is code the model is trying to reason around.
  • RejectedOne shared service principal for all writes: Simpler to operate and makes the blast radius of one compromise the entire estate.
  • Right elsewhereEmit a task for a human to perform manually: Right for a first release into a very conservative customer, and it forfeits the closed loop and most of the value.
Consequences
What it buys
  • The AI's blast radius on write is exactly zero by construction, not by policy
  • A compromise of one connector identity is bounded to one system
  • Idempotency and retry are solved once, in one place
What it costs
  • Another component in the path from approval to effect
  • Connector identities and their scopes are a real ongoing governance burden
  • Terminal-unknown outcomes must be reconciled against the target system rather than retried
Choose differently when
Collapse the plane into the application only where writes are to a system you fully own and a compromise is contained anyway. Never collapse it where the model can read untrusted content.
LessonThe safest way to stop a model misusing a credential is not to give it one. Separate the thing that decides from the thing that acts.
Shown on views02 20 26
ADR-22

Options with cost, risk and delay — not a single recommendation

Accepted

What should the platform put in front of an executive at the moment of decision?

Context
The requirement is explicit: the executive should decide with evidence rather than simply receive an AI-generated recommendation. This is a product decision as much as an architectural one. A single recommendation asks for trust the platform has not earned and gives the executive nothing to reason with; three options with their trade-offs give them the thing they are actually paid to do.
Decision
The decision engine produces a situation summary, evidence, root cause, forecast, impact and an option set — each option carrying cost, risk and expected delay — plus a recommended option and its confidence. The executive may approve, reject, request more information, delegate or escalate, and a rejection is recorded as evidence too.
How it works on-premise
The decision service composes the option set from deterministic inputs: cost from the semantic layer, risk from the risk endpoint, blast radius from the graph, precedent from retrieval. The language model orders and narrates them. Approval thresholds and dual-approval rules are tenant configuration evaluated by OPA, not code.
Options weighed
  • ChosenOption set with quantified trade-offs: Gives the executive something to decide with, and makes the platform's reasoning inspectable.
  • RejectedSingle recommendation with a confidence score: Simpler to present and asks for trust rather than earning it; a rejected recommendation also teaches the platform nothing.
  • RejectedRaw evidence, no recommendation: Defensible and unhelpful — it recreates the problem of reconciling exports, which is what the platform is replacing.
  • DeferredAutomatic execution below a threshold: Explicitly anticipated in the requirement for low-risk actions, and correct once acceptance data exists to set the threshold honestly.
Consequences
What it buys
  • The executive's judgement is supported rather than replaced, which is what gets the platform adopted
  • Rejected options are retained, so an audit sees the alternatives too
  • Confidence becomes a computed property of the inputs rather than a tone of voice
What it costs
  • Generating credible options is much harder than generating a recommendation
  • Cost, risk and delay must each come from somewhere defensible
  • More screen space and more executive reading time per decision
Choose differently when
A single recommendation is right for high-volume, low-stakes decisions where the cost of reading three options exceeds the cost of being wrong.
LessonDecision support means supplying the trade-off, not the conclusion. The conclusion is the part the executive is accountable for.
Shown on views04 19 20
ADR-23

The loop closes on the record it opened

Accepted

How does anyone find out whether an intervention worked?

Context
The requirement asks for closed-loop monitoring — decision, action, outcome, KPI monitoring, escalate or recommend next. Most platforms stop at the action, which means the organisation learns nothing and the same intervention is proposed again next quarter with the same result.
Decision
Every executed action schedules an outcome watch on the KPI it was meant to move, over a window set per action class. The verdict — improved, unchanged, worsened or inconclusive — is written to the same decision record, and a failed intervention reopens that decision rather than starting a new one.
How it works on-premise
The decision service schedules the watch on execution. A scheduled job compares the KPI against the forecast that justified the action and writes the verdict. Reopening posts a new chat message that links to the original record, so the executive sees what changed rather than a fresh, contextless alert.
Options weighed
  • ChosenOutcome watch on the same record, with reopening: The organisation accumulates evidence about which interventions actually work.
  • RejectedStop at execution: The common implementation, and it makes the platform an expensive way to send instructions.
  • RejectedA separate outcomes dashboard: Data without a loop; nobody opens it, and it cannot reopen a decision.
  • Right elsewhereAsk the executive afterwards whether it worked: Useful as a supplementary signal and worth capturing, but it is a memory, not a measurement.
Consequences
What it buys
  • Recommendation quality becomes measurable rather than asserted
  • A failed intervention escalates instead of being quietly forgotten
  • Precedent for the option set in ADR-22 accumulates naturally
What it costs
  • Attribution is genuinely hard — many things move a KPI in fourteen days
  • Slow-moving public-sector outcomes will often return inconclusive
  • A scheduling and state-management burden per executed action
Choose differently when
Skip the loop only where actions have no measurable effect on any tracked metric — in which case, ask why the action was recommended.
Lesson"Inconclusive" must be a permitted verdict. A closed loop that can only report success is not measuring anything.
Shown on views19 20 23
ADR-24

Proactive detection is deterministic and scheduled

Accepted

What actually notices that something is going wrong?

Context
The requirement asks the platform to proactively surface exceptions, emerging risks, KPI deterioration and anomalies. The obvious implementation — an agent that continuously monitors the business — is unaffordable at any real data volume, non-deterministic, and impossible to test. The harder truth is that detection is not the difficult part: triage is. A brief with forty exceptions has surfaced none.
Decision
Detection is a scheduled deterministic pipeline: thresholds from tenant configuration, anomaly models, forecast deviation, and graph propagation from one domain's signal to another's entities. Language models narrate a signal a detector has already raised. Triage — materiality scoring, deduplication to one item per root cause, and expiring logged suppressions — decides what reaches the brief.
How it works on-premise
An hourly Spark job scores gold-layer measures against baselines and calls the anomaly endpoint; supply-chain and financial-anomaly classes are additionally event-triggered from Kafka. Signals below the materiality bar are recorded rather than discarded, so a missed exception can be traced to the threshold that hid it. Suppressions are attributed, expiring records — never silent.
Options weighed
  • ChosenDeterministic scheduled detection, generative narration: Repeatable, testable, affordable, and explainable when an executive asks why they were not told.
  • RejectedAn agent that continuously monitors: Compelling in a demonstration; unbounded cost, no reproducibility, and no way to test that it would have caught something.
  • RejectedAlert on every threshold breach: Detection without triage, which is how a brief becomes forty items and then zero readers.
  • Right elsewhereLet executives configure their own alerts: A reasonable supplement for a power user, and it puts the burden of knowing what to watch on the person who hired the platform to know.
Consequences
What it buys
  • Detection is testable against historical data — you can ask whether it would have caught last year's overrun
  • Cost is a function of data volume rather than of model calls
  • Suppression is visible and expiring, so it cannot quietly become a repealed control
What it costs
  • Detectors must be authored and tuned per tenant, which is onboarding effort
  • Novel failure modes nobody wrote a detector for are missed
  • Threshold tuning is a recurring operational task, not a one-off
Choose differently when
Use a model-driven exploratory sweep where the failure modes are genuinely unknown and volumes are small enough to afford it — as a supplement to deterministic detection, never as a replacement.
LessonDetection is cheap and triage is the product. The measure of a proactive system is what it decided not to tell you.
Shown on views17 18 23

Multi-tenancyHow one product serves many organisations without becoming many products.

ADR-25

Isolation is a purchased tier, not an engineering compromise

Accepted

How isolated is one customer from another?

Context
The requirement asks for logical or physical isolation appropriate to each customer's security tier, and names a sovereign public-sector scenario. A single answer cannot serve both a commercial group that wants a low price and a ministry that will not share a database server with anyone. Choosing pooled-only loses the sovereign customers; choosing siloed-only makes the product unaffordable for everyone else.
Decision
Two tiers from one codebase. Pooled tenants share a cluster with a per-tenant search index, row-level security, per-tenant encryption keys and a per-tenant schema in the lakehouse. Siloed tenants get their own cluster, their own GPU nodes and their own database hosts, deployed from the same Terraform and Helm charts. Sovereign public-sector tenants are siloed by default. Changing tier is re-running the onboarding pipeline against new hardware.
How it works on-premise
A shared control plane holds the tenant registry — tier, site, key references, configuration version — and never tenant business data. Terraform modules parameterised by tier provision either a namespace and a schema on shared infrastructure, or a full dedicated cluster with its own accelerators, database hosts and Vault namespace. Pooled data separation is PostgreSQL row-level security plus a per-tenant key; siloed is separate hardware end to end. The siloed tier is where the on-premise economics bite hardest: GPUs cannot be shared across an air gap, so a siloed sovereign tenant is priced with its own accelerators in the quote.
Options weighed
  • ChosenTwo tiers, one codebase, tier as a deployment parameter: Serves both commercial economics and sovereign requirements without maintaining two products.
  • RejectedPooled only: Best unit economics and it disqualifies the platform from the public-sector deals it was designed for.
  • RejectedSiloed only: Simplest isolation story, and the fixed cost floor per tenant makes mid-market customers unsellable.
  • Right elsewhereA separate directory realm per customer: Right where each customer must own the identity boundary too, and it multiplies the federation and break-glass work by the tenant count.
Consequences
What it buys
  • One product, one release, two shapes — so tenant twelve costs materially less than tenant two
  • Sovereign requirements are met without special-casing the code
  • The control plane can be operated by one team without access to any tenant's data
What it costs
  • Every feature must be tested in both shapes, which roughly doubles the release matrix
  • The pooled tier's shared index is a real leak path requiring automated proof, not review
  • Two cost models and two capacity-planning exercises
Choose differently when
Pooled-only is right for a purely commercial product with no regulated customers. Siloed-only is right when every customer is regulated and the price supports it.
LessonMake isolation a product tier with a price, not an engineering promise. Then the customer chooses the trade-off, and the architecture stops pretending it can avoid one.
Shown on views08 10 21
ADR-26

Configuration, not custom code — but the configuration is engineered

Accepted

What changes when a new customer wants a different KPI, threshold or approval limit?

Context
The requirement lists eleven things each tenant must be able to configure. "Configuration, not custom code" is the right principle and is routinely undermined in practice, because the configuration is stored as untyped rows in a database, edited live through an admin screen, unversioned, unreviewed and impossible to roll back. That is custom code with worse tooling.
Decision
Tenant configuration — KPI packs, hierarchy, agent packs, data sources, policies, approval thresholds, workflows, dashboards, risk models, terminology and prompts — is a set of versioned, schema-checked artefacts in a repository, promoted through the same rings as code. Industry packs for government, banking, energy and conglomerate are the reusable starting points.
How it works on-premise
One branch per tenant in a self-hosted GitLab configuration repository. The pipeline validates against a JSON schema and rejects unknown keys, so a typo fails the build rather than silently disabling a threshold. Argo CD applies configuration alongside the image, and a rollback is a revert. The tenant console writes through the same pipeline by opening a merge request, so a customer's Tuesday change is still versioned, reviewable and reversible.
Options weighed
  • ChosenVersioned, schema-checked artefacts in Git: Configuration gets the review, diff, rollback and audit that its blast radius deserves.
  • RejectedDatabase rows edited through an admin UI: The usual approach, and the one where nobody can say what changed last Tuesday or put it back.
  • RejectedPer-tenant code branches: Honest about the divergence and fatal to the product economics by tenant five.
  • DeferredA no-code rules engine: Attractive for the workflow and threshold subset once the schema has stabilised in production.
Consequences
What it buys
  • A misconfiguration is diffable, attributable and revertible
  • Industry packs make each successive tenant cheaper, which is the commercial thesis
  • The same promotion path for code and configuration means one release process
What it costs
  • Slower than editing a row, and customers will notice
  • A schema to maintain, version and migrate as the product evolves
  • The tenant console must write through the pipeline, which is more work than writing to a table
Choose differently when
Direct database configuration is fine for genuinely cosmetic settings with no blast radius. Anything that can change what an executive is told, or what may be approved, belongs in the pipeline.
LessonIf configuration can break production, it deserves the same rigour as code. "It is only config" is how outages are introduced by people who thought they were being careful.
Shown on views05 10 22

Sovereignty and securityResidency, key custody, and the authority chain from person to row.

ADR-27

Residency is physical first, then enforced by network, admission and key custody

Accepted

What makes a data-residency claim true rather than merely stated?

Context
For a sovereign public-sector customer, residency is the reason the programme exists. In a cloud design this has to be argued: which region, which service, whose key, what the provider can see. On-premise the first layer of the argument is simply physical — the disks are in a room the customer controls — but that is not sufficient on its own, because a workload can still be scheduled somewhere unintended, a package can still be pulled from the internet at build time, and a support engineer can still copy something out.
Decision
Residency is enforced at four layers. The hardware is the customer's, in a facility they control. Admission control refuses any workload that does not carry a residency label. The single egress path is a firewall with an explicit allow-list, and package and image feeds are mirrored inside the boundary rather than fetched. The customer holds the encryption key in its own HSM, so revocation makes the data unreadable even to an operator with disk access.
How it works on-premise
OPA Gatekeeper policies deny any pod without the residency and tenant labels, and deny host-network and privileged workloads outright. Calico network policy is default-deny east-west, with Istio mTLS carrying workload identity. One egress firewall with an FQDN allow-list is the only route out, and there is no entry on it for a model API, a package index or an MCP endpoint — Harbor and an internal package mirror serve those from inside. HashiCorp Vault fronts the customer's HSM over PKCS#11 and issues per-tenant keys for the object store, the decision store, the search index and the evidence store. Vendor support access is break-glass: requested, approved by the customer, time-boxed, and recorded.
Options weighed
  • ChosenPolicy deny, private-only networking, customer-held keys: Three independent controls, each of which fails closed, and one of which the customer holds themselves.
  • RejectedContractual and configuration assurance: Cheap, common, and unable to survive the question "what stops it happening?".
  • RejectedPlatform-managed keys: Operationally simpler and removes the customer's ability to make revocation meaningful.
  • Right elsewhereOn-premises or sovereign-partner hosting: Sometimes the only acceptable answer, and it forfeits the managed AI services this design is built on. It should be priced as a different product.
Consequences
What it buys
  • A residency claim that can be demonstrated to an auditor rather than asserted
  • Key revocation is a real customer-held control
  • An air-gapped deployment is a configuration of this design rather than a different one, because nothing in it assumes an internet route
What it costs
  • Private networking makes development and diagnosis materially harder
  • Customer-managed keys introduce a customer-caused outage mode that must be understood on both sides
  • Policy exceptions become a governed process with expiry dates
Choose differently when
Lighter controls are proportionate for a commercial tenant with no residency obligation — and the tier model in ADR-25 is what lets them have that without a second product.
LessonA sovereignty claim is only as strong as the control that fails closed when someone tries to break it. Everything else is a sentence in a contract.
Shown on views21 26 28
ADR-28

The caller's authority reaches the row; no read-everything identity exists

Accepted

Whose permissions apply when an agent reads a table?

Context
The requirement is explicit that the AI must inherit the user's authorisation context, and contrasts it with the pattern where a privileged identity reads freely and the application filters afterwards. The second pattern is far easier to build and turns every application bug into a data breach, because the only thing between one executive and another's data is a correct WHERE clause in code nobody reviews as a security control.
Decision
The caller's identity is exchanged on behalf of and carried into every downstream read. Row-level security in the semantic model and group filters in the search index are evaluated as the caller. No service identity exists that can read all tenant business data.
How it works on-premise
Keycloak performs an RFC 8693 token exchange at the Experience API, producing a downstream token carrying the user's scopes, and federates to the customer's Active Directory so the groups are the ones the enterprise already manages. The tool plane passes that token to the semantic layer, where row policies apply, and builds the search filter from the caller's group claims. Mesh workload identities are used only for the platform's own operational stores. MFA, device posture and network location gate the original sign-in.
Options weighed
  • ChosenEnd-to-end delegated identity: If the person cannot see it in the source system, nothing here can surface it to them.
  • RejectedPrivileged read identity with application-side filtering: Simpler, faster, cacheable — and every filtering bug becomes a breach.
  • RejectedCopy permissions into the platform and evaluate locally: A second permissions system to keep in step with the first, which it will not stay in step with.
  • Right elsewherePer-tenant service identity with coarse role mapping: Workable where the source system genuinely cannot express per-user permissions — and then that source is limited to aggregates with no drill-down, and the limitation is stated.
Consequences
What it buys
  • Authorisation correctness is inherited from systems that already got it right
  • No privileged identity exists to be stolen
  • The security story is one sentence a CISO can check
What it costs
  • Caching becomes hard, which ADR-14 addresses at the cost of hit rate
  • Scheduled work has no user to act as and needs its own constrained pattern
  • A source that cannot express permissions constrains what the platform can offer over it
Choose differently when
A service identity is acceptable where all data is uniformly accessible to all users of the platform. In an executive context spanning HR, finance and procurement, it never is.
LessonInherit authorisation; do not re-implement it. The second implementation is the one that will be wrong, and nobody will notice until it matters.
Shown on views26 27
ADR-29

GPU supply and facility capacity are carried as a contracted risk, not an assumption

Accepted

What happens if the accelerators this design assumes cannot be obtained, powered or cooled?

Context
The design places inference on the customer's own hardware, which is what makes the sovereignty argument work without a single caveat. That hardware is the part of the estate the programme least controls: accelerator lead times run to months and move with global demand, a dense GPU rack draws several times what a standard rack draws, and many enterprise halls cannot power or cool one without work. Quietly assuming the hardware arrives is the most likely way for this architecture to be invalidated after signature, and it is precisely the kind of assumption architecture documents bury.
Decision
Accelerator supply, and the power and cooling available in the target hall, are treated as an explicit, named risk with three pre-agreed fallbacks. Both are confirmed with the customer's facilities and procurement functions before contract, and re-confirmed before each site build. The choice of fallback belongs to the customer and is recorded in the contract rather than decided by the platform team.
How it works on-premise
The tenant registry records the permitted inference location per tenant, and the orchestrator routes model calls accordingly, so a fallback is a configuration change rather than a redeployment. Admission and network policy still deny data at rest leaving the boundary regardless of where inference runs, which keeps the residency claim intact under fallback two — though the prompt and the retrieved passages do leave under that fallback, and the record says so rather than implying otherwise. Model sizes are chosen so that the design degrades rather than fails: a 70B-class model is the target, a 7B-to-13B-class model on two cards is a working system with a stated quality gap.
Options weighed
  • ChosenNamed risk with three contracted fallbacks: The honest position: the platform team cannot control global accelerator supply or the customer's facility, so the customer decides the trade-off with the facts in front of them.
  • RejectedAssume the hardware arrives and design around it later: The common approach, and it converts a known constraint into a post-signature crisis — one that takes months rather than weeks to resolve, because it is a supply chain rather than a configuration.
  • DeferredRun a smaller open-weight model on whatever accelerators exist: Fallback one. Correct where the capability gap is acceptable, and it should be a decision with a measured quality delta rather than a default.
  • Right elsewherePlace inference on a customer-approved hosted endpoint: Fallback two. A genuine answer to a hardware problem, and it trades away the strongest version of the sovereignty claim: prompts and retrieved passages leave the boundary even though stored data does not. It must be described in exactly those terms.
Consequences
What it buys
  • The largest external dependency is visible to the sponsor before commitment rather than after
  • Routing per tenant means a fallback does not require re-architecture
  • Data residency survives fallback two, because storage location is enforced separately from inference location — and the weaker claim is stated rather than glossed
What it costs
  • A commercial conversation that is easier to postpone than to have
  • Fallback two weakens the strongest version of the sovereignty claim and must be described accurately, in writing, before it is used
  • Re-confirmation is recurring work at every site build, and hardware refresh cycles put it back on the table every few years
Choose differently when
Where the customer already runs a GPU estate for other workloads and has hall capacity to spare, this becomes a routine capacity request rather than a headline risk — but it remains a check, because a shared GPU estate has its own contention and its own scheduling politics.
LessonName the dependency that could invalidate your design, and price its alternatives before signature. An architecture document that hides its biggest assumption is marketing.
Shown on views21 28
ADR-34

First-party MCP servers only, and tool descriptions are untrusted input

Accepted

Can the platform consume an MCP server it did not write?

Context
The value of a standard is the ecosystem, and vendors now ship MCP servers for their own products. Consuming one puts somebody else’s code, or a remote endpoint outside the sovereign boundary, on the path between the model and enterprise data. Worse, a tool description is text the model reads and obeys: a server that changes its description changes the model’s behaviour without any change to this platform. For a sovereign tenant that is both an egress path and an injection path, arriving with a standards logo on it.
Decision
Only first-party MCP servers, built and deployed by the platform, run inside the trust boundary. Third-party and remotely hosted servers are not consumed. Tool names, descriptions and schemas are pinned to a reviewed server version and travel through the release pipeline; they are never read live from an external source.
How it works on-premise
Servers are containers built from the platform repository, signed with Cosign, scanned by Trivy and deployed by the same pipeline as the application, from the internal Harbor registry. The egress firewall allow-list has no entry for an external MCP endpoint, so the control is a network fact rather than a policy statement — and in an air-gapped deployment there is no route at all. Where a vendor capability is genuinely needed it is wrapped: the platform writes a thin first-party server that calls the vendor's ordinary API from inside the DMZ, and the wrapper is what the model sees.
Options weighed
  • ChosenFirst-party servers only, descriptions pinned to a reviewed version: Keeps the whole model-facing surface inside the pipeline that already signs, scans and evaluates everything else.
  • ChosenWrap a vendor API in a thin first-party server: The escape hatch that keeps the rule absolute. The platform pays for an adapter and keeps one story about whose code the model talks to.
  • RejectedConsume vendor MCP servers directly: The fastest route to breadth, and it hands an outside party the ability to change what the model does, with no gate and, if remote, an egress path the sovereignty argument cannot survive.
  • RejectedAllow-list vetted third-party servers: Sounds proportionate and is hard to hold: vetting is point-in-time while a server updates continuously, so the control decays silently.
Consequences
What it buys
  • The model-facing surface is entirely first-party, scanned, signed and evaluated
  • No egress path for an external tool endpoint, so the residency claim in ADR-27 holds unchanged
  • Tool-description injection becomes a release-gate problem rather than a run-time one
  • A vendor capability is still reachable, through an adapter the platform controls
What it costs
  • The ecosystem benefit is deliberately forgone; every integration is built rather than adopted
  • A thin adapter per vendor capability is real, recurring work
  • The platform will look slower to add integrations than a competitor that consumes servers freely
Choose differently when
A commercial tenant with no residency obligation and a small blast radius can reasonably consume vetted third-party servers, if the vetting is continuous rather than one-off. Never inside a sovereign boundary.
LessonA tool description is executable text. Treat the catalogue as code that ships through your pipeline, not as configuration you fetch — an ecosystem is a supply chain, and a standard does not make other people’s code yours.
Shown on views26 28 29

Operations and assuranceReleasing safely, degrading honestly, and proving it afterwards.

ADR-30

Evaluation is a release gate, and a prompt is a release artefact

Accepted

What stops a prompt change quietly degrading every tenant's answers?

Context
A prompt change can break accuracy as thoroughly as a code change and is far easier to make casually, often by someone improving a phrase. Model upgrades have the same property and arrive on the vendor's schedule rather than ours. Without a gate, quality regression is discovered by an executive, which is the most expensive possible detection mechanism.
Decision
Prompts are versioned artefacts released with the container image and rolled back with it. Every release runs a golden-question suite, scoring retrieval and generation separately, with a no-regression rule. Tenants may pin a model version and opt in to upgrades on their own schedule. Production failures are mined back into the suite.
How it works on-premise
The prompt registry is versioned configuration in Git, mirrored into Langfuse so a production turn can be attributed to an exact prompt version. Ragas runs the golden-question suite in the GitLab pipeline against a staging deployment; retrieval metrics and generation metrics are gated separately so a regression has a stage rather than a shrug. Ring deployment through Argo CD sends the change to one pilot tenant before all tenants, and view 23's grounding-rate signal is watched for twenty-four hours afterwards.
Options weighed
  • ChosenEvaluation as a pipeline gate, prompts as artefacts: Quality regression is caught by the pipeline rather than by a chief executive.
  • RejectedManual spot checks before release: Works for the first three releases and silently stops working around release ten.
  • RejectedPrompts editable in production configuration: Fast to fix a phrasing problem, and it makes every answer's provenance unknowable.
  • RejectedVendor benchmark scores as the gate: Measures the model on someone else's questions, not the platform on this tenant's.
Consequences
What it buys
  • A regression has a stage and an owner rather than a shrug
  • A tenant's own questions become the acceptance criteria, which is also a good sales artefact
  • Model deprecation becomes a scheduled migration instead of an emergency
What it costs
  • Evaluation runs cost real money on every release and will be the first target of a cost exercise
  • Golden sets rot and must be actively maintained from production failures
  • Release cadence slows, which the team will feel before the customer feels the benefit
Choose differently when
Lighter evaluation is proportionate for an internal assistant where a wrong answer costs a few minutes. It is not proportionate anywhere an answer becomes an approval.
LessonAnything that changes the output is a release artefact. If it can be edited in production without a gate, it will be, on a Friday.
Shown on views22 24
ADR-31

Evidence is snapshotted, and the writer cannot delete

Accepted

What does an auditor see when they ask what the approver saw?

Context
Re-running the query that produced a figure gives today's number. Four months later the underlying data has been restated, a mapping has been corrected, or a period has closed — so a decision that was correct at the time looks negligent. Separately, an audit trail that the application can delete from is not an audit trail; it is a log.
Decision
Evidence is a snapshot taken at the moment of the answer and stored immutably, never re-derived. The identity that writes the audit trail has no delete permission, and immutability is enforced by the storage service rather than by application logic. Options that were shown and rejected are retained.
How it works on-premise
Evidence payloads are written to the object store under object-lock retention with legal hold, and referenced from the decision record by identifier. A per-tenant, per-day hash chain covers the record set. The writing identity holds a write-only policy; deletion requires a separate break-glass identity that cannot write, and the object lock refuses it anyway until retention expires. Access events flow to the SIEM cluster, so reading a decision record is itself audited. Object lock is what makes this a storage guarantee rather than an application promise — the same reason the evidence store is not simply a table.
Options weighed
  • ChosenImmutable snapshots, writer cannot delete: The record survives both the passage of time and a compromise of the application that wrote it.
  • RejectedRe-run the query at audit time: The intuitive approach, and it answers a different question from the one being asked.
  • RejectedStore the query and the parameters only: Compact, and it depends on the data being unchanged, which is exactly what cannot be relied on.
  • DeferredExternal notarisation of the hash chain: Worth adding where an auditor must verify without trusting the operator at all; the internal chain plus storage immutability meets the stated requirement today.
Consequences
What it buys
  • A decision can be judged on what was known at the time, which is the only fair basis
  • An application compromise cannot rewrite history
  • The evidence pack in view 30 is a mechanical export rather than a reconstruction
What it costs
  • Storage grows with every answer, and ten-year retention makes that a real, ongoing cost
  • Immutability means mistakes are permanent, including accidentally captured sensitive content
  • A separate deletion identity is operational friction, deliberately
Choose differently when
Storing the query alone is acceptable where the underlying data is genuinely immutable — an append-only event log, for instance. Very little enterprise finance data qualifies.
LessonAn audit answers "what did they know then?", not "what is true now?". Systems that re-query at audit time are answering the wrong question confidently.
Shown on views06 11 30
ADR-32

A written degradation contract, and one place the platform fails closed

Accepted

What does an executive get at 07:10 when a dependency is down?

Context
The requirement asks for graceful degradation when an enterprise system is unavailable. That is only real if it is written down per dependency and per capability, agreed with the tenant and exercised — otherwise it resolves to a 500 during the morning brief, which is when the platform's reputation is actually set.
Decision
A degradation matrix states, per dependency and per capability, what still works, what degrades and what stops. Every degraded state names itself in the answer. The platform fails open with a stated caveat everywhere except one case: if the decision cannot be recorded, it refuses to act rather than acting unrecorded.
How it works on-premise
Health of each dependency is evaluated by the tool plane before composition, and the resulting capability state is attached to the answer. Cached KPI values are served up to twenty-four hours with the as-of date shown, then withheld. Approvals wait as Temporal workflows when the execution plane is unavailable and are re-authorised on return. Degradation states are exercised in a quarterly game day per tenant tier — which on-premise includes pulling a power feed and failing a disk, not only stopping a service.
Options weighed
  • ChosenWritten matrix, fail open with a caveat, fail closed on the record: Every state is a decision someone made, and the one exception is where an unrecorded action would be worse than no action.
  • RejectedFail closed everywhere: Safest and it makes a single stale nightly feed silence the entire morning brief.
  • RejectedFail open silently: The worst option: a confidently partial answer is more dangerous than an error, because nothing signals the gap.
  • RejectedBest-effort, undocumented: The default when nobody writes the matrix, and it means the behaviour is discovered during an incident.
Consequences
What it buys
  • Degradation is a designed behaviour with an owner rather than an emergent one
  • The tenant knows in advance what an outage looks like, which is a contractual conversation had calmly
  • The one fail-closed case is defensible precisely because it is the only one
What it costs
  • Every capability needs a degraded path built and tested, which is real engineering
  • Game days cost time per tenant tier
  • More states in the user interface to design and explain
Choose differently when
Fail closed more broadly where a partial answer carries regulatory consequence — clinical or safety-critical contexts, for instance, where a caveat is not sufficient protection.
LessonDecide where you fail open and where you fail closed, write it down, and test it. A system without a degradation contract has one anyway; it just has not been reviewed.
Shown on views23 25

Every package used, in one table

Every package and open protocol on these thirty views, what it is, and what it is doing here. The last column names what was considered instead, so the table doubles as a shortlist for anyone adapting this design to a different estate — or to a cloud one.

PackageWhat it isWhat it does hereConsidered instead
Model Context Protocol (MCP) An open protocol for exposing typed tools to a language model The contract the tool plane speaks: first-party servers, one per domain, tools resolved per role A bespoke REST tool contract behind the gateway
HAProxy + Coraza A load balancer and an open-source web application firewall TLS termination, the OWASP rule set, and the only route in from the user network NGINX with ModSecurity; an enterprise load balancer appliance
Kong Gateway An API gateway with a plugin model Per-tenant rate and token quotas, contract validation, and the tool plane’s front door Apache APISIX; ingress plus custom middleware
Kubernetes A container orchestrator Runs the experience API, orchestrator, MCP servers, decision, config and governance services OpenShift for vendor support; plain hosts under Ansible
Istio A service mesh Workload identity, mutual TLS on every hop, and traffic policy between services Cilium; per-service TLS
NVIDIA GPU Operator Kubernetes lifecycle management for GPU nodes Drivers, device plugin and MIG partitioning for the inference pool Manual driver management on dedicated inference hosts
vLLM A high-throughput inference server for open-weight models Serves the reasoning model and the small routing model with continuous batching TGI, TensorRT-LLM, llama.cpp; Ollama for development
Text Embeddings Inference (TEI) A server for embedding and reranking models BGE-M3 embeddings and a cross-encoder reranker for retrieval Running the models in-process; a vector database’s built-in embedder
Llama Guard 3 and an NLI verifier A safety classifier and an entailment model Input shields, and independent verification that every claim is supported by retrieved evidence A second large model as judge; a commercial moderation service
Presidio An open-source PII detection and anonymisation library Detects personal data in a question before it reaches the model or the log Regular expressions; a commercial DLP agent
Docling and Apache Tika Document parsing and layout extraction toolkits Turn PDFs, contracts and board papers into structured, chunkable content A commercial document-understanding appliance
LangGraph A library for graph-structured, checkpointed agent execution Agent threads and plan execution, with state that survives a pod restart AutoGen; a hand-written tool-calling loop
Langfuse Self-hosted LLM observability and prompt management Turn-level traces, prompt versions, and the datasets the evaluation suite draws on Traces in the general observability stack only
Ragas An evaluation library for retrieval-augmented systems Scores retrieval and generation separately as the release gate Hand-scored spot checks; published benchmarks
Temporal A durable workflow engine Long-running investigations, and the exactly-once execution of an approved action A queue plus a hand-written state machine
Camunda 8 A BPMN workflow and decision engine Approval, escalation, delegation and SLA monitoring as business rules Approval logic inside the decision service
Apache Camel An integration framework implementing the enterprise integration patterns The connector routes the execution plane uses to write to target systems A bespoke API client per system
MinIO S3-compatible object storage for on-premise hardware The lake’s storage layer, and the immutable evidence store under object lock Ceph object gateway; an enterprise NAS
Apache Iceberg An open table format with snapshots and schema evolution Table semantics and time travel over the object store, for all three medallion zones Delta Lake; Hudi; plain Parquet
Trino A distributed SQL query engine that federates across catalogues Interactive query over the lakehouse, and reads the customer’s existing lake in place Presto; Dremio; Spark SQL alone
Apache Spark A distributed processing engine Batch transformation, entity resolution and the hourly detection sweep Trino for everything; Flink for the streaming paths
Apache Airflow + dbt A workflow scheduler and a SQL transformation framework Batch and change-data ingestion, the medallion build, and per-source freshness watermarks Dagster; scheduled scripts
Cube A semantic layer defining measures and row-level policies as code The single definition of every KPI, shared with the customer’s own reporting dbt metrics; measures defined only in the BI tool
Apache Superset An open-source BI and exploration tool Dashboards over the same semantic layer the platform answers from Metabase; the customer’s existing BI estate
OpenSearch A search engine with BM25 and k-NN vector support Enterprise retrieval with hybrid queries and index-side security trimming; a second cluster serves as the SIEM Elasticsearch; pgvector; a dedicated vector database
Neo4j A property graph database Canonical entity graph for multi-hop questions and signal propagation Apache AGE on PostgreSQL; recursive SQL
PostgreSQL + Patroni A relational database with automated high availability The decision, evidence, action and outcome record, with row-level security and point-in-time restore A commercial RDBMS the customer already licenses
MLflow + KServe A model registry and a Kubernetes model-serving runtime Forecasting, anomaly detection and risk scoring as versioned, served artefacts Seldon Core; Ray Serve; notebooks on a schedule
Redis An in-memory data store Conversation state and a semantic cache keyed on the caller’s permission fingerprint Valkey; state in the decision store
Apache Kafka + Debezium An event log and a change-data-capture connector Event ingestion and low-latency change feeds from source databases RabbitMQ; NATS; nightly extracts only
Keycloak An identity and access management server Sign-in, federation to Active Directory, and the token exchange that carries the caller’s authority to the row The directory’s own federation service; a commercial IAM product
HashiCorp Vault + HSM A secrets manager fronting a hardware security module Customer-held encryption keys per tenant, and short-lived credentials for the execution plane OpenBao; the cluster’s own secret store
Open Policy Agent A general-purpose policy engine, with Gatekeeper for admission control Admission rules that enforce residency and isolation, and the authorisation decisions behind approval thresholds Kyverno for admission; policy asserted in code review
OpenMetadata + OpenLineage A data catalogue and an open lineage specification Sensitivity labels captured at ingestion, and source-to-measure lineage in the evidence pack DataHub; lineage inferred from pipeline metadata
Prometheus, Grafana, Loki, Tempo Metrics, dashboards, logs and distributed traces End-to-end turn traces following OpenTelemetry GenAI conventions, plus GPU-hours per interaction A commercial APM installed on-premise
Falco + Trivy + Cosign Runtime threat detection, vulnerability scanning and artefact signing Posture across the cluster, and proof that only reviewed images run Wazuh; a commercial container security platform
Harbor A container registry with scanning and replication The internal registry and the mirror through which every image and package enters the boundary Nexus; a registry on the build server
GitLab + Argo CD A self-hosted source and CI platform, and a GitOps reconciler Two pipelines — product code and tenant configuration — and a cluster rebuilt from Git Jenkins with Helm; deployment by hand at a change window
Terraform, Helm, Ansible Infrastructure, package and configuration automation Tenant onboarding, with the isolation tier as a deployment parameter Manual provisioning; a per-tenant repository fork
Velero Backup and restore for Kubernetes cluster state Cluster-level recovery alongside WAL archive for the database and replication for objects Storage-array snapshots; an enterprise backup product
Mattermost A self-hosted team chat platform Delivers the morning brief and approvals as interactive messages inside the boundary Rocket.Chat; email deep-links; the web app only
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.