Multimodal RAG Platform

Solution Architecture v1.0 · Enterprise Architecture · 2026-08

A platform that ingests text, PDFs, images, tables, audio and video, and answers natural-language questions across all of it with grounded, cited responses.

16 views 16 HTML views16 SVG16 draw.io 2 documents Updated 2026-10-07
Architecture views

16 views, each in three formats.

Open a view to read it in full. Every SVG carries its diagram source inside it, so it opens in diagrams.net fully editable with no import step; the draw.io files are the same diagrams as plain source.

  1. 01
    System Context

    C4 Level 1. Who uses the platform, which content sources feed it, and which external services it depends on.

  2. 02
    High-Level Architecture

    The whole platform in six stages: ingest, understand, index, retrieve, generate, serve. The picture a sponsor should understand without narration.

  3. 03
    Layered Architecture

    Six layers with dependencies pointing downward only. Establishes what may call what.

  4. 04
    Container Architecture

    C4 Level 2. Deployable units, their technology, and the synchronous query path between them.

  5. 05
    Multimodal Ingestion Pipeline

    One lane per modality across five stages. The lanes run in parallel and converge on a single chunk record.

  6. 06
    Indexing & Embedding Pipeline

    How a normalised chunk becomes a retrievable, governed index entry.

  7. 07
    Retrieval & Generation Flow

    The query path from admitted question to verified, cited answer.

  8. 08
    Data Architecture

    Four zones by ownership and rebuildability. Only one of them is irreplaceable.

  9. 09
    Integration Architecture

    Every inbound and outbound interface, with protocol, cadence and direction.

  10. 10
    Security Architecture

    Five trust zones in decreasing exposure, with every boundary crossing labelled — including the attacker's path.

  11. 11
    Deployment & Infrastructure

    What runs where, across two availability zones and a warm standby region.

  12. 12
    CI/CD & Environment Promotion

    How a change reaches production, and what stops a bad one — including a bad prompt.

  13. 13
    Observability & Evaluation Matrix

    Six signal types across five stages. A matrix rather than a flow, because it makes the gaps visible.

  14. 14
    RAGOps Lifecycle

    The improvement loop, drawn as a loop. A RAG system is never finished at launch.

  15. 15
    Critical Flow — Question to Cited Answer

    The end-to-end request as ordered messages across eight participants, including the abstain path.

  16. 16
    Core Data Model

    Nine entities carrying multi-tenancy, multimodal provenance and citation lineage.

Documents

The written architecture, on the page.

The view index above is the map; this is the argument. The one-pager and the decision record are part of the deliverable, so they are printed here in full — each also opens as its own page with a table of contents.

Document 1 of 2 · 10 min read

Solution Architecture Document — Multimodal RAG Platform

Version 1.0 · Status For review · Date 2026-08 · Owner Enterprise Architecture


1. Purpose and scope

Build a platform that ingests heterogeneous enterprise content — text, PDFs, images, tables, audio and video — and answers natural-language questions against it with grounded, cited responses.

In scope: multimodal ingestion and understanding, embedding and indexing, hybrid retrieval, LLM generation with citation enforcement, a web interface for upload and query, multi-tenancy, security, observability and the evaluation loop.

Out of scope for v1: real-time streaming ingestion of live media, agentic multi-step tool use beyond retrieval, fine-tuning of foundation models, and write-back into source systems.

2. Drivers

Driver Consequence for the architecture
Answers must be grounded and attributable Provenance is carried in the data model, not bolted on. A grounding gate can refuse to answer.
Content is genuinely heterogeneous One extraction path per modality, converging on one chunk record.
Enterprise corpora are full of identifiers Lexical retrieval is a first-class channel, not a fallback.
Questions include aggregation over tables Tables are stored twice: as prose for retrieval, as typed rows for SQL.
Quality drifts and must be managed Evaluation is a release gate and a running loop, not a launch activity.
Data is sensitive and multi-tenant Isolation enforced at the retriever, not at the caller.

3. Architecture overview

The platform separates a write path (asynchronous, minutes-scale, bursty) from a read path (synchronous, seconds-scale, latency-sensitive). They share only storage. This is the single most important structural decision: the two have different scaling profiles, different failure semantics and different cost drivers.

Write path   sources → intake → raw store → per-modality understanding →
             chunk + enrich → embed → vector | lexical | typed tables

Read path    question → guard → rewrite/route → hybrid search →
             fuse + rerank → parent expand → generate → grounding gate → answer

See view 02-high-level-architecture for the picture, 05 and 06 for the write path, 07 and 15 for the read path.

4. Key design decisions

4.1 One chunk record for every modality

Every source, regardless of type, normalises into a common record:

Field Purpose
text The retrievable text projection — the prose, the caption, the transcript, the linearised table
modality What it originally was, used for routing and filtering
provenance page, char_start/char_end, bbox, timecode_ms
parent_chunk_id The larger section returned to the model at generation time

This is what makes a question answerable across modalities at all: the retrieval layer only ever sees one kind of thing. It is also what makes a citation clickable — the provenance columns point at a specific page, region or moment.

4.2 Text-projection-first retrieval, with a native image channel alongside

Images are indexed twice. A vision model writes a description of the image, and any text in it is OCR'd; both are embedded as text. Separately, a joint image-text embedding is stored.

Rationale: for the dominant question type — "what does the Q3 revenue chart show" — a generated description retrieves far better than a joint image embedding, because it captures semantics the embedding compresses away. The native image vector is retained for the minority case: "find the diagram that looks like this". Two channels, fused at rerank.

4.3 Tables are stored twice

A linearised markdown chunk enters the vector and lexical indexes. The actual typed rows land in an analytical table addressable by SQL.

Rationale: retrieval over a linearised table cannot answer "what was the total across all regions" — the number is not in any chunk. Numeric questions route to generated SQL instead. Without this, the platform silently produces confident wrong arithmetic.

4.4 Hybrid retrieval, not vector-only

Dense and sparse searches run concurrently and are merged with reciprocal rank fusion, then reranked by a cross-encoder.

Rationale: dense retrieval fails on part numbers, error codes, policy identifiers and personal names — exactly the tokens enterprise users search for. RRF needs no score calibration between the channels, which is why it is preferred over weighted score fusion.

4.5 Contextual chunk enrichment

At index time, each chunk is prefixed with a short generated line situating it in its document. This is a one-off indexing cost that materially raises retrieval accuracy, because a chunk that says "the rate increased to 4.2%" becomes findable by a query about the specific policy it belongs to. It runs on Claude Haiku 4.5 through the Batch API at half the standard token rate.

4.6 Small-to-big retrieval

Small chunks are embedded for retrieval precision; the larger parent section is what reaches the model. Precision where it is searched, context where it is read.

4.7 The grounding gate

After generation, every claim is checked against a retrieved span. Unsupported claims trigger one retry with a rewritten query; failing that, the platform abstains and says so.

Rationale: the requirement is grounded answers. Grounding that is hoped for is not grounding. Abstention is a logged, measured outcome — the abstention rate is a quality metric, not a defect.

4.8 A model gateway, not direct SDK calls

Every model call leaves through one gateway that owns routing, fallback, semantic caching, per-tenant budget caps, payload redaction and retry policy.

Rationale: without it, model choice, cost control and redaction are scattered across a dozen services and cannot be changed or audited in one place.

4.9 Isolation enforced at the retriever

Tenant and ACL filtering is applied inside the retriever, using claims from the verified token — never by a filter the caller supplies. A caller that can request an unfiltered search is a caller that can leak another tenant's corpus.

5. Model selection

Every generation call goes through the model gateway, which routes by task:

Task Model Rationale
Answer generation claude-opus-5 Strongest synthesis and instruction-following; 1M context leaves room for expanded parents. Citations enabled on document blocks.
Vision — image and chart description claude-opus-5 Indexing quality compounds; a poor description is permanently unretrievable.
Contextual chunk headers claude-haiku-4-5 High volume, mechanically simple. Runs on the Batch API at 50% cost.
Query rewrite, routing, classification claude-haiku-4-5 Latency-critical, low complexity.

Adaptive thinking (thinking: {type: "adaptive"}) is on for generation, with output_config.effort tuned per route — high for answer synthesis, low for classification. Prompt caching holds the stable system prompt and tool definitions ahead of the volatile retrieved context, so the cacheable prefix is not invalidated per request.

Embeddings are not an Anthropic surface; a dedicated embedding provider supplies the text and image vectors, called through the same gateway.

6. Security architecture

See view 10-security-trust-zones.

Five zones in decreasing exposure: Internet, perimeter, application, data, control plane.

Threats explicitly modelled:

  • Malicious upload. Files are quarantined and scanned before any parser touches them. Parsers run with no network access and a strict resource budget.
  • Prompt injection inside document text. This is the defining threat of RAG. Retrieved content is fenced and labelled as data in the prompt; the content guard strips instruction-shaped text at index time; the generation prompt states that retrieved content is never an instruction. Defence is layered because no single layer is reliable.
  • Cross-tenant leakage. Namespace per tenant in the vector index, row-level security on metadata, filtering enforced inside the retriever.

Controls: TLS 1.3 inbound, mTLS between services, OIDC with JWKS verification at the gateway, customer-managed encryption keys, secrets rotated every 90 days, write-once audit log forwarded to the SIEM, egress restricted to an allow-list.

7. Non-functional requirements

Attribute Target How it is met
Query latency p95 < 6 s end to end; first token < 1.5 s Concurrent retrieval fan-out, streaming responses, semantic cache
Retrieval latency p95 < 400 ms for search + rerank HNSW index, GPU reranker, top-50 candidate cap
Ingestion throughput 10k documents/hour sustained Queue-depth autoscaling on stateless workers
Ingestion latency 95% of documents queryable within 15 min Parallel per-modality lanes
Availability 99.5% monthly for the query path Two AZs, stateless services, managed data services
RPO / RTO 15 min / 4 h Cross-region object replication, index rebuild from derived zone
Groundedness > 95% of claims traceable to a span Grounding gate as release criterion
Retrieval recall@10 > 0.85 on the golden set Hybrid retrieval, contextual enrichment, eval gate in CI
Tenant isolation Zero cross-tenant retrievals Enforced in retriever; asserted in automated tests

8. Cost model

Cost is dominated by two lines: generation tokens on the read path and enrichment plus embedding on the write path.

Lever Effect
Route classification and rewrite to Haiku Removes the majority of calls from the Opus rate
Batch API for chunk enrichment 50% reduction on the largest write-path token line
Prompt caching of the stable prefix Large reduction in repeated input tokens per query
Semantic cache on the read path Removes repeat questions entirely
Top-8 after rerank rather than top-20 Directly reduces input tokens per generation

Per-tenant cost is metered at the model gateway and enforced as a budget cap. Cost per query is a CI gate (view 12), so a change that improves quality by tripling spend does not ship unnoticed.

9. Risks

# Risk Impact Mitigation
R1 Prompt injection via ingested content High Layered: index-time stripping, prompt fencing, grounding gate
R2 Poor extraction on scanned or complex PDFs silently degrades answers High Extraction confidence scored and logged; low-confidence documents flagged, not silently indexed
R3 Vector index rebuild time exceeds RTO as the corpus grows Medium Rebuild time measured monthly; switch to replication when it approaches 4 h
R4 Generation cost grows faster than usage Medium Cost gate in CI, per-tenant caps, routing and cache tuning
R5 Golden set ages and stops representing real queries Medium Curation is a standing weekly activity, fed from real failures
R6 Embedding model upgrade requires full reindex Medium Model version pinned in the catalog; derived zone makes reindex a batch job, not a re-ingest

10. Assumptions

  • Source systems expose delta or change APIs; full re-crawls are not required per cycle.
  • Document ACLs are available at ingest and are authoritative for retrieval filtering.
  • A corporate OIDC provider exists and can issue tokens carrying tenant and group claims.
  • Initial corpus is on the order of 10^6 chunks per tenant, not 10^9.

11. Roadmap

Phase Duration Outcome
0 — Foundation 4 weeks Landing zone, IaC, CI/CD skeleton, identity integration
1 — Text and PDF vertical slice 6 weeks Upload, parse, index, hybrid retrieve, cited answer. Golden set v1 and the eval harness.
2 — Images and tables 5 weeks Vision descriptions, table extraction, text-to-SQL for numeric questions
3 — Audio and video 4 weeks Transcription, timecoded chunks, media citations that seek to the moment
4 — Hardening 5 weeks Multi-tenancy at scale, DR exercise, penetration test, cost tuning
5 — Operate ongoing RAGOps loop (view 14) as standing practice

Phase 1 is deliberately a full vertical slice rather than a horizontal layer. It proves the grounding gate and the eval harness — the two things most likely to invalidate the design — before the modality-specific work multiplies the cost of changing them.

12. Traceability

Requirement Views Section
Ingest text, PDF, images, tables, audio, video 05 4.1
Embedding and indexing 06 4.2, 4.3, 4.5
Intelligent retrieval 07, 15 4.4, 4.6
LLM generation layer 07, 15 4.7, 5
Simple upload and query interface 01, 03, 04 3
Grounded answers with citations 07, 15, 16 4.1, 4.7
Cross-source questions 02, 05, 08 4.1
Document 2 of 2 · 6 min read

Architecture Decision Records — Multimodal RAG Platform

Each record states the decision, what was rejected, and what would make us revisit it.


ADR-001 — Normalise every modality into one chunk record

Status Accepted

Context Six input modalities with nothing structurally in common. Retrieval, ranking, citation and access control each need a uniform object to operate on.

Decision Every extraction path emits the same record: a text projection, a modality tag, a provenance pointer and an optional parent reference. Modality-specific detail lives in the asset entity, not in the retrieval path.

Rejected Per-modality indexes with a federated ranker. Rejected because ranking across independently-scored indexes requires score calibration that drifts, and because citation and ACL logic would be implemented six times.

Consequences Retrieval, reranking and generation are modality-agnostic. The cost is that each new modality must be reduced to text well enough to be findable — which is a real constraint on, for example, purely visual content with no OCR-able text.

Revisit if a modality arrives that genuinely cannot be projected to text without losing the thing users search for.


ADR-002 — Index images as generated descriptions plus a native image vector

Status Accepted

Context Images carry meaning that OCR alone misses — chart trends, diagram topology, photographs. A joint image-text embedding retrieves visually similar images but poorly captures "what does this chart say".

Decision Two channels. A vision model writes a description; that description plus any OCR text is embedded as text and becomes the primary retrieval channel. A joint image-text embedding is stored as a secondary channel for visual-similarity queries. The channels are fused at rerank.

Rejected Image embeddings only — fails the dominant question type. Descriptions only — fails "find the diagram that looks like this".

Consequences Indexing cost per image is higher and includes a model call. Description quality is now permanent index quality: a bad description makes an image permanently unretrievable, so this call uses the strong model, not the cheap one.

Revisit if joint embedding models close the semantic gap enough to make descriptions redundant.


ADR-003 — Store tables twice: linearised for retrieval, typed for SQL

Status Accepted

Context Users ask both "what does the pricing table say about tier 3" (retrieval) and "what is the total across all regions" (aggregation). The second has no answer in any chunk.

Decision Tables are linearised to markdown and indexed as ordinary chunks, and separately loaded as typed rows into an analytical store. Numeric and aggregate questions are routed to generated SQL against those rows.

Rejected Retrieval-only over linearised tables. Rejected because it produces confident arithmetic that is wrong — the worst possible failure mode for a system whose value proposition is groundedness.

Consequences A routing decision is now required at query time, and generated SQL must be sandboxed and validated. Duplicate storage of table content is accepted.

Revisit if table sizes make row-level loading impractical at corpus scale.


ADR-004 — Hybrid retrieval with reciprocal rank fusion and cross-encoder reranking

Status Accepted

Context Enterprise queries contain identifiers — part numbers, error codes, policy references — where dense retrieval reliably fails. Pure lexical retrieval fails on paraphrase.

Decision Dense and sparse retrieval run concurrently, each returning 50 candidates. Results merge by reciprocal rank fusion, then a cross-encoder reranks to the top 8.

Rejected Weighted score fusion — requires calibration between incomparable score distributions and drifts as either index changes. RRF operates on ranks and needs none.

Consequences Two indexes to maintain and keep consistent. A GPU reranker on the latency-critical path, budgeted at 400 ms for search plus rerank.

Revisit if measured recall from a single channel matches the hybrid within tolerance on the golden set.


ADR-005 — Prepend generated context to each chunk at index time

Status Accepted

Context An isolated chunk often lacks the terms a user would search by. "The rate increased to 4.2%" does not mention which policy or which year.

Decision At index time, generate one short line situating the chunk in its document and prepend it before embedding.

Rejected Larger chunks — improves context at a direct cost to retrieval precision. Query-time expansion only — does not fix a chunk that is unfindable in the first place.

Consequences One model call per chunk at index time. Mitigated by using Claude Haiku 4.5 via the Batch API at half rate. Reindexing is required if the enrichment prompt changes, which is why the derived zone exists.


ADR-006 — Small-to-big retrieval via parent references

Status Accepted

Context Precision favours small chunks; generation quality favours large context.

Decision Embed small chunks; on a hit, return the parent section to the model.

Consequences chunk.parent_chunk_id is load-bearing. Deduplication is needed when several child hits share a parent.


ADR-007 — Enforce grounding before an answer is released

Status Accepted

Context The requirement is grounded, cited answers. Generation with retrieved context raises groundedness but does not guarantee it.

Decision A verification step checks each claim against a retrieved span. Failure triggers one retry with a rewritten query; a second failure returns an explicit abstention.

Rejected Trusting the generation prompt. Rejected because an unenforced instruction is not a control, and because groundedness then cannot be measured as a release gate.

Consequences Added latency and an extra model call on the read path. Abstention rate becomes a tracked quality metric — a rising rate means retrieval degraded, which is exactly the signal wanted.


ADR-008 — All model calls through a single gateway

Status Accepted

Context Model choice, cost control, redaction, caching, fallback and retry policy would otherwise be reimplemented in every service that calls a model.

Decision One gateway owns routing by task, semantic caching, per-tenant budget caps, payload redaction, provider fallback and retry. No service holds a provider credential.

Consequences The gateway is on the critical path and must be highly available. In exchange, a model swap, a price change or a redaction rule is a single-place change, and egress is auditable at one point.


ADR-009 — Separate the asynchronous write path from the synchronous read path

Status Accepted

Context Ingestion is bursty and minutes-scale; querying is steady and seconds-scale.

Decision Two independently scaled domains sharing only storage. Ingestion is queue-driven with idempotent, resumable workers; querying is a synchronous request path.

Consequences Eventual consistency between upload and queryability, bounded at 15 minutes and stated as an NFR. Users are shown ingestion status rather than left to guess.


ADR-010 — Enforce tenant isolation inside the retriever

Status Accepted

Context Multi-tenant corpora in a shared index. A filter supplied by the caller is a filter an attacker can omit.

Decision The retriever derives tenant and ACL filters from verified token claims and applies them itself. Vector namespaces are partitioned per tenant; metadata uses row-level security.

Consequences The retriever cannot be called in an unfiltered mode, including by internal batch jobs, which must present a service identity scoped to a tenant. Cross-tenant leakage is asserted in automated tests as a release gate.


ADR-011 — Treat prompts, chunking configuration and eval sets as versioned code

Status Accepted

Context A prompt edit or a chunk-size change moves answer quality as much as a code change, but conventionally bypasses every release control.

Decision These artefacts live in the repository, are versioned with the code, and pass the same CI gates — including the RAG eval gate and the cost gate.

Consequences No hot-editing prompts in production. The eval harness must be fast enough to run per pull request, which caps golden-set size and argues for a tiered set: a fast tier in CI, a full tier nightly.


ADR-012 — Rebuild indexes on failover rather than replicate them

Status Accepted

Context RTO is 4 hours. Cross-region replication of vector and lexical indexes roughly doubles their cost.

Decision Replicate the object store and the metadata database. Rebuild vector and lexical indexes in the DR region from the derived zone on failover.

Consequences Failover is slower but within RTO. The rebuild path must be exercised, not assumed — measured quarterly, and the decision reverts to replication if rebuild time approaches the RTO.

The package

Everything as it was delivered.

These files are served exactly as they were produced — the diagram pages keep their own house style because that is the artifact, not a rendering of it.