Solution Architecture Document — Multimodal RAG Platform
Solution Architecture v1.0 · Enterprise Architecture · 2026-08
Version 1.0 · Status For review · Date 2026-08 · Owner Enterprise Architecture
1. Purpose and scope
Build a platform that ingests heterogeneous enterprise content — text, PDFs, images, tables, audio and video — and answers natural-language questions against it with grounded, cited responses.
In scope: multimodal ingestion and understanding, embedding and indexing, hybrid retrieval, LLM generation with citation enforcement, a web interface for upload and query, multi-tenancy, security, observability and the evaluation loop.
Out of scope for v1: real-time streaming ingestion of live media, agentic multi-step tool use beyond retrieval, fine-tuning of foundation models, and write-back into source systems.
2. Drivers
| Driver | Consequence for the architecture |
|---|---|
| Answers must be grounded and attributable | Provenance is carried in the data model, not bolted on. A grounding gate can refuse to answer. |
| Content is genuinely heterogeneous | One extraction path per modality, converging on one chunk record. |
| Enterprise corpora are full of identifiers | Lexical retrieval is a first-class channel, not a fallback. |
| Questions include aggregation over tables | Tables are stored twice: as prose for retrieval, as typed rows for SQL. |
| Quality drifts and must be managed | Evaluation is a release gate and a running loop, not a launch activity. |
| Data is sensitive and multi-tenant | Isolation enforced at the retriever, not at the caller. |
3. Architecture overview
The platform separates a write path (asynchronous, minutes-scale, bursty) from a read path (synchronous, seconds-scale, latency-sensitive). They share only storage. This is the single most important structural decision: the two have different scaling profiles, different failure semantics and different cost drivers.
Write path sources → intake → raw store → per-modality understanding →
chunk + enrich → embed → vector | lexical | typed tables
Read path question → guard → rewrite/route → hybrid search →
fuse + rerank → parent expand → generate → grounding gate → answer
See view 02-high-level-architecture for the picture, 05 and 06 for the write path,
07 and 15 for the read path.
4. Key design decisions
4.1 One chunk record for every modality
Every source, regardless of type, normalises into a common record:
| Field | Purpose |
|---|---|
text |
The retrievable text projection — the prose, the caption, the transcript, the linearised table |
modality |
What it originally was, used for routing and filtering |
| provenance | page, char_start/char_end, bbox, timecode_ms |
parent_chunk_id |
The larger section returned to the model at generation time |
This is what makes a question answerable across modalities at all: the retrieval layer only ever sees one kind of thing. It is also what makes a citation clickable — the provenance columns point at a specific page, region or moment.
4.2 Text-projection-first retrieval, with a native image channel alongside
Images are indexed twice. A vision model writes a description of the image, and any text in it is OCR'd; both are embedded as text. Separately, a joint image-text embedding is stored.
Rationale: for the dominant question type — "what does the Q3 revenue chart show" — a generated description retrieves far better than a joint image embedding, because it captures semantics the embedding compresses away. The native image vector is retained for the minority case: "find the diagram that looks like this". Two channels, fused at rerank.
4.3 Tables are stored twice
A linearised markdown chunk enters the vector and lexical indexes. The actual typed rows land in an analytical table addressable by SQL.
Rationale: retrieval over a linearised table cannot answer "what was the total across all regions" — the number is not in any chunk. Numeric questions route to generated SQL instead. Without this, the platform silently produces confident wrong arithmetic.
4.4 Hybrid retrieval, not vector-only
Dense and sparse searches run concurrently and are merged with reciprocal rank fusion, then reranked by a cross-encoder.
Rationale: dense retrieval fails on part numbers, error codes, policy identifiers and personal names — exactly the tokens enterprise users search for. RRF needs no score calibration between the channels, which is why it is preferred over weighted score fusion.
4.5 Contextual chunk enrichment
At index time, each chunk is prefixed with a short generated line situating it in its document. This is a one-off indexing cost that materially raises retrieval accuracy, because a chunk that says "the rate increased to 4.2%" becomes findable by a query about the specific policy it belongs to. It runs on Claude Haiku 4.5 through the Batch API at half the standard token rate.
4.6 Small-to-big retrieval
Small chunks are embedded for retrieval precision; the larger parent section is what reaches the model. Precision where it is searched, context where it is read.
4.7 The grounding gate
After generation, every claim is checked against a retrieved span. Unsupported claims trigger one retry with a rewritten query; failing that, the platform abstains and says so.
Rationale: the requirement is grounded answers. Grounding that is hoped for is not grounding. Abstention is a logged, measured outcome — the abstention rate is a quality metric, not a defect.
4.8 A model gateway, not direct SDK calls
Every model call leaves through one gateway that owns routing, fallback, semantic caching, per-tenant budget caps, payload redaction and retry policy.
Rationale: without it, model choice, cost control and redaction are scattered across a dozen services and cannot be changed or audited in one place.
4.9 Isolation enforced at the retriever
Tenant and ACL filtering is applied inside the retriever, using claims from the verified token — never by a filter the caller supplies. A caller that can request an unfiltered search is a caller that can leak another tenant's corpus.
5. Model selection
Every generation call goes through the model gateway, which routes by task:
| Task | Model | Rationale |
|---|---|---|
| Answer generation | claude-opus-5 |
Strongest synthesis and instruction-following; 1M context leaves room for expanded parents. Citations enabled on document blocks. |
| Vision — image and chart description | claude-opus-5 |
Indexing quality compounds; a poor description is permanently unretrievable. |
| Contextual chunk headers | claude-haiku-4-5 |
High volume, mechanically simple. Runs on the Batch API at 50% cost. |
| Query rewrite, routing, classification | claude-haiku-4-5 |
Latency-critical, low complexity. |
Adaptive thinking (thinking: {type: "adaptive"}) is on for generation, with
output_config.effort tuned per route — high for answer synthesis, low for
classification. Prompt caching holds the stable system prompt and tool definitions ahead of
the volatile retrieved context, so the cacheable prefix is not invalidated per request.
Embeddings are not an Anthropic surface; a dedicated embedding provider supplies the text and image vectors, called through the same gateway.
6. Security architecture
See view 10-security-trust-zones.
Five zones in decreasing exposure: Internet, perimeter, application, data, control plane.
Threats explicitly modelled:
- Malicious upload. Files are quarantined and scanned before any parser touches them. Parsers run with no network access and a strict resource budget.
- Prompt injection inside document text. This is the defining threat of RAG. Retrieved content is fenced and labelled as data in the prompt; the content guard strips instruction-shaped text at index time; the generation prompt states that retrieved content is never an instruction. Defence is layered because no single layer is reliable.
- Cross-tenant leakage. Namespace per tenant in the vector index, row-level security on metadata, filtering enforced inside the retriever.
Controls: TLS 1.3 inbound, mTLS between services, OIDC with JWKS verification at the gateway, customer-managed encryption keys, secrets rotated every 90 days, write-once audit log forwarded to the SIEM, egress restricted to an allow-list.
7. Non-functional requirements
| Attribute | Target | How it is met |
|---|---|---|
| Query latency | p95 < 6 s end to end; first token < 1.5 s | Concurrent retrieval fan-out, streaming responses, semantic cache |
| Retrieval latency | p95 < 400 ms for search + rerank | HNSW index, GPU reranker, top-50 candidate cap |
| Ingestion throughput | 10k documents/hour sustained | Queue-depth autoscaling on stateless workers |
| Ingestion latency | 95% of documents queryable within 15 min | Parallel per-modality lanes |
| Availability | 99.5% monthly for the query path | Two AZs, stateless services, managed data services |
| RPO / RTO | 15 min / 4 h | Cross-region object replication, index rebuild from derived zone |
| Groundedness | > 95% of claims traceable to a span | Grounding gate as release criterion |
| Retrieval recall@10 | > 0.85 on the golden set | Hybrid retrieval, contextual enrichment, eval gate in CI |
| Tenant isolation | Zero cross-tenant retrievals | Enforced in retriever; asserted in automated tests |
8. Cost model
Cost is dominated by two lines: generation tokens on the read path and enrichment plus embedding on the write path.
| Lever | Effect |
|---|---|
| Route classification and rewrite to Haiku | Removes the majority of calls from the Opus rate |
| Batch API for chunk enrichment | 50% reduction on the largest write-path token line |
| Prompt caching of the stable prefix | Large reduction in repeated input tokens per query |
| Semantic cache on the read path | Removes repeat questions entirely |
| Top-8 after rerank rather than top-20 | Directly reduces input tokens per generation |
Per-tenant cost is metered at the model gateway and enforced as a budget cap. Cost per query
is a CI gate (view 12), so a change that improves quality by tripling spend does not ship
unnoticed.
9. Risks
| # | Risk | Impact | Mitigation |
|---|---|---|---|
| R1 | Prompt injection via ingested content | High | Layered: index-time stripping, prompt fencing, grounding gate |
| R2 | Poor extraction on scanned or complex PDFs silently degrades answers | High | Extraction confidence scored and logged; low-confidence documents flagged, not silently indexed |
| R3 | Vector index rebuild time exceeds RTO as the corpus grows | Medium | Rebuild time measured monthly; switch to replication when it approaches 4 h |
| R4 | Generation cost grows faster than usage | Medium | Cost gate in CI, per-tenant caps, routing and cache tuning |
| R5 | Golden set ages and stops representing real queries | Medium | Curation is a standing weekly activity, fed from real failures |
| R6 | Embedding model upgrade requires full reindex | Medium | Model version pinned in the catalog; derived zone makes reindex a batch job, not a re-ingest |
10. Assumptions
- Source systems expose delta or change APIs; full re-crawls are not required per cycle.
- Document ACLs are available at ingest and are authoritative for retrieval filtering.
- A corporate OIDC provider exists and can issue tokens carrying tenant and group claims.
- Initial corpus is on the order of 10^6 chunks per tenant, not 10^9.
11. Roadmap
| Phase | Duration | Outcome |
|---|---|---|
| 0 — Foundation | 4 weeks | Landing zone, IaC, CI/CD skeleton, identity integration |
| 1 — Text and PDF vertical slice | 6 weeks | Upload, parse, index, hybrid retrieve, cited answer. Golden set v1 and the eval harness. |
| 2 — Images and tables | 5 weeks | Vision descriptions, table extraction, text-to-SQL for numeric questions |
| 3 — Audio and video | 4 weeks | Transcription, timecoded chunks, media citations that seek to the moment |
| 4 — Hardening | 5 weeks | Multi-tenancy at scale, DR exercise, penetration test, cost tuning |
| 5 — Operate | ongoing | RAGOps loop (view 14) as standing practice |
Phase 1 is deliberately a full vertical slice rather than a horizontal layer. It proves the grounding gate and the eval harness — the two things most likely to invalidate the design — before the modality-specific work multiplies the cost of changing them.
12. Traceability
| Requirement | Views | Section |
|---|---|---|
| Ingest text, PDF, images, tables, audio, video | 05 | 4.1 |
| Embedding and indexing | 06 | 4.2, 4.3, 4.5 |
| Intelligent retrieval | 07, 15 | 4.4, 4.6 |
| LLM generation layer | 07, 15 | 4.7, 5 |
| Simple upload and query interface | 01, 03, 04 | 3 |
| Grounded answers with citations | 07, 15, 16 | 4.1, 4.7 |
| Cross-source questions | 02, 05, 08 | 4.1 |