File Upload & Scanning Pipeline

Solution Architecture v1.0 · Microsoft Azure · Integration Platform Architecture · 2026-10

Everyone has used this system without being told its name. The paperclip in the email client. The drag-and-drop into a chat channel that posts a thumbnail a second later. The folder that syncs from a laptop on hotel wifi and is somehow complete in the morning. The CV uploaded to a company that has never heard of you. The 40 GB video edit dropped into a shared drive the night before a deadline. The screens are easy. What is hard is that each of them is a stranger's bytes entering a system that will later hand them to someone who trusts the system more than they trust the stranger — and the obvious design, scan it inline and store it if it is clean, fails three ways at once: it cannot survive hours of bad network, it cannot survive a scanner that is permanently slower than the uploaders, and it has no answer at all when a signature published tomorrow matches a file declared clean today. This package is the internal upload plane for an assumed collaboration SaaS: 40 million monthly active members across 12 product tenants, 60 million objects a day at a mean of 2.4 MB and a maximum of 50 GB, 150 TB/day of ingress, 12 PB under management — on two Azure Blob storage accounts split into an untrusted and a serving plane, user-delegation SAS as the only upload and download credential, Event Grid and three Service Bus lanes for scan fan-out, ephemeral Container Apps jobs in their own subscription as sandboxed scan workers, Cosmos DB for object state and verdicts, and immutable blob storage for the transition log.

21 views 21 HTML views21 SVG21 draw.io 2 documents Updated 2026-10-05
Architecture views

21 views, each in three formats.

Open a view to read it in full. Every SVG carries its diagram source inside it, so it opens in diagrams.net fully editable with no import step; the draw.io files are the same diagrams as plain source.

  1. 01
    System Context

    One upload plane for twelve product tenants, and the four consumers that are deliberately outside it.

  2. 02
    High-Level Architecture

    Five stages, and the seam between Decide and Serve that the rest of the set exists to explain.

  3. 03
    Actors and Their Core Journeys

    Nine actors, three of them machines, and the one whose experience is the design centre.

  4. 04
    Journey — Attach a File to a Conversation

    The ordinary case: four megabytes, five phases, and two seconds that must not feel like doubt.

  5. 05
    Journey — A 40 GB Master from a Bad Network

    The case that kills naive designs: hours of transfer, an interruption that is certain, and a verdict measured in minutes.

  6. 06
    Journey — Review and Release a Blocked File

    The journey that decides whether quarantine is a security control or a help-desk queue.

  7. 07
    Layered Architecture

    Eight layers, two of them storage, and one crossing rule that makes the whole set legible.

  8. 08
    Platform Components

    Three subscriptions, because the scan plane's blast radius is a subscription boundary rather than a network rule.

  9. 09
    Interface Catalogue

    Four contracts in, three out, and the one that carries bytes without touching the platform.

  10. 10
    Data Flow — One File to the First Read

    Six stages, the 35% that never reaches a scanner, and the only arrow that runs backwards.

  11. 11
    Storage Classes

    Three classes ordered by what loss costs, and only one with an RPO worth arguing about.

  12. 12
    Data Model

    Twelve entities — and the thirteenth, which every first draft adds and this design refuses.

  13. 13
    Critical Flow — Initiate to First Read

    Twenty-one messages, and the two that most designs leave out.

  14. 14
    Scan Pipeline — Three Object Classes

    Five rows, six stages, and the blank cells that are the argument.

  15. 15
    Back-Pressure — Which Promise Breaks First

    Four load bands, three lanes, and the one invariant that is never traded.

  16. 16
    Download Path

    Four checks before any byte moves, and the five minutes that bound a revocation.

  17. 17
    Deployment Architecture

    Two regions inside one residency boundary, and the arrow that is deliberately absent.

  18. 18
    Observability

    Six signal families across six stages, reduced to four alarms that wake a human.

  19. 19
    Object Lifecycle

    Seven states, and the closing arrow most implementations never draw.

  20. 20
    Security Zones

    Six zones, and the two crossings that carry the entire security argument.

  21. 21
    Identity & Credential Flow

    Nineteen messages establishing that no credential outlives the decision that justified it.

Documents

The written architecture, on the page.

The view index above is the map; this is the argument. The one-pager and the decision record are part of the deliverable, so they are printed here in full — each also opens as its own page with a table of contents.

Document 1 of 2 · 13 min read

Architecture One-Pager

File Upload & Scanning Pipeline · Solution Architecture v1.0 · Microsoft Azure · Integration Platform Architecture · 2026-10

The object's lifecycle state — not the presence of its bytes in storage — is the sole authority on access; the untrusted plane and the serving plane are separate, and nothing crosses without a current verdict.

Someone drags a file into a conversation. That gesture crosses the hardest boundary a product has: a stranger's bytes entering a system that will later hand them to someone else who trusts the system more than they trust the stranger. The naive design — scan it inline, store it if it is clean — fails three ways at once. It cannot handle the 40 GB master from a hotel network, because an upload that must survive hours of bad wifi cannot also block on a scan. It cannot handle the scanner being slower than the uploaders, which it permanently is, so the pipeline's availability becomes the antivirus vendor's availability. And it has no answer at all when a signature published tomorrow matches a file declared clean today, because a design with no quarantine state has nowhere to put an object it has changed its mind about. Twelve product teams each solving this separately produces twelve different answers to "is this file safe", twelve support queues for attachments stuck in limbo, and one estate-wide liability that nobody owns.

Accept bytes directly into an untrusted storage plane on a short-lived, path-scoped, write-only credential, so platform compute never sits in the data path at 45 GB/s. Finalise with an explicit assertion of the chunk set and the whole-object hash, establish content identity by hashing and content type by inspecting the bytes, and enqueue the object exactly once onto a durable priority queue. Scan in ephemeral, egress-restricted sandboxes under hard bounds on time, memory, archive depth and expansion ratio, composing multiple engines' results into one versioned verdict that names every engine and signature version it rests on. Record the verdict, move the object's lifecycle state, and only then promote it into a separate serving plane from which a short read credential — minted after re-checking state and authorisation — lets the store serve the bytes directly. When the signature set moves, re-judge recent objects and revoke the verdicts that no longer hold.

What it is, and what it is not

  • Lifecycle state as the sole authority on access — not Presence of bytes in a container, with a flag each read path must remember to check
  • Two storage planes with separate identity and network controls — not One container with an is_clean column
  • A verdict as a versioned, timestamped, revocable claim — not Clean as a permanent property of a file
  • Ingest availability independent of scan availability — not An upload path that fails when the antivirus vendor does
  • Direct-to-store transfer on a minted credential — not A streaming gateway sized to aggregate ingress
  • Bounds breached yields indeterminate — not A timeout treated as a pass
  • A published, lane-by-lane shedding order — not An unbounded queue and the word flake
  • Content-hash dedup inside a declared isolation level — not Platform-wide reuse that makes the index an oracle

The decisions that are the architecture

  1. State grants access, storage does not (ADR-01) — The lifecycle state machine is the only thing a read path consults. Bytes in the untrusted plane have no issuable read credential, so unreachability is structural rather than a convention every caller must honour.
  2. Two storage planes, separate accounts (ADR-02) — Untrusted and serving are separate storage accounts with distinct identity and network controls, which makes the isolation claim verifiable — and makes promotion a priced copy at p99 object size.
  3. Bytes go around the platform (ADR-03) — Clients write directly to the store on a write-only, path-scoped, 60-minute credential. No platform compute is sized to aggregate ingress, and resumability is inherited from the store.
  4. Ingest availability is independent of scan availability (ADR-04) — An object is acceptable and durably stored while the scan tier is wholly unavailable, remaining unreachable until scanned. A scanner outage is a latency incident, not an availability one.
  5. A verdict is versioned, timestamped and revocable (ADR-05) — Verdicts are keyed by content hash and engine version. Re-scan writes a new row; revocation moves the object out of the serving plane. Clean is never permanent.
  6. Bounds breached yields indeterminate, never clean (ADR-06) — Every scan is bounded in wall-clock time, memory, archive depth and expansion ratio. A breach is a typed indeterminate verdict naming the bound, which reaches a human within 24 hours.
  7. The scanner is the least-trusted compute in the platform (ADR-07) — Ephemeral per-object jobs in a separate subscription, no inbound reachability, egress allow-listed to two destinations, and no credential able to change an object's state.
  8. Back-pressure is declared, not discovered (ADR-08) — Three priority lanes with per-tenant admission limits and a published shedding order: bulk refuses first, background throttles, interactive is rejected at initiation. Rejection never happens after bytes move.
  9. Verdict reuse is bounded by a declared isolation level (ADR-09) — Content-hash dedup avoids re-scanning 35% of objects. Reuse defaults to within-tenant because platform-wide reuse turns the index into an oracle for whether another tenant holds a specific file.
  10. Finalisation is an assertion the platform can refuse (ADR-10) — The client asserts the chunk set and the whole-object hash. A mismatch is a failed upload, not an assembled object: nothing unverified ever enters the scan queue.
  11. Every object reaches a terminal state or a report (ADR-11) — A reconciler re-drives, dead-letters or reports any object past its expected dwell time. Silence about an object is a defect, because a stuck object is both a support ticket and a security hole.
  12. Residency is enforced by absence of replication (ADR-12) — Each residency boundary is an independent stack with no replication relationship to any other, so a failover cannot move bytes out of the geography a tenant declared.
  13. Size is a workload class (ADR-13) — Three execution profiles by object size rather than one sized for the worst case: a 50 KB screenshot and a 50 GB archive share an API, not a worker shape or a latency target.

Why this should still be right in ten years

Object stores, scan engines, container runtimes and credential formats will all be replaced inside a decade. What should survive is the set of claims about where authority lives, because none of them names a product.

  • "State grants access" is not a technology. It is a claim about which component is allowed to answer the question "may this person have these bytes". It survives replacing the object store, the credential mechanism and the state store, because none of those changes who is entitled to decide.
  • The seam between seeing and deciding. The worker reads bytes and cannot change state; the control plane changes state and never touches bytes. That separation outlives whatever sandbox technology currently implements it, and it is the claim a reviewer in 2036 will still want to check.
  • Revocability is permanent. Threat intelligence will always arrive after some uploads. A design where clean is a revocable claim rather than a property will remain correct however good detection becomes, because the arrival order of knowledge does not improve.
  • Bounds will matter more, not less. Formats will keep gaining nesting, compression and indirection. A design that treats a breached bound as indeterminate rather than clean degrades gracefully as content gets more hostile; one that treats a timeout as a pass degrades silently.
  • The direct-to-store decision follows economics that are not changing. Compute in a byte path costs proportionally to bytes. As objects get larger, the case for keeping platform compute out of the data path strengthens rather than weakens.
  • What will date. The engine mix, the dedup isolation default, and the re-scan window are all tied to current detection economics and current liability expectations. Expect all three to be revisited; none of them changes the boundary.

Non-functional targets

Every figure below is a stated assumption for this design, sized for the reference workload and intended to be argued with rather than believed. The view column points at the diagram where the figure is visible as a constraint.

Quality Target How it is met View
Ingest availability ≥ 99.99% monthly, measured as chunks durably stored and acknowledged Direct-to-store transfer, zone-redundant storage, control plane independent of scan plane 02
Download credential availability ≥ 99.99% monthly for already-available objects Stateless issuance against a strongly-consistent metadata read, zone-redundant 16
Scan control-plane availability ≥ 99.95% monthly Durable queues absorb orchestration outages; objects wait rather than fail 08
Initiate latency p99 ≤ 120 ms Admission decided from cached policy and a single metadata write 13
Chunk acknowledgement p99 ≤ 400 ms at 16 MB, in-region Block write direct to storage; platform not in the path 10
Time-to-verdict, ≤ 10 MB p50 ≤ 2 s, p95 ≤ 8 s, p99 ≤ 30 s on the interactive lane Fast signature engine on reserved capacity, small execution profile, no unpack stage 14
Time-to-verdict, ≤ 1 GB p95 ≤ 180 s Large execution profile, assembled scan, deep engine only on trigger 14
Time-to-verdict, ≤ 50 GB p95 ≤ 20 min Background lane, large profile, bounded expansion 14
Revocation propagation Out of the serving plane within 60 s; in-flight credentials expire within 5 min State re-checked per issuance, 5-minute read credential, edge invalidation 16
Throughput 700 initiations/s steady, 3,000/s for 10 min (4.3×), 45 GB/s aggregate ingress Control plane scales on request rate; data path scales with the store 02
Scan rate 4,000 objects/s entering the queue at peak, 35% deduplicated Three lanes, per-tenant admission, dedup short-circuit before fetch 15
Volume 60 M objects/day, 150 TB/day ingress, 12 PB under management growing 4 PB/year Tenant-partitioned metadata, tiered storage, session garbage collection 11
Retention Sessions 7 d; evidence 180 d; verdicts, provenance and transition log 7 y Lifecycle policy per container; immutable blob with legal hold for the log 11
Durability & RPO ≥ 11 nines; RPO 0 for finalised bytes, metadata and verdicts; RPO 60 s for open-session chunk state Zone-redundant in region, geo-redundant within the residency boundary 17
RTO 15 min for the read path in the secondary region; 60 min for full scan capacity Control scaled to zero in secondary, verdicts replicated so failover does not re-scan 17
Detection correctness False positives ≤ 0.01% of clean objects; indeterminate ≤ 0.2%, each resolved within 24 h Composition rule, appeal path within 4 business hours, bounded expansion 06
Safety invariant Zero tolerated cases of an object reaching a reader without a current clean verdict State-gated issuance, two storage planes, alarm on any serve attempt against a non-available object 18
Cost ≤ $0.85 per 1,000 objects ingested, scanned and stored 30 days; scan compute ≤ 35% of total No compute in the byte path, dedup avoiding ≥ 25% of scan compute, bulk lane on interruptible capacity 18

Scope

In scope

  • Upload initiation, admission and short-lived path-scoped credential minting for both end-user and server-to-server callers
  • Resumable chunked transfer from 1 byte to 50 GB, with per-chunk and whole-object integrity verification
  • Content identity by cryptographic hash, content type by byte inspection, and deduplicated verdict reuse within a declared isolation level
  • Scan orchestration across a fast signature engine and a conditional deep engine, with bounded archive expansion
  • The lifecycle state machine, quarantine with evidence retention, audited release, and verdict revocation
  • Priority lanes, per-tenant admission limits and a published shedding order
  • Authorised download credential issuance and edge delivery of available objects
  • Per-tenant policy for ceilings, allowed types, required engines, reuse isolation and evidence retention
  • Verdict and state-change events, the append-only transition log, and the operational signals that make the pipeline legible

Explicitly out of scope

  • Preview, thumbnail and transcode generation — subscribers to the available event
  • Data-loss-prevention classification and content indexing — consumers of verdicts, not gates on them
  • Digital-rights management and watermarking
  • The product surfaces that render attachments and the sharing model above them
  • Long-term archival policy and legal-hold workflow beyond evidence retention
  • Building or maintaining scan engines — they are licensed, pulled images

What a four-week prototype should prove

The prototype's job is to falsify the two claims the whole design rests on: that unreachability can be made structural rather than conventional, and that a verdict can be revoked fast enough to matter. Everything else in the design is comparatively ordinary engineering.

  1. One tenant, two destination scopes, one fast engine, both storage planes as separate accounts
  2. Resumable upload of a 20 GB object with three deliberate interruptions and one credential expiry
  3. The full lifecycle state machine with promotion, quarantine, audited release and revocation
  4. A durable queue with at-least-once delivery, idempotent verdict writes and a dead-letter path
  5. A decompression bomb and a malformed container, to prove the bounds hold and the lane does not block
  6. A revocation measured end to end: verdict revoked to last successful read
  • A read credential is requested one second after a verdict is revoked, and is refused
  • A read credential issued one second before revocation still works, and stops working within 5 minutes
  • The scan tier is stopped entirely for an hour: uploads still succeed, nothing becomes readable, and the backlog drains without loss
  • A worker is killed between reading the bytes and writing the verdict: the object is re-leased and the verdict is written once
  • A 200× decompression bomb yields indeterminate naming the bound, and the next object in the lane is unaffected
  • An upload credential is replayed against a different blob path, and is rejected by the store
  • Promotion of a 20 GB object is measured, in wall-clock time and in money, to price the two-plane decision honestly

Open risks, carried rather than hidden

Risk If it lands Response
Promotion cost at p99 object size makes the two-plane split unaffordable Either the latency between verdict and availability grows beyond the journey's tolerance, or the design collapses to in-place promotion and loses the structural isolation claim Measure promotion in the prototype at 20 GB; if the cost is prohibitive, move to in-place promotion with access-policy rewrite and record the weakened isolation claim explicitly rather than implying the strong one
Time-to-verdict at p99 object size is unachievable by assembled scanning The published 20-minute target becomes a routine breach, and the honest response is a worse product promise rather than a better pipeline Prototype streamed scanning for prefix-judgeable formats with provisional verdicts, and publish per-size-band targets instead of one number
Platform-wide dedup is adopted for cost reasons without closing the existence oracle A probe can learn that another tenant holds a specific file — a disclosure that no amount of encryption prevents Keep within-tenant reuse as the default, make platform-wide reuse an explicit knowing opt-in, and apply reuse only on the server's ingest path so it is never observable in a response
The re-scan window is quietly dropped because cold-storage reads are expensive The liability window for newly-discovered threats becomes unbounded while the product continues to imply files are checked Fund the recent-window re-scan explicitly as a line item, publish how long a file stays trusted, and make the rolling sweep's progress an operational signal rather than a background hope
Optimistic availability spreads from one scope to the default The safety invariant — nothing reaches a reader without a verdict — stops being an invariant, and the architecture's central claim is no longer true Keep the opt-in per destination scope, require the revocation path to be proven before it is enabled, and alarm on any serve attempt against a non-available object
An engine zero-day achieves code execution in the detonation zone One tenant's one object is exposed, plus whatever the egress allow-list permits — bounded, but not nothing Ephemeral per-object jobs, no inbound reachability, two-destination egress allow-list, no state-changing credential, and engine diversity so one vendor's flaw is not the whole platform's

The reasoning behind every component and technology choice is in the Architecture Decision Record: 13 records across 5 areas, each with the alternatives that lost and what the choice costs.

Document 2 of 2 · 49 min read

Architecture Decision Record

File Upload & Scanning Pipeline · Solution Architecture v1.0 · Microsoft Azure · Integration Platform Architecture · 2026-10

The argument these decisions serve is summarised in the Architecture One-Pager.

Thirteen decisions make up this architecture. Everything else across the twenty-one views is a consequence of one of them. Each record states the question it answers, the context that made the question hard, the decision, how it is realised on Microsoft Azure, the options rejected and why, what the decision costs, and what would have to change for it to be revisited.

Status of this document. This is a design, not a report on a running system. Every rate, latency, volume, retention and cost figure in this package is a stated assumption, sized for an assumed collaboration SaaS of 40 million monthly active members across 12 product tenants, 60 million objects a day and 12 PB under management. The numbers are stated explicitly so they can be argued with and corrected — not because they were measured.

How to read a record

  • Question: The forcing question: why a decision was needed at all.
  • Context: The requirement, the scale and the constraint that make it hard.
  • Decision: What this architecture does, stated so it can be checked.
  • How it is realised on AWS: The concrete mechanism: which service or package, configured how, in which subscription.
  • Options weighed: Chosen, rejected, deferred, or right elsewhere, with the reason for each.
  • Consequences: What the choice buys and what it costs, both kept visible.
  • Choose differently when: The conditions that would flip the decision for your system.
  • Why it holds up over time: What keeps the decision right as scale, staff and technology change.
  • Lesson: The principle that transfers beyond this platform.

Decision map

Authority and access: The decisions that make unreachability structural rather than a convention every read path has to remember.

  • ADR-01 · The object's lifecycle state, not the presence of its bytes, is the sole authority on access
  • ADR-02 · Untrusted and serving are separate storage accounts, and promotion is a real, priced step

The data path: Where bytes go, who touches them, and why the platform stays out of the way.

  • ADR-03 · Clients write bytes directly to the store; only control passes through the platform
  • ADR-04 · Ingest availability is independent of scan availability

Judgement: What a verdict is, what bounds it, and what happens when the world’s knowledge changes.

  • ADR-05 · A verdict is a versioned, timestamped, revocable claim — not a property of the file
  • ADR-06 · A breached bound yields indeterminate, never clean
  • ADR-07 · The scan worker is the least-trusted compute in the platform

Load and capacity: What the platform does when inspection cannot keep up with acceptance, and how size and duplication change the bill.

  • ADR-08 · Back-pressure is declared in advance, lane by lane
  • ADR-09 · Verdict reuse is bounded by a declared isolation level, defaulting to within-tenant
  • ADR-13 · Object size is a workload class, not a parameter

Integrity and completeness: How the platform knows an upload is whole, that no object was quietly lost, and that a promise about geography is true.

  • ADR-10 · Finalisation is an assertion the platform is allowed to refuse
  • ADR-11 · Every object reaches a terminal state or appears on a report
  • ADR-12 · Residency is enforced by the absence of any replication relationship

Technology by capability

Microsoft Azure was chosen for this exercise deliberately. The four most recent use cases in this practice were built on Google Cloud and Amazon Web Services, so Azure is the rotation this library exists for — and the topic genuinely suits it: user-delegation SAS gives short-lived, identity-derived, path-and-verb-scoped upload credentials without handing a client an account key, which is exactly the primitive the direct-to-store decision needs, and Event Grid with queue-triggered Container Apps jobs gives per-object ephemeral scan sandboxes without a scheduler to operate. Every requirement in the ask stays vendor-neutral; the table below is where the architecture commits.

Capability Choice Origin Credible alternative Why this one Record
Object storage, two planes Two Azure Storage accounts (untrusted, serving) with private endpoints, ZRS in region and GRS within the residency boundary Microsoft Azure S3 with bucket policies; Google Cloud Storage; MinIO on-premises Separate accounts give separate identity and network boundaries, which is what makes the isolation claim structural rather than configured ADR-02
Upload credential User-delegation SAS derived from the platform's managed identity — write-only, one blob path, 60 minutes, renewable in session Microsoft Azure S3 pre-signed PUT; GCS signed URL; STS session credentials Derived from a platform identity rather than an account key, scoped to one path and one verb, and revocable by revoking the delegation key ADR-03
Resumable transfer Block blobs, 4–256 MB blocks, up to 16 parallel, explicit commit of the block list Microsoft Azure S3 multipart upload; GCS resumable session; tus.io over a gateway Resumability and parallelism come from the store, and the explicit block list is what makes finalisation refusable ADR-10
Object metadata and state Cosmos DB, strong consistency in region, partitioned by tenant Microsoft Azure Azure SQL; DynamoDB; Spanner; PostgreSQL with read replicas The state machine is on the read critical path and must be strongly consistent; partitioning by tenant keeps one tenant's growth off another's partitions ADR-01
Verdict store Cosmos DB partitioned by content hash, immutable rows except the revocation field Microsoft Azure DynamoDB; Bigtable; an append-only log with a projection Keying by content hash and engine version makes a re-scan a new row and makes reuse an explicit lookup ADR-05
Transition log Append-only blobs with immutability policy and legal hold, 7 years Microsoft Azure S3 Object Lock; Azure Data Explorer; an append-only relational table Tamper-evident storage is what makes a release override and a break-glass read defensible years later ADR-01
Scan work intake Event Grid for finalisation, three Service Bus queues for the priority lanes Microsoft Azure SQS with three queues; Pub/Sub with three subscriptions; Kafka with three topics Durable at-least-once delivery with per-lane visibility timeouts sized to the slowest permitted scan, and dead-lettering for poison objects ADR-08
Scan execution Container Apps jobs in a separate subscription, one object per instance, no ingress, egress allow-listed Microsoft Azure Kubernetes Jobs with gVisor; AWS Fargate tasks; Cloud Run jobs; Firecracker microVMs Per-object ephemeral sandboxes with no inbound reachability and no state-changing credential bound the blast radius of an engine zero-day ADR-07
Scan engines Two licensed vendor engine images — fast signature always, deep inspection on trigger Third party Microsoft Defender for Storage as the managed on-upload scanner; ClamAV as the open-source signature engine Engine diversity means one vendor's gap is not the whole platform's; Defender for Storage was considered and does not express the quarantine state machine, bounded expansion or revocation this design requires ADR-06
Edge delivery Azure Front Door, serving-plane origin only, no pre-promotion caching Microsoft Azure CloudFront; Cloud CDN; Fastly Keeps bytes out of platform compute on the read path while never being able to reach an object that is not available ADR-01
Identity Microsoft Entra ID for workforce and customer identities; managed identities for workloads Microsoft Azure Okta; Auth0; Cognito; Keycloak The storage accounts trust only platform identities, and every client credential is derived from one ADR-03
Key management Key Vault, per residency boundary, with Managed HSM for tenants requiring customer-managed keys Microsoft Azure AWS KMS; Cloud KMS; HashiCorp Vault Per-boundary keys mean an erroneous cross-boundary copy would be unreadable at its destination ADR-12
Reconciliation Scheduled Container Apps job sweeping non-terminal objects past their class dwell time Microsoft Azure Durable Functions; Step Functions; a Kubernetes CronJob Silent loss is the default failure mode of every hop in this pipeline; the sweep is what makes it detectable ADR-11
Observability Azure Monitor and Log Analytics, alarming on oldest-unscanned age per lane, dwell-time breach, serve-attempt-on-non-available and signature-set age Microsoft Azure Prometheus with Grafana; Datadog; Cloud Monitoring The platform is scaled and paged on time-to-verdict, not throughput, so queue depth is deliberately not an alarm ADR-08

The decisions, and the alternatives that lost

Authority and access

The decisions that make unreachability structural rather than a convention every read path has to remember.

ADR-01 · The object's lifecycle state, not the presence of its bytes, is the sole authority on access

Status: Accepted · Shown on views: 02, 07, 16, 19

When a read path wants to know whether someone may have an object, what does it ask?

Context. The cheapest thing to build is a container of files and a column on a row saying whether each one passed its scan. It works until the second read path is written. From then on, correctness depends on every caller — the download API, the preview generator, the mobile sync endpoint, the admin export tool, the migration script somebody writes at 2 a.m. — remembering to check the column, interpreting a null the same way, and re-checking it rather than caching the answer. Each of those is a place where a malicious object becomes reachable, and none of them is visible in a diagram. The question is not whether the check exists but whether it can be omitted, and in a flag-based design it can always be omitted.

Decision. A single lifecycle state machine is the only authority on access. A read path never inspects storage, never interprets a verdict, and never caches an access decision: it asks the state machine, which answers from the object's current state and the authorisation of the requesting principal at that moment. The state machine is also the only component that may move an object between storage planes, and every transition it makes is recorded with actor, cause and timestamp.

How it is realised on AWS. Object state lives in Cosmos DB with strong consistency inside the region. The download API resolves state and authorisation on every credential request — there is no cached grant and no long-lived link. The untrusted storage account has no identity entitled to issue a read credential for it, so the absence of access is enforced by the storage account's own access model and not only by the API's logic. Transitions append to an immutable blob log with legal hold.

Option Verdict Reasoning
Lifecycle state machine as the sole gatekeeper Chosen One place to be right, one place to audit, and a read path that cannot accidentally skip the check.
is_clean flag on the object row, checked by each read path Rejected Correct on the day it is written; one new read path away from wrong, with no structural signal that it has gone wrong.
Access decided by storage-layer ACLs alone, with no state model Rejected Cannot express scanning, quarantined or revoked — and cannot explain afterwards why access was refused.
Signed capability tokens issued at upload, valid until expiry Rejected Fast and stateless, and unable to honour a revocation: the capability outlives the judgement that justified it.

What it buys

  • "Is this object reachable?" has one answer in one place, which a reviewer can verify without reading every caller.
  • Revocation becomes expressible: changing state is sufficient to stop new access, with no cache to invalidate in the authorisation path.
  • Adding a read path costs nothing in security review, because it has no choice but to ask.

What it costs

  • Every download request pays a strongly-consistent metadata read, which is why credential issuance has an 80 ms p99 budget rather than being free.
  • The state store is on the critical path for reads, so its availability target is the platform's download availability target.
  • A client cannot be handed a durable link, which is a real product constraint on anything that wants to embed an object URL.

Choose differently when. If every read path were owned by one team in one codebase, permanently, and the product never needed to revoke access to an object already declared clean, a flag would be adequate and cheaper. The decision should also be revisited if the state store's availability ever becomes the dominant cause of download failures, in which case the answer is a read-through cache with a lifetime shorter than the revocation target — not abandoning the gatekeeper.

Why it holds up over time. This is a statement about which component is entitled to answer a question, not about a product. It survives replacing the metadata store, the credential format and the object store, because none of those changes who decides.

Lesson. Decide early whether a safety property is enforced structurally or by convention. A convention is correct when written and decays with every new caller; a structure is checkable by someone who has not read the code.

ADR-02 · Untrusted and serving are separate storage accounts, and promotion is a real, priced step

Status: Accepted · Shown on views: 08, 11, 20

Where does an object live while nobody is allowed to read it, and what does becoming readable cost?

Context. Three arrangements are available. Keep everything in one container and distinguish by a flag or a metadata field: free promotion, instant availability, and an isolation claim that rests entirely on access control being correct everywhere. Keep two containers in one account: slightly better, still one identity boundary. Or keep two accounts with separate identity and network controls, and copy on promotion: an isolation claim a reviewer can verify by reading the access model rather than the code — and a copy whose cost at p99 object size is not small. At a mean of 2.4 MB the copy is noise. At the p99 of 240 MB it is visible. At 50 GB it is a minutes-long operation with its own failure mode sitting between the verdict and availability.

Decision. The untrusted plane and the serving plane are separate storage accounts with distinct managed identities, distinct network controls and no identity entitled to read the untrusted plane except the scan workers. An object becomes available by being copied into the serving plane under the state machine's direction, and the copy is treated as a real step with a latency budget, a failure path and a cost line — not as an implementation detail.

How it is realised on AWS. Two Azure Storage accounts behind private endpoints. The untrusted account permits write via user-delegation SAS scoped to one blob path, and read only to the scan plane's workload identity. The serving account is the only one Front Door is permitted to reach. Promotion is a server-side copy initiated by the control plane, with the object entering an explicit promoting sub-state so a failure mid-copy is visible to the reconciler rather than leaving an object that is clean and unreachable.

Option Verdict Reasoning
Two accounts, copy on promotion Chosen Isolation is structural and auditable; the price is a copy that grows with object size and a new state to reconcile.
One container, state flag only Rejected Free and instant promotion, and an isolation story that is only as good as every access-control rule in the estate.
Two containers in one account Rejected Cheaper than two accounts and shares one identity and network boundary, which is the boundary that matters.
In-place promotion by rewriting access policy rejected for now Instant and free at any size; the fallback if the prototype shows promotion cost at 20 GB is prohibitive, accepted with the weaker claim stated.

What it buys

  • The claim "an unscanned object cannot be read" is enforced by the storage account's access model, so it survives a bug in the API layer.
  • A compromised serving-plane credential yields no access to unscanned or quarantined objects at all.
  • Quarantine and evidence live naturally in the untrusted account, with break-glass as a separate identity rather than a separate code path.

What it costs

  • Promotion is priced per byte and takes time proportional to object size, which lands directly in the journey of the 40 GB upload.
  • A second account doubles the lifecycle, key and network configuration surface.
  • There is a window between verdict and availability in which an object is clean and unreachable — a state the reconciler must own.

Choose differently when. If the measured promotion cost at p99 and maximum object size exceeds what the product will pay in latency or money, switch to in-place promotion with an access-policy rewrite — and say plainly that isolation is then configured rather than structural, instead of continuing to make the stronger claim.

Why it holds up over time. The claim is that an isolation boundary a reviewer can verify is worth paying for. That outlives the specific cloud's notion of an account, though the price of the copy is a current-economics number and should be re-measured.

Lesson. A security boundary that costs nothing usually is not one. Price the boundary honestly, then decide whether the product can afford the real thing or must settle for the configured version — and name which one shipped.

The data path

Where bytes go, who touches them, and why the platform stays out of the way.

ADR-03 · Clients write bytes directly to the store; only control passes through the platform

Status: Accepted · Shown on views: 02, 09, 10, 21

Do the bytes traverse platform-owned compute, or does the platform only hand out the right to write them?

Context. A streaming gateway is attractive because it can act on bytes as they arrive: enforce the real size rather than the declared one, determine the true content type before anything lands, reject mid-stream, and hide the storage topology entirely. It is also compute that must scale with aggregate ingress — 45 GB/s at the assumed peak — and that must itself implement resumability, parallel chunking and retry semantics that the object store already has. Direct-to-store removes that component, inherits block-level resumability, and halves the byte movement, at the price of admitting bytes before anything has looked at them and of expressing all enforcement in the store's own credentialing model.

Decision. Clients write directly to the object store using a short-lived credential scoped to exactly one object path and one verb. No platform component sits in the byte path on upload or on download. Everything the gateway would have enforced is moved either earlier — to admission, before a credential is minted — or later, to finalisation and scanning, where the bytes are already durable and can be inspected without holding a connection open.

How it is realised on AWS. User-delegation SAS derived from the platform's managed identity, write-only, single blob path, 60 minutes, renewable within an open session. Clients use block-blob uploads of 4 MB to 256 MB with up to 16 parallel connections. The declared length and type are checked at initiation; the real length is checked at finalisation against the committed block list, and the real type is determined by inspecting the stored bytes.

Option Verdict Reasoning
Direct-to-store on a minted, path-scoped credential Chosen No compute scaled to bytes, resumability for free, and enforcement relocated to admission and finalisation.
Streaming gateway for all uploads Rejected Can reject mid-stream and hide the topology; costs a tier sized to 45 GB/s and a reimplementation of resumable transfer.
Hybrid: gateway for small, direct for large Rejected Defensible, and doubles the number of upload paths to secure, test and reason about for a benefit that admission already delivers.
Long-lived storage key handed to trusted product backends Rejected Simplest integration; a single leak exposes every object in the account, and nothing about it is revocable per object.

What it buys

  • The data path costs storage and network, not compute, which is where the pipeline's cost model earns most of its margin.
  • Resumability, parallelism and retry are the store's problem and are already correct.
  • A compromised platform control plane cannot read bytes it was never in the path of.

What it costs

  • Bytes land before anything has inspected them, which is precisely why the untrusted plane and the state machine must exist.
  • All credential scoping must be expressible in the store's model; anything it cannot express, the platform cannot enforce at write time.
  • A client that ignores the declared length can waste storage up to the ceiling before finalisation refuses it — garbage-collected, and not free.

Choose differently when. If a policy ever genuinely requires acting on bytes before they are durable — a regulatory prohibition on storing a class of content even transiently, for example — a gateway becomes necessary for that scope, and the decision should be revisited for that scope alone rather than for the platform.

Why it holds up over time. Compute in a byte path costs proportionally to bytes, and objects are getting larger rather than smaller. The economics behind this decision strengthen with time.

Lesson. Before putting compute in a data path, ask what it will do that cannot be done earlier on metadata or later on durable bytes. Usually the answer is nothing, and the tier exists because it felt like control.

ADR-04 · Ingest availability is independent of scan availability

Status: Accepted · Shown on views: 02, 13, 15

When the scan tier is down, does the upload fail?

Context. If scanning is synchronous with upload, the pipeline's availability is the product of every engine's availability and the signature feed's. A vendor pushes a bad engine build and uploads stop estate-wide; a licence server has a bad afternoon and twelve product teams file incidents. Worse, the scan tier is not merely occasionally unavailable — it is permanently slower than the ingest tier, because inspecting a byte costs more than storing one. A design that couples them is therefore not just fragile in incidents, it is wrong at steady state: the uploader waits on the slowest component in the system by construction.

Decision. An object is acceptable, transferable and durably stored while the scan tier is wholly unavailable. It simply remains unreachable. Ingest and scanning have separate availability targets, separate scaling signals and separate failure handling, joined only by a durable queue. A scanner outage is a time-to-verdict incident, never an upload incident.

How it is realised on AWS. Finalisation publishes to Event Grid and returns 202 with state scanning; the Service Bus lanes hold work until workers exist. Ingest targets ≥ 99.99% monthly; the scan control plane ≥ 99.95%. The scan tier scales on oldest-unscanned age per lane, so a recovering scanner drains the backlog without any coordination with the ingest path.

Option Verdict Reasoning
Asynchronous scanning behind a durable queue Chosen Uploads survive any scanner failure; the cost is that an object exists in a state nobody may read.
Synchronous scan inside the upload request Rejected No quarantine, no state machine, no revocation — and upload latency equal to scan time at every object size.
Synchronous for small objects, asynchronous above a threshold Rejected Two correctness models for one platform, and the small-object path still fails when the engine does.
Asynchronous, but reject new uploads while the scanner is down Rejected Protects time-to-verdict by sacrificing the one promise the platform can always keep.

What it buys

  • A vendor engine outage degrades time-to-verdict and nothing else; members can still upload and nothing unsafe becomes readable.
  • A 4× burst is a capacity question rather than a correctness one, because the accept path owes the caller only durability.
  • The two tiers can be scaled, priced and operated on their own signals.

What it costs

  • Objects exist that are durable and unreadable, which is a product state the UI must express honestly rather than as a spinner.
  • A backlog is now possible, which makes oldest-unscanned age the platform's most important operational number.
  • The queue becomes a correctness dependency: at-least-once delivery with idempotent verdict writes is mandatory, not a nicety.

Choose differently when. For a scope where no object may ever exist unscanned — some regulated intake flows genuinely require this — the right answer is to refuse the upload at initiation when the scanner is unavailable, not to make scanning synchronous. That is a per-scope policy, and it is still this architecture.

Why it holds up over time. Inspection will always cost more than storage. The asymmetry that motivates this decision is physical rather than technological.

Lesson. When two tiers have structurally different costs per unit of work, coupling their availability makes the cheap one as fragile as the expensive one. Put a durable queue between them and decide what the user is told in the gap.

Judgement

What a verdict is, what bounds it, and what happens when the world’s knowledge changes.

ADR-05 · A verdict is a versioned, timestamped, revocable claim — not a property of the file

Status: Accepted · Shown on views: 12, 16, 19

What exactly is recorded when a scan finishes, and what happens when tomorrow's knowledge contradicts it?

Context. The natural model is a boolean on the object: scanned, clean, done. It cannot answer the two questions that are actually asked. An auditor asks which engine and which signature version declared a specific file clean on a specific date. An incident responder asks which files were declared clean by the signature set that turned out to be missing a family, and needs them re-judged. A boolean answers neither, and worse, it cannot be corrected without destroying the evidence that the earlier judgement was made — so "we were wrong yesterday" becomes an update that erases its own history.

Decision. A verdict is a row keyed by content hash and engine-version tuple, carrying the engine identity and build, the signature-set version and its publication time, the determined content type, the outcome, any detection identifier, the bound hit if one was, the scan duration, the decision time and a revocation field. Re-scanning under a new signature set writes a new verdict rather than overwriting the old one. An object's state points at its current verdict, and revoking that verdict moves the object out of the serving plane.

How it is realised on AWS. Verdicts in Cosmos DB, partitioned by content hash, immutable once written except for the revocation field. Engine and signature versions are recorded from the worker's own runtime rather than from configuration, so provenance reflects what actually ran. Objects clean within the last 30 days are re-judged within 24 hours of a material signature change; older objects on a rolling 180-day sweep. Revocation demotes the object within 60 seconds and lets in-flight read credentials expire within 5 minutes.

Option Verdict Reasoning
Versioned, revocable verdict rows keyed by content and engine version Chosen Answers the auditor and the responder, and makes being wrong a recordable event rather than a destructive edit.
Boolean is_clean on the object Rejected Cheapest; cannot express provenance, cannot be revoked without erasing history, cannot drive a re-scan.
Verdict as a log entry only, with no current pointer Rejected Perfect history and a read path that must reduce a log to a decision on every request.
Verdict with a fixed expiry, requiring periodic re-scan of everything Rejected Honest about staleness and prices a full-corpus re-read on a timer rather than on evidence.

What it buys

  • The platform can state, for any object and any date, which engine and signature version cleared it.
  • New intelligence is actionable: the set of objects to re-judge is a query, not an archaeology project.
  • Dedup reuse becomes safe to reason about, because the thing being reused names the exact conditions under which it was decided.

What it costs

  • Verdict storage grows with re-scans as well as uploads, retained seven years.
  • The read path must resolve current-verdict-not-revoked rather than reading one field.
  • Re-scan is a real recurring cost in cold-storage reads and scan compute, and has to be funded rather than assumed.

Choose differently when. If detection ever became genuinely stable — signature sets no longer materially changing — the revocation machinery would be dead weight. Nothing suggests that, and a design that assumes it fails silently rather than loudly.

Why it holds up over time. Threat intelligence will always arrive after some uploads. A design in which clean is revocable remains correct however good detection becomes, because the arrival order of knowledge does not improve.

Lesson. If a judgement can turn out to be wrong, store it as a claim with its provenance, not as a property of the thing judged. The cost is a row; the alternative is a design that cannot change its mind.

ADR-06 · A breached bound yields indeterminate, never clean

Status: Accepted · Shown on views: 14, 19

What is the verdict when the scanner runs out of time, memory, or patience with a nested archive?

Context. Every scan has limits, and a hostile object is designed to find them. A zip nested forty deep, an archive that expands a thousandfold, a container file whose index crashes a parser, a file crafted to take an engine ninety seconds instead of ninety milliseconds. The implementation that treats a timeout as a pass is both the most common and the most dangerous: it converts an attacker's ability to make scanning expensive into an ability to bypass scanning entirely. The implementation that treats a timeout as infected is safer and unusable, because legitimate large archives hit the same bound and the false-positive rate destroys the product.

Decision. Every scan attempt is bounded in wall-clock time, memory, archive recursion depth and total expanded size. Breaching any bound produces a typed indeterminate verdict naming the bound that was hit. Indeterminate is a first-class outcome with its own state, its own handling path and its own budget: it is never readable as clean, and every indeterminate object reaches a human or an explicit policy decision within 24 hours.

How it is realised on AWS. Depth 12, 200× expansion ratio, per-class wall-clock and memory ceilings enforced by the Container Apps job's own limits, and retry at most twice in a fresh sandbox before dead-lettering. The bound hit is recorded on the verdict row. The tenant's quarantine console shows indeterminate objects distinctly from infected ones, and the release path for indeterminate differs from the one for infected — which cannot be released through the ordinary API at all.

Option Verdict Reasoning
Typed indeterminate verdict naming the bound Chosen Safe and arguable: the admin can see why, and the platform has not pretended to an opinion it does not hold.
Timeout treated as clean Rejected Turns making scanning expensive into bypassing scanning, which is the cheapest attack on the whole pipeline.
Timeout treated as infected Rejected Safe and unusable: legitimate large archives hit the same bound and the block is unarguable.
Unbounded scanning with no limits Rejected A single crafted object consumes a lane indefinitely, which is a denial of service with a progress bar.

What it buys

  • A decompression bomb is a policy outcome with a named cause rather than a crashed worker and a blocked lane.
  • The admin reviewing a block can distinguish "we found something" from "we could not finish looking", which is what makes an appeal possible.
  • The indeterminate rate becomes a measurable quality signal for both the engines and the bounds.

What it costs

  • A third outcome means a third path in the product, the console and the policy model.
  • Legitimate deep archives are inconvenienced and need a human within 24 hours — a real operational load at 0.2% of 60 million objects a day.
  • The bounds themselves are now a tuning surface with security consequences in both directions.

Choose differently when. If an engine family emerges that is provably bounded in time and memory for arbitrary input, the time and memory bounds become unnecessary — though depth and expansion bounds would remain, because those are properties of the content rather than the engine.

Why it holds up over time. Content formats keep gaining nesting, compression and indirection. A design that degrades to indeterminate rather than to clean gets safer as content gets more hostile.

Lesson. Name the failure outcome before building the success path. A system with two outcomes where reality has three will quietly map the third onto whichever of the two is cheaper, and that is almost always the unsafe one.

ADR-07 · The scan worker is the least-trusted compute in the platform

Status: Accepted · Shown on views: 08, 20, 21

What is assumed about a process whose entire job is to open files that may be designed to attack it?

Context. Scan engines are large C and C++ parsers for hundreds of formats, and parsers of hostile input are where remote code execution lives. Treating the worker as ordinary application compute — in the platform's network, with the platform's identity, in a long-lived pool with a shared filesystem — means a single engine vulnerability yields an attacker a foothold inside the system that decides what is safe. The worker is the one component in this architecture that should be designed on the assumption that it will be compromised, because its input is chosen by the adversary.

Decision. Scan workers run in their own subscription, in ephemeral jobs destroyed after one object whatever the outcome, with no inbound network reachability, an egress allow-list of exactly two destinations, no shared mutable filesystem between objects, and a workload identity that can read the untrusted plane and write verdicts and nothing else. In particular, no credential held by a worker can change an object's lifecycle state.

How it is realised on AWS. A separate Azure subscription with its own network boundary and private endpoints. Container Apps jobs, one object per job instance, no ingress. Egress restricted to the signature feed and the verdict sink. The worker's identity has read on the untrusted storage account and write on the verdict collection; it has no permission on the serving account, the object metadata store or the transition log.

Option Verdict Reasoning
Ephemeral per-object sandbox in an isolated subscription Chosen A compromise sees one object, one tenant, two egress destinations, and cannot declare anything clean.
Warm worker pool with per-object process isolation Rejected Cheaper per scan and shares a kernel and a filesystem across tenants' objects.
Scanning inside the control plane Rejected Simplest to build and places hostile-input parsing next to the component that grants access.
Third-party scanning SaaS, objects shipped out Rejected No sandbox to operate and breaks residency, since the bytes leave the tenant's declared geography.

What it buys

  • An engine zero-day is bounded to one object and one tenant, with no path to change a verdict or reach the serving plane.
  • Seeing an object and deciding its fate are held by different identities, which is the assurance claim the security views rest on.
  • Residency is preserved, because bytes never leave the boundary for scanning.

What it costs

  • Per-object job start-up is real latency and real money against a warm pool, paid 60 million times a day.
  • An extra subscription to govern, with its own identity, network and cost boundary.
  • Engine images must be pulled and cached per boundary rather than fetched on demand.

Choose differently when. If job start-up latency ever threatens the interactive lane's p50, the right move is a warm pool per tenant with per-object process and filesystem isolation — accepting a weaker claim and saying so — not a shared pool across tenants.

Why it holds up over time. Parsers of adversarial input will remain the highest-risk code in any content platform. Designing their host as expendable is not a technology choice.

Lesson. Identify the component whose input an attacker chooses, and design it as if it is already lost. The useful question is not how to keep it safe but what it is allowed to do once it is not.

Load and capacity

What the platform does when inspection cannot keep up with acceptance, and how size and duplication change the bill.

ADR-08 · Back-pressure is declared in advance, lane by lane

Status: Accepted · Shown on views: 14, 15, 18

When the scan tier cannot keep up, which promise breaks first — and who decided?

Context. The scan tier is the bottleneck by construction, so saturation is a certainty rather than an incident. Four responses are available and each is a different product. Reject new uploads: protects time-to-verdict, visibly fails the user. Widen the verdict target: accepts everything and lets attachments sit in scanning with no published bound. Admit optimistically under load: trades security posture for throughput at exactly the moment an attacker would choose. Degrade to signature-only: keeps latency, lowers detection. A system that has not chosen in advance will choose under pressure, at 3 a.m., by whichever path happens to be least resistant — and that path is almost always the unbounded queue.

Decision. Three priority lanes — interactive, background and bulk — each with per-tenant admission limits, and a published shedding order: bulk initiations are refused first, then background work is throttled and its published target widened, and only then are interactive initiations rejected with a typed retry hint. Rejection always happens at initiation, where it costs the client nothing. The one promise never traded is that nothing reaches a reader without a current verdict.

How it is realised on AWS. Three Service Bus queues with per-tenant admission counters. The scan tier scales on oldest-unscanned age per lane, and the same signal drives the shedding bands: under 1× target is normal, 1–2× elevated, 2–5× saturated, above 5× critical. A widened target is published in the API response rather than implied. Interactive capacity is reserved; bulk capacity is interruptible.

Option Verdict Reasoning
Published lane-by-lane shedding order, rejection at initiation Chosen The decision is made in daylight by the people who own the product promise, and the client is told the truth.
Unbounded queue, widen targets silently Rejected Nothing fails and nothing works; the backlog becomes the product and no number is true.
Admit optimistically under load Rejected Trades the safety invariant precisely when load may itself be the attack.
Degrade to signature-only scanning under load rejected as an automatic response Keeps latency and lowers detection without telling anyone; acceptable only as a declared per-tenant posture, never as an automatic fallback.

What it buys

  • A weekend migration cannot starve a Monday morning attachment, because the lanes have separate admission.
  • The failure mode is a typed rejection with a retry hint rather than an object stuck in scanning indefinitely.
  • Capacity planning has one number — oldest-unscanned age per lane — that is both the scaling trigger and the alarm.

What it costs

  • Three lanes multiply the admission, quota and monitoring surface by three.
  • A refused bulk initiation is visible unhappiness for a tenant mid-migration, which has to be a conversation rather than a surprise.
  • Reserved interactive capacity is paid for whether or not it is used.

Choose differently when. If scan capacity ever became elastic enough that oldest-unscanned age never left the normal band at 4× burst, the lanes would still earn their place for cost reasons — bulk on interruptible capacity — but the shedding order would become theoretical. It should still be published.

Why it holds up over time. Any pipeline whose inspection stage is more expensive than its acceptance stage will face this choice. Naming the order in advance is a governance act, not a technical one, and it does not age.

Lesson. Write down which promise you will break before you have to break one. The order is a product decision, and a system that has not been told will decide it by accident.

ADR-09 · Verdict reuse is bounded by a declared isolation level, defaulting to within-tenant

Status: Accepted · Shown on views: 10, 14

If two tenants upload byte-identical files, may the second reuse the first's verdict?

Context. 35% of objects in the assumed workload are duplicates: the same policy PDF, the same installer, the same meme, the same quarterly template. Platform-wide deduplication is therefore the single largest cost lever in the design — it removes a third of all scan compute for the price of an index lookup. It also turns the dedup index into an oracle. An attacker who can upload and observe behaviour can learn whether a specific file already exists in the platform, which across a tenant boundary means learning something about another organisation's contents: whether they hold a particular leaked document, a particular contract, a particular build. Encryption does not help; the hash is the disclosure.

Decision. Verdict reuse is a per-tenant policy with a declared isolation level: within a destination scope, within a tenant, or platform-wide. The default is within-tenant. Platform-wide reuse is available only as an explicit, knowing opt-in, and in every case reuse is applied on the platform's own ingest path and never made observable in a response — a caller cannot distinguish a reused verdict from a fresh scan by timing or by content of the reply.

How it is realised on AWS. The dedup index in Cosmos DB is partitioned by the tenant's declared isolation level, so a cross-boundary lookup is not merely forbidden but unroutable. Every reuse is recorded against the originating verdict so the audit trail shows how a decision propagated. Where reuse applies, the object still passes through the lane and still reports a time-to-verdict consistent with its class rather than returning instantly.

Option Verdict Reasoning
Declared isolation level, within-tenant default, reuse not observable Chosen Keeps most of the cost benefit inside a tenant and closes the cross-tenant oracle by construction.
Platform-wide dedup by default Rejected Maximum saving and a cross-tenant existence oracle that no later control removes.
No reuse at all; scan every object every time Rejected Simplest to reason about and spends a third of the platform's scan budget re-deriving answers it already has.
Platform-wide reuse with rate limiting on the probe path Rejected Raises the cost of the oracle without removing it, and makes the security argument a quantitative one nobody can audit.

What it buys

  • The expected ≥ 25% scan-compute saving is realised inside the trust boundary where it is free of disclosure risk.
  • The side channel is closed structurally — a cross-boundary lookup cannot be issued, not merely should not be.
  • Reuse is auditable: every reused verdict names the scan it came from.

What it costs

  • A large multi-tenant platform re-scans the same popular installer once per tenant.
  • Isolation level becomes another policy dimension a tenant has to understand, and a default that may be questioned at renewal.
  • Hiding reuse means giving up the obvious latency win of answering instantly on a hash hit.

Choose differently when. A tenant may knowingly opt into platform-wide reuse for a scope where existence disclosure is not meaningful — public marketing assets, for instance. The default should change only if the oracle can be removed entirely rather than priced.

Why it holds up over time. Content-addressed storage will always create this tension between cost and existence disclosure. The resolution — declare the boundary, default to the narrow one, keep reuse unobservable — outlives any particular index technology.

Lesson. A deduplication boundary is a confidentiality boundary. Decide it on the threat model, not on the cost model, and then take whatever saving the threat model leaves you.

ADR-13 · Object size is a workload class, not a parameter

Status: Accepted · Shown on views: 05, 14, 18

Does a 50 KB screenshot go through the same pipeline as a 50 GB archive?

Context. They share an API and nothing else. The screenshot needs no assembly, no unpack stage, a few hundred milliseconds of engine time and a worker with a small memory footprint; a human is waiting for it. The archive needs assembly to scratch storage, a bounded container walk, tens of gigabytes of working space, minutes of engine time, and nobody is watching the screen. A single execution profile sized for the archive makes every screenshot expensive and wastes most of the platform's compute; one sized for the screenshot cannot scan the archive at all; one sized in the middle fails at both ends. The same applies to latency targets: a single published number is either a lie about large objects or a betrayal of small ones.

Decision. Three execution profiles by object size — small, large, and archive or container — each with its own memory and time ceilings, its own default lane, and its own published time-to-verdict target. The dispatcher routes on the finalised object's actual size and determined type rather than on the client's claim. Targets are published per size band, never as one platform number.

How it is realised on AWS. Container Apps job definitions per profile with distinct memory and CPU allocations. Small objects stream from storage and skip assembly entirely. Large and archive objects assemble to ephemeral scratch and are walked under the depth-12 and 200× expansion bounds. Published targets: p95 8 s under 10 MB, 180 s under 1 GB, 20 min under 50 GB.

Option Verdict Reasoning
Three size-based execution profiles with per-band published targets Chosen Median cost is set by the small profile, and no published number is a lie about either end of the range.
One profile sized for the largest object Rejected Simplest to operate and spends worst-case resources on the 99% of objects that are small.
One profile with a size ceiling, large objects refused Rejected Cheap and removes a product capability the heavy-media journey depends on.
Continuous autoscaling of worker size per object Rejected Theoretically optimal and adds a scheduling problem and an unbounded failure surface for a marginal gain over three buckets.

What it buys

  • The platform's unit cost is dominated by the cheap profile, which is what makes ≤ $0.85 per 1,000 objects plausible.
  • Large objects get the resources they genuinely need without the small path paying for them.
  • Published targets per band mean the interactive journey and the 40 GB journey can both be honest.

What it costs

  • Three job definitions, three capacity pools and three sets of targets to operate and explain.
  • A boundary case near a profile threshold can land in the wrong class and miss its target, which has to be tolerated rather than tuned away.
  • Scratch storage for assembly is a real cost and a real cleanup obligation on the large and archive profiles.

Choose differently when. If the object size distribution ever lost its long tail — a product where nothing exceeds a few megabytes — a single profile would be correct and the dispatcher would be overhead. The distribution assumed here, with a mean of 2.4 MB and a maximum of 50 GB, is what makes three classes right.

Why it holds up over time. The gap between the median and the maximum object in a file platform has widened every decade. Treating size as a class rather than a parameter gets more right over time, not less.

Lesson. When one input dimension spans five orders of magnitude, it is not a parameter — it is a set of different problems wearing one API. Name the classes and publish a promise for each.

Integrity and completeness

How the platform knows an upload is whole, that no object was quietly lost, and that a promise about geography is true.

ADR-10 · Finalisation is an assertion the platform is allowed to refuse

Status: Accepted · Shown on views: 05, 10, 13

How does the platform know a resumable upload of 10,000 blocks is complete and uncorrupted?

Context. With bytes written directly to the store by the client, the platform does not observe the transfer. It learns about it only from the client's claim that the transfer is finished. A design that treats that claim as true will, eventually and quietly, assemble and scan a truncated object: a sync client with a bug, a mobile client killed by the OS mid-upload, a flaky proxy that silently dropped a block, or a deliberately malformed finalise that commits a subset of blocks chosen so the malicious portion is absent from the scan and present in the delivered file.

Decision. Finalisation is an explicit call in which the client asserts the complete block set and the whole-object hash. The platform verifies the committed block list against the asserted set and the computed whole-object hash against the asserted hash, and refuses to finalise on any mismatch rather than assembling what it has. Per-chunk checksums are verified on receipt; an object that fails either check is a failed upload and never enters the scan queue.

How it is realised on AWS. Block-blob commit with an explicit block list, per-block checksum verification at write time, and a whole-object hash computed from the committed object after assembly rather than taken on trust from the client. Mismatches return a typed error naming which assertion failed, and the session remains resumable so a client can repair rather than restart.

Option Verdict Reasoning
Explicit finalise asserting block set and whole-object hash, refusable Chosen Nothing unverified enters the scan queue, and a buggy client gets a specific error instead of a corrupt file.
Implicit finalisation on a timeout after the last block Rejected No client change needed and silently assembles whatever arrived before the timer.
Trust the client's declared length only Rejected Catches truncation but not reordering, omission or substitution of blocks.
Verify by re-reading and hashing asynchronously after finalisation Rejected Cheaper on the critical path and allows a window in which an unverified object is already queued.

What it buys

  • The scan queue contains only objects whose bytes are known to be the bytes the client meant to send.
  • Deliberate partial commits as a scan-evasion technique are closed off at the only point where they could be attempted.
  • A client bug produces a named failure rather than a corrupt file discovered by a user months later.

What it costs

  • Finalisation costs a whole-object hash computation, which at 50 GB is neither instant nor free.
  • The client must track and assert its block set, which is a real client-side obligation and a source of integration friction.
  • A mismatch strands storage until the session is repaired or garbage-collected.

Choose differently when. If the object store ever offered a verifiable whole-object digest computed at write time with the same guarantees, the platform could verify that instead of recomputing, which would remove the finalisation cost without weakening the check.

Why it holds up over time. Whenever a client writes to storage a platform does not observe, the completeness claim has to be verified rather than believed. That is a structural property of the direct-to-store decision.

Lesson. When you remove yourself from the data path, you also remove yourself from knowing what happened. Put the verification back explicitly, and make it refusable.

ADR-11 · Every object reaches a terminal state or appears on a report

Status: Accepted · Shown on views: 18, 19

What happens to the object that is in scanning and has been for nine hours?

Context. Every stage in this pipeline can lose an object without anything failing. A worker dies after reading bytes and before writing a verdict. A finalise event is published and the subscription is not yet healthy. A promotion copy fails halfway. An engine returns nothing and the lease simply expires repeatedly. In each case no error is raised anywhere, there is no failed request to retry, and the object sits in a non-terminal state looking exactly like an object that is merely busy. It is simultaneously a support ticket — the member sees scanning forever — and a security hole, because an object nobody is tracking is an object nobody is judging.

Decision. Every object has an expected dwell time for its state and size class. A reconciler sweeps for objects past that dwell time and re-drives, dead-letters or reports each one. Silence about an object is a defect by definition: an object either reaches a terminal state or appears, named, on a reconciliation report. Oldest dwell time past expectation is one of the four alarms that wake a human.

How it is realised on AWS. A scheduled Container Apps job reads object metadata for non-terminal states past their class dwell time, re-publishes the finalise event or re-enqueues the scan where that is safe and idempotent, and dead-letters with an indeterminate verdict where it is not. Queue delivery is at-least-once with idempotent verdict writes, so a re-drive cannot double-count. The reconciler's own output is a first-class operational signal, not a log line.

Option Verdict Reasoning
Reconciler as a first-class component with per-state dwell times Chosen The only design in which losing an object is detectable rather than merely unlikely.
Rely on at-least-once delivery and queue redelivery Rejected Covers the worker-death case and not the lost event, the failed promotion or the engine that returns nothing.
Client-side polling and retry Rejected Puts recovery in the hands of whichever client happens to still care, and leaves abandoned objects invisible.
A periodic report with no automated re-drive Rejected Honest and converts every lost object into manual work at a volume no team can absorb.

What it buys

  • An object stuck between the planes is found by the platform rather than by a user, which is the difference between an operational signal and a support escalation.
  • The promotion window — clean but not yet available — has an owner.
  • The safety invariant becomes checkable: the reconciler is the thing that can assert no object is unaccounted for.

What it costs

  • A full sweep over 12 PB of object metadata is itself a cost and must be incremental and partitioned to stay affordable.
  • Every re-drive path must be provably idempotent, which constrains how verdicts and promotions are written.
  • Dwell times become a tuning surface: too tight produces noise, too loose hides the problem the component exists to find.

Choose differently when. Nothing short of exactly-once delivery across every hop, including the storage copy, would remove the need for this component — and that does not exist. The sweep's frequency and partitioning are tunable; its existence is not.

Why it holds up over time. Distributed systems lose work silently. A component whose job is to notice is permanent, whatever the queue technology underneath it.

Lesson. For any state an object can be in, ask what detects the object that stays there. If the answer is a user complaint, the design has outsourced its correctness to its customers.

ADR-12 · Residency is enforced by the absence of any replication relationship

Status: Accepted · Shown on views: 17, 20

How is a promise that a tenant's files stay in their geography made true during a failover?

Context. A residency promise is easy to make and easy to break by accident. The breaking mechanisms are all mundane: a geo-redundant storage pair configured with the wrong secondary, a global load balancer failing requests to the nearest healthy region regardless of boundary, a scan tier that bursts into whichever region has spare capacity, a disaster-recovery runbook written before the residency commitment existed. Each of those is a configuration away from a boundary crossing nobody intended and nobody notices until a regulator asks.

Decision. Each residency boundary is an independent stack — its own storage accounts, its own control plane, its own scan capacity, its own keys — with no replication relationship to any other boundary. Failover happens between regions inside a boundary, never across boundaries. Scanning never transfers an object's bytes outside its boundary, including to spare capacity elsewhere. There is no configuration that could be changed to make a cross-boundary copy happen, because there is no relationship to misconfigure.

How it is realised on AWS. Two regions per boundary: primary active, secondary warm with control scaled to zero. Zone-redundant storage in region, geo-redundant within the boundary only. Verdicts replicate with the bytes so a failover does not re-scan the corpus. Front Door routes within the boundary. Keys are per boundary in Key Vault, so even an erroneous copy would be unreadable at its destination.

Option Verdict Reasoning
Independent stacks per boundary, no cross-boundary relationship Chosen The promise is kept by there being no mechanism to break it, rather than by configuration being correct.
One global stack with residency tags enforced in application logic Rejected Cheapest to operate and one bug away from a reportable breach, with no structural signal that it happened.
Global stack with policy-as-code guardrails Rejected Better, and still a boundary enforced by the correctness of a rule rather than by the absence of a path.
Cross-boundary replication for disaster recovery only Rejected Creates exactly the relationship this decision exists to remove, for an event the in-boundary secondary already covers.

What it buys

  • A failover cannot move bytes out of the declared geography, because there is nowhere out of the geography to move them to.
  • Per-boundary keys mean an erroneous copy is unreadable even if one somehow occurred.
  • The residency claim can be verified by reading the infrastructure definition rather than by auditing application logic.

What it costs

  • Every boundary pays for its own warm secondary and its own reserved scan capacity; there is no pooling of spare capacity across boundaries.
  • Three boundaries mean three of everything to deploy, patch and observe.
  • A boundary-wide outage has no cross-boundary fallback, which is a deliberate availability sacrifice.

Choose differently when. If a tenant with no residency requirement at all represented enough volume, a shared no-residency boundary could pool capacity — as an additional boundary, never by relaxing an existing one.

Why it holds up over time. Jurisdictional boundaries are getting more numerous and more specific, not fewer. A design whose boundaries are structural absorbs that; one whose boundaries are configured accumulates risk with every new rule.

Lesson. To make a promise unbreakable, remove the mechanism rather than adding the rule. A guardrail prevents the mistakes you imagined; an absent path prevents the ones you did not.

Every package used, in one table

These terms are used precisely in this package. Several are used loosely in the wider literature on upload and anti-malware pipelines, and the loose readings are what make two upload designs disagree about what they guarantee.

Package What it is What it does here Considered instead
Object One uploaded file as the platform tracks it: an identity, an owning destination scope, a content identity, a policy version, a lifecycle state and a pointer to its current verdict. The unit of state, access control and audit. A "file", which conflates the bytes with the record of what the platform believes about them.
Content identity A cryptographic hash of the bytes, independent of filename, path and owning tenant. The key for deduplication and for verdict reuse, and the join between an object and the judgements made about identical bytes. A "file hash", which usually implies a checksum used for integrity rather than an identity used for decisions.
Lifecycle state The object's position in an explicit state machine: initiated, transferring, finalised, scanning, available, quarantined, deleted. The sole authority on whether any reader may obtain the object. A "status", which is usually a display string with no authority attached.
Untrusted plane Storage holding finalised-but-unscanned, quarantined and evidence objects, with no identity entitled to issue a read credential for it. Where an object is allowed to be dangerous. A "staging bucket", which implies a workflow step rather than a trust boundary.
Serving plane Storage holding only objects whose current verdict is clean and whose state is available. The only storage any reader or edge cache can reach. "Production storage", which says where it runs rather than what is true of its contents.
Verdict A versioned claim about a content identity, naming the engine build, signature-set version, outcome, any detection identifier, any bound hit, the decision time and a revocation field. The evidence that justifies an object's state, and the thing a re-scan replaces rather than overwrites. "Scan result", which suggests a transient output rather than a durable, citable claim.
Indeterminate A terminal verdict meaning the scan could not complete within its declared bounds, naming the bound that was hit. The third outcome that stops a timeout being silently read as a pass. "Unknown" or "error", both of which invite a caller to treat the object as fine.
Promotion The state-machine-directed move of an object from the untrusted plane to the serving plane after a clean verdict. The priced step at which an object becomes reachable. "Publishing", which implies a product action rather than a storage and access-control transition.
Revocation Withdrawing a clean verdict and returning the object to quarantine when later knowledge contradicts the earlier judgement. What makes clean a claim rather than a property, and what bounds the liability window for new intelligence. "Re-quarantine", which describes the movement and omits that a recorded judgement has been withdrawn.
Lane A priority class — interactive, background or bulk — with its own admission limits, published time-to-verdict target and position in the shedding order. How a weekend migration is prevented from starving a Monday attachment. A "queue", which is the mechanism rather than the promise.
Oldest-unscanned age The age of the oldest object awaiting a verdict in a lane. The scaling trigger, the alarm, and the input to the shedding bands — chosen over queue depth because the promise being kept is time-to-verdict. "Backlog" or "queue depth", which cannot distinguish a deep queue that is draining from a shallow one that is stuck.
Dwell time The expected maximum time an object should remain in a given non-terminal state for its size class. What the reconciler compares against, making silent loss detectable. A "timeout", which implies an automatic failure rather than an investigation trigger.
Reuse isolation level The declared boundary within which a verdict may be reused for identical content: destination scope, tenant, or platform-wide. The single policy field that decides whether the dedup index can act as a cross-tenant existence oracle. "Dedup scope", which sounds like a storage optimisation rather than a confidentiality boundary.
Residency boundary An independently deployed stack for a declared geography, with no replication relationship to any other boundary. How the residency promise is kept by absence of mechanism rather than by correctness of configuration. A "region", which is a cloud construct that says nothing about jurisdiction.
The package

Everything as it was delivered.

These files are served exactly as they were produced — the diagram pages keep their own house style because that is the artifact, not a rendering of it.