Document 11 min read

Architecture One-Pager

Solution Architecture v1.0 · Microsoft Azure · Cost & Efficiency Architecture · 2026-09

Cost Allocation & Showback Platform · Solution Architecture v1.0 · Microsoft Azure · Cost & Efficiency Architecture · 2026-09

Every published number is a frozen artefact produced by a pure, versioned function of three immutable inputs — never a live query over mutable data.

The provider bills an account, a subscription or a project. The organisation wants to know what a team, a product and a customer cost. Between the two sits an allocation problem that is mostly metadata and politics and only partly arithmetic, and it is made hard by three things at once: cloud billing data is restated for weeks after the usage occurred; a large share of spend is incurred by shared infrastructure that no single team owns; and the organisation whose costs are being attributed is itself moving, as teams split, resources change hands and the apportionment rules are argued over and revised. A platform that answers by querying the current state of all three will give a different answer to the same question on two consecutive days and be unable to say why. That is the single failure that ends a showback programme, because a number nobody can reproduce is a number nobody will accept.

Land every provider export immutably and exactly as given. Normalise beside it, never over it, onto one FOCUS-conformed cost record, and reconcile the total to the provider's own invoice before anything downstream may publish. Keep ownership as an effective-dated model resolved by a declared rule chain, so a snapshot can be taken as of any instant. Keep allocation policy as a versioned, simulatable artefact outside the engine. Then make allocation a pure function of exactly three pinned inputs — billing snapshot, ownership snapshot, policy version — whose output is derived, disposable and rebuildable. Freeze that output into a statement that carries all three versions, supersede it rather than overwrite it when the inputs move, and let nothing downstream of the freeze read anything that can still change.

What it is, and what it is not

  • A pure function of three pinned, immutable inputs — not A dashboard that queries live billing and current ownership
  • A raw zone that is the complete, immutable source of everything — not A staging area that is cleared once the warehouse is loaded
  • An unallocated remainder published as its own line — not A proportional spread that balances to 100% and hides the gap
  • Restatement as the normal path run a second time — not Restatement as an exception workflow with its own code
  • A frozen statement version that supersedes its predecessor — not A statement table updated in place at month end
  • A read-only measurement platform — not A cost-optimisation tool that resizes or deletes resources

The decisions that are the architecture

  1. Allocation is a pure function of three pinned inputs (ADR-01) — Billing snapshot, ownership snapshot and policy version are resolved and hashed before any arithmetic happens, and every allocated row and every statement carries all three. 'Why did my number change?' is answerable in one sentence: which of the three moved.
  2. The bill is landed as given, forever (ADR-02) — Every provider export is written immutably to the raw zone, versioned by provider, period and export version. Normalisation happens beside it. Everything downstream is derived from it, which is what makes any period rebuildable years later.
  3. Allocated facts are materialised, not resolved at query time (ADR-03) — A statement must be an artefact that exists, can be frozen, and can be pointed at during a dispute. Query-time resolution makes policy iteration free and makes publication meaningless.
  4. Restatement supersedes; it never overwrites (ADR-04) — A re-issued export lands as a new version, diffs at record grain, re-runs the same pure function, and produces statement v2. v1 survives for seven years, and so does the reason for the delta.
  5. Everything about the organisation is effective-dated (ADR-05) — Ownership resolves by a declared precedence chain — tag, parent scope, mapping table, exception — as of a point in time, and records which rule fired. A reorganisation changes the future, not the meaning of a closed period.
  6. The remainder is a named line, not a spread (ADR-07) — Unallocated spend is published as its own line in every view at every level, and charged to a named accountable owner rather than distributed proportionally. Balancing to 100% is exactly why tag coverage never improves.
  7. Cluster idle lands on whoever sized the cluster (ADR-09) — Shared compute is apportioned on max(request, usage), and the reserved-but-unused remainder goes to the platform owner — the only party who can act on it — rather than being averaged across tenants who cannot.
  8. Reconciliation is a gate, not a metric (ADR-11) — A period whose ingested total does not match the provider invoice within 0.1% or $500 does not publish. Showback may be late; it may not be quietly short.
  9. Frozen and current are never the same figure (ADR-12) — The current period is served from the same pipeline, explicitly labelled as an estimate, and never rendered in the same number as a closed statement. One path, two labels.
  10. Access is scoped by accountability, server-side (ADR-16) — The scope predicate is resolved from the effective-dated org tree and injected into every query; a parameter cannot widen it. Out-of-scope reads are answered at aggregate grain rather than refused.

Why this should still be right in ten years

Cloud service catalogues, billing formats and organisational shapes will all change inside the life of this platform. These are the properties that should outlast them.

  • Purity outlives the engine. The contract is that identical inputs produce identical output. Whether the engine is Spark, a warehouse or something that does not exist yet, the contract is unchanged and every past period stays reproducible.
  • Immutable raw survives every format change. FOCUS will be revised; providers will add charge categories and retire SKUs. Because the conformed layer is derived and the raw layer is not, a new conformer re-derives history rather than losing it.
  • Effective-dating survives every reorganisation. The one thing guaranteed to happen repeatedly is that the organisation changes shape. A model that stores ownership as of a time rather than as of now does not need to be rebuilt when it does.
  • Policy as data survives the argument. Apportionment rules are the part of this system people will keep changing, because they encode a negotiation rather than a fact. Keeping them versioned and outside the engine means the argument costs a review, not a release.
  • A derived, disposable fact store survives growth. The largest store is the one that can be thrown away. As volume grows forty per cent a year, the cost decision about retention stays reversible, because nothing irreplaceable lives there.
  • Showback and chargeback share one ledger. The move from informational to billing is a governance change, not a re-platforming, because enabling chargeback changes who acts on the number rather than how it is computed.

Non-functional targets

Every figure below is a stated assumption for this design. They are chosen to be defensible and arguable; a reviewer who changes one can follow it to the decision that depends on it.

Quality Target How it is met View
Reporting availability ≥ 99.5% monthly; publication and posting ≥ 99.9% around close Zone-redundant serving tier separated from batch, so no interactive request waits on an allocation run 17
Data freshness Conformed by 09:00 T+1 p95; allocated by 10:00 T+1 p95, ceiling T+2 Hourly export polling with a declared cutoff, then a daily incremental allocation run 13
Allocation run time Full month-to-date re-run p95 ≤ 45 min, hard ceiling 90 min Partition-parallel Spark on interruptible capacity, checkpointed per partition 12
Query latency Dashboard p95 ≤ 2 s over 13 months; drill-through p95 ≤ 10 s Pre-aggregated roll-up marts for dashboards, full-grain fact store reserved for drill-through 09
Attribution completeness ≥ 98% of spend to a named owner at close; remainder ≤ 2% Rule-chain resolution plus an owner-facing reconciliation surface ranked by cost 16
Reconciliation accuracy Within 0.1% or $500 of the provider invoice, whichever is larger Per-provider invoice comparison as a publication gate, not a dashboard metric 12
Reproducibility Identical inputs produce bit-identical output Run digest over the three pinned input ids; a matching digest returns the existing result 12
Durability RPO 0 for statements, disputes, ownership and policy; RPO 24 h for conformed facts Synchronous commit and geo-replication for human-authored and published data only 10
Recovery RTO 4 h reporting, 1 h on close days, 24 h for full fact re-derivation Cold serving in the recovery region; the fact store is rebuilt from the replicated raw zone 17
Scale 1.2 bn cost records and 2.6 bn usage samples per month, growing 40% a year Columnar fact store partitioned by period and provider; telemetry downsampled on ingest 07
Detection Sustained ≥ 30% day-over-day increase routed within 24 h; ≤ 5 false positives per team per quarter Detect on allocated facts so the alert already has an owner; suppression windows for declared events 16
Platform cost ≤ 0.5% of the spend it measures Lifecycle-tiered raw storage, spot batch capacity, pre-aggregated serving; published as a line in its own report 17

Scope

In scope

  • Ingestion and FOCUS conformance of billing data from three clouds, plus licence and SaaS spend
  • Effective-dated ownership resolution and tag-coverage measurement in currency
  • Shared-cost apportionment, commitment amortisation and remainder handling under versioned policy
  • Frozen monthly statements with drill-through, restatement and a dispute workflow
  • Budgets, forecasts and anomaly detection routed to the resolved accountable owner
  • Unit economics from team-reported business drivers
  • A GL posting path for chargeback, generated only from frozen statements

Explicitly out of scope

  • Procurement, contract negotiation and the decision to buy a commitment
  • The general ledger itself — the ERP remains the book of record
  • Rate optimisation actions: the platform reports opportunities and never changes a resource
  • Capitalisation, depreciation and headcount cost
  • Tag enforcement at resource creation, which belongs to the cloud governance platform

What a four-week prototype should prove

The prototype's job is to falsify the central claim — that allocation can be a pure, reproducible function over data that is still moving — on one cloud and one month of real spend.

  1. Land one provider's detailed export immutably, then land a restated version of the same period and show the record-grain diff
  2. Resolve ownership by the rule chain against real tags, and publish the coverage gap in currency rather than resource count
  3. Run allocation twice on identical pinned inputs and show bit-identical output and a matching digest
  4. Re-run the same period after a single policy change and produce a per-team delta report
  5. Freeze a statement, then restate it, and retrieve both versions with the delta explained to resource level
  6. Apportion one shared Kubernetes cluster from real usage telemetry and show where the idle went
  • The export for one provider does not arrive: the period publishes as visibly incomplete, naming the provider
  • The export schema gains an unrecognised column: the partition is quarantined and the provider's period is blocked
  • A team is split mid-period: both successors receive the correct share and the closed prior period is unchanged
  • Usage telemetry is lost for six hours: the affected allocations are marked degraded and excluded from chargeback
  • A bad apportionment rule is adopted: it is rolled back in one command and the affected periods re-run inside the batch window

Open risks, carried rather than hidden

Risk If it lands Response
Tag coverage never improves, and the remainder stays above the 2% target The number stays arguable, chargeback never becomes credible, and the platform is treated as a report Charge the remainder to a named accountable owner rather than spreading it; publish coverage per team in currency; make the reconciliation surface rank gaps by cost so fixing metadata is cheaper than disputing the bill
Apportionment becomes a permanent political negotiation rather than a rule Policy churns every period, statements restate constantly, and the ledger loses the stability that made it trusted Policy is versioned, simulated against a closed period before adoption, published to tenants before it applies, and reviewed on the close cadence rather than on demand
Provider restatements arrive after the correction window and after the ledger has been reported Finance holds a number the platform can no longer reproduce, and the two records diverge permanently A 90-day correction window with supersede-not-replace versioning; restatements past the window are carried as a labelled adjustment in the current period rather than silently applied
The fact store's growth outruns the 0.5% platform-cost target The platform that measures efficiency becomes a line item someone else has to justify The fact store is derived and disposable: retention can be shortened without data loss because the raw zone is the archive of record; roll-ups absorb dashboard traffic
Chargeback is enabled before the ledger is trusted Every dispute becomes a budget dispute, the workflow saturates, and teams route around the platform with their own spreadsheets Showback earns chargeback: the same ledger, gated on sustained attribution completeness and dispute volume rather than on a date in a programme plan

The reasoning behind every component and technology choice is in the Architecture Decision Record: 16 records across 6 areas, each with the alternatives that lost and what the choice costs.