# Leaderboard & Counting Service

**Solution Architecture v1.0 · Google Cloud · Platform Architecture · 2026-10 · 21 views · 17 architecture decision records**

A member finishes a lesson, a ride or a round, opens the app, and looks for their own row in a list. That number — the view count under a video, the upvote tally on a post, the weekly league standing, the segment leaderboard, the end-of-year "top 2% of listeners" card — is the most-looked-at piece of software in a consumer product and the one least often designed. This package is the shared platform that produces those numbers for every product team in one company: an assumed 80 million monthly active members across 25 product tenants, 250,000 counting events a second with a 4× burst for 120 seconds, 400,000 leaderboard reads a second split 70/30 between top-N and rank-of-member, 3.5 billion distinct live counters in 120,000 leaderboard scopes, and an event log of 25 TB per 90 days — built on Pub/Sub and a Cloud Storage archive as the record, Dataflow for event-time aggregation, Bigtable for bucket aggregates and versioned ranked views, Spanner for definitions, seasons and sealed standings, Memorystore for the serving cache, and Cloud Run for the request services.

The product is not the ability to add up numbers quickly. It is **the ability to be wrong and recover** — because a user-visible aggregate at this scale will eventually be wrong, and the only question that matters is whether that is a repair or an incident.

The design rests on one rule: **the event log is the system of record, and every counter, window and ranking is a disposable, versioned projection of it — with the platform's single piece of spent freshness going to the member's own contribution.**

The decisions that carry the design:

- **The log is the truth; everything else is a projection.** Bucket aggregates, ranked views, histograms and caches are all derived, versioned and droppable. A corrupt partition, a bad ranking deploy, a wrong tie-break rule and a fraud campaign discovered six weeks late therefore have one remedy — rebuild — instead of four disaster procedures, and rolling back a ranking is a pointer move (ADR-01).
- **A write is acknowledged on durability, never on being counted.** The accept path's only synchronous dependency is the log, so write latency is stable through pipeline restarts, projection rebuilds and cache loss, and a 4× burst is a capacity question rather than a correctness one (ADR-02).
- **Rebuild is routine, not a disaster procedure.** Every tenant's largest board is rebuilt monthly at ten times real time and diffed against live, which makes the 30-minute RTO a measured number and makes a diff the earliest possible warning that an accumulator has stopped being associative (ADR-03).
- **Exactly-once effect comes from the data model.** A client-supplied idempotency key deduplicated at admission and again in the pipeline, because the most common duplicate in a consumer product is the one the broker never sees — a client that retried before the first request was acknowledged (ADR-04).
- **Event time, bounded, with the overflow made visible.** A ride that syncs four hours late counts in the week it happened, up to a declared horizon; past it, a visible late bucket rather than a silent drop or a silent fold-in. Bounded lateness is what lets a window be declared complete, which is the precondition for sealing a standing at all (ADR-05).
- **Accuracy is declared per counter and carries a price.** Exact, approximate or unique-cardinality is part of the versioned definition, enforced by the platform, with the unit cost per million events shown where the choice is made — which is what makes a distinct-viewer counter affordable and stops an estimate quietly ending up behind a number someone is paid against (ADR-06).
- **The hottest key pays for itself.** Per-key write rate is detected at admission and state is split across shards with read fan-in, adapting as the key heats and cools. The alternatives make either the median key or the 5-second RPO pay for the viral one (ADR-07).
- **A rank outside the materialised head is a percentile.** An exact rank at position 41,208 of 10 million is a count-less-than query nobody will pay for at 120,000 reads a second — and "top 3%" is usually the better answer anyway. Scoped boards exist to give an exact rank that means something (ADR-08).
- **Freshness is spent where it is perceived.** The acting member's own accepted-but-unprojected deltas are applied at read time over a stale projection. One extra keyed read on 30% of requests buys the only consistency a member can actually detect (ADR-10).
- **Stale but honest, and stable within a session.** Every response carries its projection version and as-of, and a session is pinned to one version so a rank cannot oscillate between two reads with no action taken (ADR-11).
- **Removal is a compensating event, and closure is published.** Counters are never mutated; a deletion, a ban, a disqualification or a fraud takedown is a compensating delta with cause, requester and authority. After an explicit closure point a retraction produces a correction record beside the original standing rather than a silent rewrite (ADR-12, ADR-13).
- **Suspect contributions are held, not judged.** An inflation signal is probabilistic, so the contribution is quarantined — outside the counters, retained with original timestamps, reversible — and a freeze stops publication without stopping ingestion. A false positive then costs a delay rather than a member's effort (ADR-14).
- **Counters and the control plane never share a store.** 3.5 billion rebuildable rows and a few million exact ones have different recovery stories, and a store that is excellent at both at a 1000:1 ratio does not exist (ADR-15).
- **One write region, with a read standby.** Active-active counting is a merge-semantics problem, not a replication one: commutative counters merge, and the ranked views, seasons and closures built on them do not (ADR-16).

The architecture one-pager (including why the design should still hold up in ten years, and the five risks that would change it) and the full decision record appear on the landing page of the diagram set, directly below the index of views. The same content is published as [docs/architecture-one-pager.md](docs/architecture-one-pager.md) (~15 min) and [docs/decision-record.md](docs/decision-record.md) (~65 min).

---

## What is here

| Path | Contents |
|---|---|
| `diagrams/index.html` | The landing page: 21 views in seven acts with every format linked, then the **architecture one-pager** and the **decision record** |
| `diagrams/*.html` | One self-contained page per view: the inlined diagram plus the reasoning cards, with copy / PNG / PDF export |
| `diagrams/svg/*.svg` | The same views as SVG with the diagram XML embedded; they re-open fully editable in diagrams.net |
| `diagrams/drawio/*.drawio` | draw.io source |
| `docs/architecture-one-pager.md` | The one-pager as markdown |
| `docs/decision-record.md` | The 17 decision records, the capability-to-technology table and the package glossary as markdown |
| `specs/part-a..c.json` | Diagram specifications, the source of truth for every view |
| `specs/manifest-a..b.json` | Acts, page titles, subtitles and reasoning cards |
| `specs/adr-onepager.json`, `specs/adr-records-a..b.json` | The one-pager, the decision records, the capability-to-technology table and the glossary |
| `scripts/build.sh` | Rebuilds every deliverable from the specs (Node only, no network) |
| `ask.md` | The requirement |

## The twenty-one views

| # | View | What it answers |
|---|---|---|
| 01 | System Context | Who generates the numbers, who consumes them, and the one decision the platform refuses to make — what earns a point |
| 02 | High-Level Architecture | Six stages, with the caller's acknowledgement at the end of the second one and replay drawn as a routine path |
| 03 | Actors and Their Core Journeys | Nine actors, three of them machines, including the analyst who removes a cheater and the service that pays against a standing |
| 04 | Journey — Where Do I Stand | The trough is not latency: it is a member's own contribution missing from a number that is otherwise correct |
| 05 | Journey — Ship a New Leaderboard | The trough is a cost-and-accuracy choice the engineer cannot make blind, which is why the price is shown at declaration |
| 06 | Journey — Remove a Cheater | Contain, prove, retract, reopen — and what the platform owes when the season has already closed and a reward was paid |
| 07 | Layered Architecture | Six layers, with Processing below Service because nothing in the request path waits on it |
| 08 | Components and Boundaries | Two planes and one rule: the request plane reads the processing plane's output and never waits on it |
| 09 | Integration Surface | Four surfaces — count, read, configure, operate — and why the client write path is the narrowest of them |
| 10 | Counting Data Flow | One stream, two destinations: the aggregation that feeds the screen and the archive that makes it disposable |
| 11 | Storage Zones by Ownership | Three zones with three recovery stories — backed up, rebuilt, simply lost — and nothing allowed to sit between them |
| 12 | Data Model | Twelve entities, and the two that are versioned because changing them in place would reinterpret history |
| 13 | From One Event to a Rank | Seventeen messages, exactly one of which the caller waits on, and a red one for the stall that widens staleness |
| 14 | Aggregation Pipeline | Six stages, every accumulator commutative and associative, which is what makes a rebuild identical to the original run |
| 15 | Read Paths by Query Type | Five paths, one authorisation model, and the one difference that matters: only the member's own rank gets freshness spent on it |
| 16 | Deployment Architecture | Two regions, one write region, and a standby whose cache is rebuilt on cutover rather than replicated |
| 17 | Observability Matrix | Six signal classes across five stages, with projection lag singled out as the only one a member can perceive |
| 18 | Projection Lifecycle | Declared, materialised, published, advanced, rebuilt — and released when nobody has read it for thirty days |
| 19 | Trust Zones | Five zones; a client may increase its own counter and nothing else, and nothing outside the restricted zone holds a key |
| 20 | Who May Increase a Counter | Four independent refusals before anything is logged, and only the fourth — the inflation signal — is reversible |
| 21 | Assumed Failure Classes | Seventeen named failures in three tiers, sorted by who notices: nobody, the reader as a widened as-of, or an analyst |

## Rebuilding

```bash
bash scripts/build.sh
```

Node 20+ and nothing else. No draw.io Desktop, no browser, no network. The build assembles
`specs/views.json` and `specs/manifest.json` from the authoring parts, generates the draw.io
sources, fails on any geometry or routing defect, renders the editable and plain SVGs, writes the
HTML pages and the index, injects the one-pager and the decision record, and proves every relative
link in `diagrams/` resolves.

At v1.0 the geometry gate reports 0 errors and 0 warnings across 21 files, the routing gate
0 errors and 11 clutter warnings across 168 edges, all 348 nodes resolve to an icon with no weak
matches, and the link check passes on 166 relative links. The 11 warnings are label stacking on
the context and integration views, where every relationship edge converges on one centre node;
they were reduced by shortening labels and cutting edges that other views already carry, and the
remainder is inherent to those two layouts.

## A note on the numbers

Every rate, latency, ratio, threshold and retention figure in this package is a **stated
assumption**, chosen to be defensible and arguable rather than measured. They are stated precisely
so that a reviewer can disagree with one and follow it to the decision that depends on it.
`ask.md` marks them as assumptions section by section; the decision record's evidence note
restates the operating context in one paragraph.
