Distributed Job Scheduler

Solution Architecture v1.0 · Google Cloud · Platform Architecture · 2026-10 · 20 views · 16 architecture decision records

A developer opens the settings screen, types 0 9 1-5, picks a time zone, and expects a digest to go out every weekday morning. The box is four words wide. Behind it is a durable multi-tenant timer fleet that has to fire tens of millions of independent triggers at the right instant while a partition owner is being replaced, a node's clock is wrong, a zone is draining, and a tenant has just resumed four thousand triggers that slept through a day. This package is that service: the scheduled-trigger platform inside one developer product — an assumed 50,000 tenants, 20 million active triggers, 500 million fires a day at a mean of 5,800 a second, with 60% of all fires landing in the first second of a minute and a design peak of 45,000 a second for five seconds at the top of the hour — built on a globally consistent relational store for the registry, the due index and the fire ledger with commit timestamps as time authority, a rate-shaped task queue for lane-split dispatch, serverless containers for the control, timing and dispatch tiers, a wide-column store for fire history, and a columnar warehouse for lateness and cost reporting.

20 views 20 HTML views20 SVG20 draw.io 2 documents Updated 2026-10-07
Architecture views

20 views, each in three formats.

Open a view to read it in full. Every SVG carries its diagram source inside it, so it opens in diagrams.net fully editable with no import step; the draw.io files are the same diagrams as plain source.

  1. 01
    System Context

    What is inside the boundary, who touches it, and what the service refuses to own.

  2. 02
    High-Level Architecture

    The fire path as a spine: declare, remember, decide, commit, shape, deliver, account.

  3. 03
    Actors and Their Core Journeys

    Six parties, their goals in their own voice, and what each of them gets to do.

  4. 04
    Journey — Ship a Scheduled Job

    A developer with a four-word requirement meets the one screen where it gets hard.

  5. 05
    Journey — The Morning After an Outage

    The on-call engineer who finds out from a customer, and what the platform owes them.

  6. 06
    Layered Architecture

    Seven layers, and the single seam between the control plane and the timing plane.

  7. 07
    Platform Components

    The containers inside one region, and the only component that leaves the perimeter.

  8. 08
    Integration Surface

    Three inbound APIs, deliberately separated, and four outbound target classes.

  9. 09
    Data Flow — Definition to Evidence

    Seven states of one instant, of which only two are authoritative.

  10. 10
    Storage Zones

    Backed up, rebuildable, or cold — and why that is the only classification that matters here.

  11. 11
    Data Model

    Eleven entities, and one four-part primary key that does the work of a deduplication service.

  12. 12
    Critical Flow — One Instant, Decided and Delivered

    Eighteen messages, and the one that is the architecture: the ledger insert.

  13. 13
    Catch-up After an Outage

    The expected failure mode, drawn as a pipeline rather than described as an incident.

  14. 14
    Dispatch Lanes

    Five classes of fire through one pipeline, with a published order for who loses first.

  15. 15
    Deployment Architecture

    One write region across three zones, and a standby that deliberately holds no leases.

  16. 16
    Observability

    Signal type by fire-path stage, and the one row that exists because the others cannot see this failure.

  17. 17
    Trigger Lifecycle

    Seven states, and why the loop only closes if the tenant can see lateness.

  18. 18
    Security Trust Zones

    Five zones, and the two guards that carry most of the risk.

  19. 19
    Identity and Privilege Flow

    Who proves what, to whom, in what order — including the call that is deliberately refused.

  20. 20
    Failure Classes and Their Answers

    Ten classes, each with what it looks like, what answers it structurally, and what is left over.

Documents

The written architecture, on the page.

The view index above is the map; this is the argument. The one-pager and the decision record are part of the deliverable, so they are printed here in full — each also opens as its own page with a table of contents.

Document 1 of 2 · 18 min read

Architecture One-Pager

Distributed Job Scheduler · Solution Architecture v1.0 · Google Cloud · Platform Architecture · 2026-10 · 20 views · 16 architecture decision records

The timing plane decides; the execution plane works; and the only thing crossing between them is an immutable, uniquely-keyed fire record that is committed before it is delivered.

Every developer platform ships a four-word scheduling box, and behind it is a problem that looks solved and is not. A multi-tenant timer fleet has to fire tens of millions of independent triggers at the right instant while a partition owner is being replaced, a node's clock is wrong, a zone is draining, and a tenant has just resumed four thousand triggers that slept through a day. The convenient implementation — fire from a node's system clock and record the result afterwards — fails in three ways that each look like success at the time. It fires early when a clock runs fast, producing work that reads a window of data which has not closed yet and reports a short answer. It loses fires that happened but were never recorded, which no amount of later reconciliation can recover because the record that was supposed to prove it is the thing that is missing. And when it comes back from an outage it releases everything it owes at once, turning a scheduling incident into an incident for every executor downstream. On top of those, it cannot answer the one question tenants actually ask — "did my job run last night?" — because a job that did not run produces no signal at all, in a service that is up, healthy on every dashboard, and silent.

Separate the plane that decides from the plane that works, and make the scheduler's durable output a decision rather than a result. A partition owner asks a time source for a bounded uncertainty interval, scans its share of a time-ordered due index for instants due at or before the lower bound of that interval — so it is deliberately ε late and never early — applies the trigger's declared overlap policy, and commits an immutable fire record whose primary key is derived deterministically from the tenant, the trigger and the scheduled instant. That key is the idempotency key, so two owners mid-handover compute the same one and the second insert loses to a uniqueness constraint instead of producing a second fire. Only then is anything dispatched, through one of three separate queues — on-time, catch-up, backfill — each with its own drain rate, so a backlog released by a recovery is rate-shaped rather than flooded and recovery is deliberately slower than the failure that caused it. Delivery is at-least-once with the key attached and a published residual duplicate rate, not an exactly-once claim nobody can verify. Everything derived — the due index, the caches, the fire history — is a rebuildable projection of a registry and a ledger that are the only systems of record, and the rebuild runs continuously in a shadow partition rather than waiting to be needed. What the tenant gets is a per-fire record with its scheduled, decided and dispatched instants and the cause of any non-dispatch; what the platform watches is the age of the oldest undispatched due instant, which is the only signal that can tell a working scheduler from a silent one.

What it is, and what it is not

  • A durable decision, committed before delivery — not A timer that fires and then writes down what happened
  • Deliberately ε late, never early, at any percentile — not As punctual as the node's clock allows, in both directions
  • At-least-once with a key and a published duplicate rate — not Exactly-once delivery
  • Idempotency as a uniqueness constraint in the data model — not A deduplication service on the dispatch path
  • Imperfect leases made survivable by the fire key — not Leader election strong enough to be trusted alone
  • Post-outage behaviour declared per trigger beforehand — not An operator decision taken during the recovery
  • Recovery rate-shaped slower than the failure — not Drain the backlog as fast as the pipeline allows
  • A scheduler that hands over work — not A workflow or DAG orchestrator
  • Fire history as a tenant-facing product surface — not A platform log that support can search

The decisions that are the architecture

  1. Commit before dispatch, always (ADR-01) — The fire record is written before any attempt is made, so a dispatcher that dies mid-attempt loses an attempt and never a fire. Of the two orderings available, one failure is unrecoverable and the other is a retry; the ordering of two writes decides which one the system can have.
  2. The control plane cannot take the fire path down (ADR-02) — Two planes, separately deployed and scaled, sharing only the state of record. Console traffic, a bad management deploy and an identity outage become latency incidents instead of missed fires. One write region, because two regions evaluating "is the previous run still going" produce two answers no key reconciles.
  3. Idempotency lives in the data model (ADR-03) — The fire's primary key is (tenant, trigger, scheduled instant, sequence) — computed, not generated. Split ownership becomes a constraint violation rather than a detection problem, which is what lets leases be short, takeover fast, and leader election imperfect.
  4. Late is a budget; early is a bug (ADR-04) — The node clock is untrusted. The scheduler takes a bounded uncertainty interval, waits it out, and scans for instants due at or before its lower bound — paying about 10 ms of deliberate lateness to make earliness impossible. A node whose uncertainty exceeds the ceiling sheds its leases rather than firing on a clock it cannot justify.
  5. Recurrences are recomputed, not materialised (ADR-05) — One due-index row per trigger, holding one next instant, recomputed on advance against the named zone. An amendment, a pause, a validity change and a time-zone database update are all picked up with no invalidation sweep, because nothing was cached to invalidate — and dry-run is exact rather than indicative.
  6. The post-outage decision is declared before the outage (ADR-06) — Every trigger declares its missed-fire policy from a closed set of three, defaulting to firing once for the most recent missed instant. A billing run and a cache warm have opposite correct answers and look identical to the platform; the only party who knows which is which is the tenant, and the only time they can say so is in advance.
  7. Wall-clock time means what the tenant meant (ADR-07) — Expression plus IANA zone, resolved at computation, with declared rules for the day a time does not exist and the day it happens twice, and the time-zone database version recorded on every fire — so a disputed instant six months later is explainable from data rather than from memory.
  8. Every policy is bounded by a horizon (ADR-08) — An instant older than the catch-up horizon is recorded as expired with its cause and never dispatched, whatever the tenant declared — the one place the platform overrules them. Caught-up fires carry their original scheduled instant, because the work reads a window defined by it.
  9. Recovery is slower than failure, by design (ADR-09) — Three queues with independent drain rates and a per-tenant catch-up cap, plus a published shedding order: backfill, catch-up, over-quota, then on-time. A one-hour backlog takes about ten hours to drain, which is the price of not making the recovery into the next outage for every executor downstream.
  10. The peak is absorbed, not purchased (ADR-10) — Sixty per cent of fires land in the first second of a minute. Deterministic opt-out jitter smearing and the lateness budget absorb a 45,000/s peak that lasts five seconds an hour, instead of provisioning eight times the steady-state capacity to be idle for the rest of it.
  11. Rebuildable means rebuilt on a schedule (ADR-11) — The due index, the leases and every cache are projections of the registry and are not backed up. A shadow rebuilder recomputes and compares continuously, and partition ownership moves routinely rather than only during failure — because a recovery path that is only exercised in an incident is a claim, not a capability.
  12. unknown is a state, not an absence (ADR-12) — Overlap is enforced against recorded outcome state bounded by a window. A fire still unknown at the end of its window counts as finished, and the decision is recorded. The alternative — treating it as still running — turns one lost callback into a permanently stopped trigger that looks perfectly healthy.
  13. Alert on the gap, not on health (ADR-13) — The paging signal is oldest undispatched due instant age, because a scheduler that is up and not firing passes every other check. The same measurement, per fire and per trigger, is a tenant-facing product surface — so "did my job run" is answerable without reading platform logs or opening a ticket.
  14. A dispatch target is a privilege, not a parameter (ADR-14) — Targets must be verifiably owned, checked at definition and again at dispatch; no tenant request reaches the fire zone; and credentials are minted per attempt with the attempt's lifetime. A service that will make authenticated requests to any URL on a timer, with a retry budget, is an attack tool until ownership is proved.
  15. The operations that multiply work are separate grants (ADR-15) — Backfill, horizon extension and quota change are grantable apart from trigger authorship, and audited with prior values. Each of them is most attractive at the moment judgement is worst, which is during the incident — so the routine role's worst case stays routine.
  16. Opaque data gets the strictest boundary (ADR-16) — The tenant payload exists in two stores, encrypted under a per-tenant key, and never in history, reporting, metrics or logs. The platform cannot classify what it does not interpret, so it handles the payload as the most sensitive thing it could be and debugging starts from a fire key.

Why this should still be right in ten years

Clouds, queue services, container runtimes and consistency models will all be replaced inside this platform's life. These are the properties that should outlast them.

  • The boundary is not a technology. "The timing plane's product is a committed decision" is a statement about where authority lives, not about which store holds it. Spanner, a successor nobody has shipped, or a different strongly consistent store can each satisfy it, and every decision downstream — the key as the deduplication mechanism, lanes as retries against records, history as a projection — survives the substitution unchanged.
  • Deterministic identity outlives every protocol. Naming a fire from its own semantics rather than generating an identifier for it is the oldest reliable trick in distributed systems. It depends on no store, no broker and no transport, and it is why imperfect leader election is affordable here. What could change is executors' willingness to deduplicate — which would call for an added suppression store, not a redesign.
  • The asymmetry between late and early is permanent. Scheduled work reads windows defined by its scheduled instant. That is a property of the work, not of the scheduler, and no improvement in clocks makes an early fire safe. If clocks get better, ε shrinks and the deliberate lateness disappears on its own; if they get worse, the lease-shedding ceiling catches it.
  • Recovery will always need to be slower than failure. A system that accumulates obligations while it is down and then discharges them into dependencies it does not control has a feedback loop into its own failure. That is control theory, not engineering fashion. The cap's value should move with measured executor capacity; the existence of a cap should not.
  • Silence will always need a gap measurement. Any system whose failure produces no output is invisible to health checks, forever. Alerting on the distance between what should have happened and what did is the only class of signal that sees it, and that remains true whatever the monitoring stack looks like.
  • The declared-policy habit transfers. Collecting, in advance, the information a system will need to make a decision it cannot make correctly during an incident is a general property of operable systems. Here it is the missed-fire policy; elsewhere it is a failover preference or a data-loss tolerance. The mechanism is configuration; the principle is foresight.

Non-functional targets

Every figure below is a stated assumption. They are listed with the mechanism that is supposed to deliver them and the view where that mechanism is drawn, so a reviewer can disagree with the number and the means separately.

Quality Target How it is met View
Fire punctuality p50 ≤ 1 s, p99 ≤ 5 s, p99.9 ≤ 30 s of lateness Due scan at 1 s tick on leased partitions, with the dispatch lane's drain rate sized for the mean rather than the peak 12
Earliness Zero at every percentile Scan for instants due at or before the lower bound of the clock's uncertainty interval; ε is waited out, not averaged 12
Degraded punctuality p99 ≤ 60 s for up to 10 minutes during zone loss or reassignment Partitions spread across three zones; reassignment exercised routinely rather than only on failure 15
Dispatch plane availability ≥ 99.95% monthly, measured as due instants dispatched inside their p99.9 budget Punctuality as the availability definition; stateless dispatch tier scaled independently of the timing tier 06
Control plane availability ≥ 99.9% monthly, failing independently of the fire path Separate deployment, scaling and service accounts, sharing only the state of record 06
Trigger population 20,000,000 active triggers across 50,000 tenants 256-partition keyspace with one due-index row per trigger; partition count changeable online 07
Steady-state throughput 500,000,000 fires/day — 5,800/s mean One transactional fire insert per fire; lane queues with declared drain rates 09
Peak throughput 45,000 fires/s for 5 s at the top of the hour Deterministic opt-out jitter smearing plus the lateness budget, instead of provisioned peak capacity 14
Clock uncertainty ε ≤ 10 ms; leases shed above 100 ms Bounded-uncertainty time source; per-node ε recorded at lease grant and alerted per node, not averaged 12
Duplicate dispatch ≤ 1 in 10^6 fires delivered more than once Deterministic fire key as the ledger primary key; executors deduplicate on it 11
Durability RPO 0 in-region; cross-region RPO ≤ 15 s Synchronous three-zone replication for registry and ledger; asynchronous standby replica 15
Recovery time RTO 5 min dispatch plane, 30 min control plane; partition reassignment ≤ 15 s p99 Declared region promotion; short leases made safe by the fire key; routine reassignment 15
Catch-up bound Horizon 1 h default / 24 h max; catch-up ≤ 10% of tenant quota Horizon filter ahead of the policy; separate catch-up queue with its own drain rate 13
Configuration propagation Create, amend, pause, resume, delete effective within 5 s p99 Recomputed due-index row written in the same transaction as the version 09
History 90 days queryable, 13 months archived; 30-day single-trigger query p95 ≤ 500 ms Bigtable projection keyed tenant#trigger#instant with TTL; archive-class object storage beyond it 10
Detection Oldest undispatched due instant > 60 s pages immediately Derived per-partition metric from the due index, alerted independently of process health 16
Cost ≤ $0.60 per million fires at steady state Mean-sized capacity, no per-trigger always-on resource, tiered history, retry and catch-up charged to the originating tenant 16

Scope

In scope

  • Trigger definition and lifecycle — calendar, interval and one-shot recurrences in named zones, with versioning, pause, resume, delete, manual run and exact dry-run
  • Time semantics: a bounded-uncertainty clock, declared daylight-saving rules, and a recorded time-zone database version per fire
  • A partitioned, time-ordered due index with lease-held ownership and clock-gated instant claiming
  • An append-only fire ledger whose primary key is the idempotency key, committed before any dispatch
  • At-least-once dispatch to an authenticated HTTP target, a tenant queue or an internal executor, with bounded jittered retry truncated at the next instant
  • Missed-fire policies, a platform-enforced catch-up horizon, and a bounded authorised backfill
  • Overlap policies with unknown as a terminal state, and per-trigger and per-tenant concurrency limits
  • Per-tenant quotas, fair dispatch queues, deterministic jitter smearing and a published shedding order
  • Tenant-facing per-fire history with causes, and the oldest-undispatched-due-age alert
  • Target ownership verification, per-attempt dispatch credentials, signed dispatch and verified callbacks, and the privilege split for backfill, horizon and quota

Explicitly out of scope

  • Executing the triggered work — the scheduler hands over a record and a credential and nothing more
  • Workflow and DAG orchestration, retries of business logic, and compensation; the separate distributed-workflow-orchestration-platform package covers that ground
  • Trigger-to-trigger dependencies, deferred to Phase 3 and only if they can be added without becoming an orchestrator
  • Sub-minute recurrences, which need a separately budgeted punctuality contract
  • Tenant-supplied recurrence predicates such as business calendars and market holidays, deferred to Phase 3 behind a sandboxed evaluation budget
  • Active-active multi-region scheduling, deferred because two authorities evaluating overlap produce answers no key reconciles
  • The product surfaces that render a schedule, and the tenant's own alerting on its work's outcome

What a four-week prototype should prove

The prototype's job is to falsify the two central claims — that imperfect leases plus a deterministic key really do make duplicate fires harmless, and that a bounded clock lets the scheduler be reliably late and never early — on one tenant and a few hundred triggers, not to build a platform.

  1. One tenant, 500 triggers across all three recurrence forms, in three time zones including one with daylight saving, with the two awkward days exercised deliberately rather than waited for.
  2. The timing plane on leased partitions against a strongly consistent store, with the fire key as the ledger's primary key and the clock gate reading a bounded uncertainty interval.
  3. Two dispatch lanes — on-time and catch-up — with independent drain rates and a per-tenant cap, against one deliberately slow HTTP target and one that times out without answering.
  4. The shadow rebuilder, running from day one, so the rebuild claim is under measurement for the whole four weeks rather than tested at the end.
  5. Fire history with scheduled, decided and dispatched instants and the cause of every non-dispatch, queryable by the tenant — because the prototype has to answer "did it run" to be worth anything.
  6. Measured cost per million fires on the prototype's shape, compared against the assumed $0.60, which is the assumption with the least evidence behind it.
  • A partition owner is killed mid-cycle, repeatedly, under load: every instant has exactly one fire record and at most one extra dispatch attempt, with the measured duplicate rate compared to the assumed 1 in 10^6
  • A lease handover happens while a scan is in flight: the second insert loses to the uniqueness constraint, and no application code resolves the conflict
  • A node's clock is stepped forward by two seconds: no fire goes out early, at any percentile
  • A node's reported ε is pushed past 100 ms: it sheds its leases and the partition is reassigned inside the 15-second budget, rather than firing on the bad clock
  • A 1-minute trigger is paused for three hours and resumed under each of the three missed-fire policies: the dispatch rate stays inside the 10% catch-up cap, and the horizon expires what it should with a cause the tenant can read
  • A dispatch is held open past its timeout with no callback: the fire is recorded unknown at the end of its window, the overlap decision is recorded, and the next fire is not withheld indefinitely
  • The due index is dropped entirely and rebuilt from the registry: the result matches what the shadow rebuilder had been reporting all along
  • A trigger is pointed at a domain the tenant has not verified: the definition is refused at write time, not at the first fire

Open risks, carried rather than hidden

Risk If it lands Response
An executor that cannot deduplicate on the fire key At-least-once stops being a contract and becomes a defect: a duplicate fire produces duplicate work, and the whole leases-may-be-imperfect argument collapses into a correctness problem Make deduplication an onboarding requirement with a conformance test rather than a documentation note; publish the duplicate rate on the target's own page; and keep a platform-side suppression store as a costed Phase 2 option for targets that genuinely cannot.
Fleet-wide clock drift, where ε is wrong everywhere at once rather than on one node Lease-shedding has nothing to compare against, the clock gate approves a bad interval, and the design's central guarantee — no early fires — fails silently and uniformly Alert on per-node ε rather than an aggregate, cross-check the time service against the store's commit timestamps as an independent source, and treat a correlated ε excursion as a platform incident with scheduling suspended rather than degraded.
The missed-fire default is wrong for most tenants, and they accepted it without reading it Either duplicated billable work or silently skipped work, discovered by a customer, with the platform's correct behaviour as the explanation — the trough in view 05 Instrument override rates by tenant segment and treat a high override rate as evidence the default is wrong; show the policy's consequence in a sentence at the point of choice; and make dry-run replay what each policy would have done to the last recorded gap.
The oldest-undispatched-due-age alert is loud during every legitimate catch-up drain The one signal that detects the system's characteristic failure becomes the one the on-call rotation learns to ignore, which is worse than not having it Compute the metric separately per lane so a catch-up drain does not mask an on-time stall, suppress it against a declared drain with an expected completion time, and page on the on-time lane alone.
Recomputing instants for 20 million triggers every cycle becomes the scaling wall Scan cost grows with the trigger population rather than with the fires actually due, and punctuality degrades as the population grows rather than as load does Measure compiler cost per trigger per cycle from day one as a first-class capacity metric; keep the short-horizon materialisation variant of ADR-05 costed and ready, since it is a localised change to the due index's contract and leaves the ledger, the key and the lanes untouched.

The reasoning behind every component and technology choice is in the Architecture Decision Record: 16 records across 5 areas, each with the alternatives that lost and what the choice costs.

Document 2 of 2 · 63 min read

Distributed Job Scheduler — Architecture One-Pager and Decision Record

Distributed Job Scheduler · Solution Architecture v1.0 · Google Cloud · Platform Architecture · 2026-10 · 20 views · 16 architecture decision records

The argument these decisions serve is summarised in the Architecture One-Pager.

The architecture behind one small box in a developer platform's settings screen: "every weekday at 09:00". The one-pager is the argument in fifteen minutes; the decision record is the same argument with its working shown, sixteen decisions deep.

Status of this document. This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure is a stated assumption, chosen to be defensible and arguable rather than measured. The operating context assumed throughout is a developer platform with 50,000 tenants and 20 million active triggers, firing 500 million times a day at a mean of 5,800 fires a second, with 60% of all fires landing in the first second of a minute and the top of the hour carrying eight times the mean minute — a design peak of 45,000 fires a second for five seconds. Punctuality is assumed at p50 ≤ 1 s, p99 ≤ 5 s and p99.9 ≤ 30 s of lateness, with zero tolerance for earliness at any percentile; the clock uncertainty bound is assumed at ε ≤ 10 ms with leases shed above 100 ms; the residual duplicate dispatch rate is assumed at 1 in 10^6. Where a number came from nowhere, it is marked as an assumption in the view cards and in the records, and a reviewer should treat each of them as an invitation to disagree.

How to read a record

  • Question: The forcing question: why a decision was needed at all.
  • Context: The requirement, the scale and the constraint that make it hard.
  • Decision: What this architecture does, stated so it can be checked.
  • How it is realised on AWS: The concrete mechanism: which service or package, configured how, in which subscription.
  • Options weighed: Chosen, rejected, deferred, or right elsewhere, with the reason for each.
  • Consequences: What the choice buys and what it costs, both kept visible.
  • Choose differently when: The conditions that would flip the decision for your system.
  • Why it holds up over time: What keeps the decision right as scale, staff and technology change.
  • Lesson: The principle that transfers beyond this platform.

Decision map

Decision and delivery: The boundary that defines the architecture, and the two decisions that follow from it directly.

  • ADR-01 · The scheduler's durable output is a committed fire record, written before any dispatch attempt
  • ADR-02 · The control plane and the timing plane share only the state of record, and there is one write region
  • ADR-03 · Idempotency lives in the data model: the fire's primary key is the idempotency key, and delivery is at-least-once

Time: What time it is, how wrong that might be, and what a wall-clock expression actually means.

  • ADR-04 · The node clock is untrusted; dispatch waits out a bounded uncertainty interval, so lateness is a budget and earliness is a bug
  • ADR-05 · Recurrences are recomputed each cycle from a versioned definition, not materialised ahead into the due index
  • ADR-07 · Wall-clock recurrences resolve against a named zone with declared DST rules, and the time-zone database version is recorded on every fire

Policy after the outage: What a trigger does about the fires it missed — declared before it misses them, not decided during the incident.

  • ADR-06 · The missed-fire policy is declared per trigger before the outage, and defaults to firing once for the most recent missed instant
  • ADR-08 · A catch-up horizon bounds every missed-fire policy, and the original scheduled instant travels with the fire
  • ADR-09 · Catch-up, backfill and on-time fires use separate queues with a per-tenant catch-up cap and a published shedding order
  • ADR-12 · Overlap is enforced against recorded outcome state, with unknown as a first-class terminal state bounded by a window

Scale, cost and recovery: Absorbing the top-of-hour peak, and treating the due index as something rebuilt rather than restored.

  • ADR-10 · The top-of-hour peak is absorbed by opt-out jitter smearing and the lateness budget, not by provisioned capacity
  • ADR-11 · The due index is a rebuildable projection, verified continuously by a shadow rebuild rather than backed up

Evidence and trust: What the tenant can see, what a schedule is allowed to point at, and who is allowed to multiply work.

  • ADR-13 · Fire history is a tenant-facing product surface, and the platform alert is oldest undispatched due age
  • ADR-14 · A target must be verifiably owned by its tenant, no tenant request reaches the fire zone, and dispatch credentials are minted per attempt
  • ADR-15 · Backfill, horizon extension and quota change are grants separate from trigger authorship
  • ADR-16 · Tenant payloads live only in the registry and the ledger, encrypted under a per-tenant key, and never in evidence or logs

Technology by capability

Google Cloud was chosen for this exercise for two reasons that point the same way. The first is rotation: across this repository's use cases Microsoft Azure and self-hosted open source dominate, with Amazon Web Services next, and Google Cloud carries the smallest share — six of forty-four packages before this one. Reaching for the same cloud each time teaches a service catalogue rather than architecture. The second reason is that this particular topic genuinely belongs here. A scheduler's hardest requirement is a defensible answer to "what time is it, and how wrong could that be", and this is the platform that exposes a bounded clock-uncertainty interval on a transactional store rather than asking the design to assume one. ADR-04 is the decision that depends on it, and it would read as hand-waving on a stack that cannot report its own clock error. Everything else below is replaceable; the requirement document deliberately stays vendor-neutral so that it survives a change of cloud.

Capability Choice Origin Credible alternative Why this one Record
Trigger registry, due index, fire ledger Spanner regional instance, synchronous across three zones, commit timestamps as time authority Google Cloud CockroachDB or YugabyteDB self-hosted; DynamoDB with a separate time service; Azure Cosmos DB with strong consistency One store that enforces the fire key's uniqueness constraint and reports bounded clock uncertainty — the two properties ADR-01, ADR-03 and ADR-04 all depend on ADR-01
Bounded-uncertainty time Platform time service plus Spanner commit timestamps, with per-node ε recorded at lease grant Google Cloud AWS Time Sync with clock bound; a self-run PTP fabric; chrony with a measured error estimate The decision to wait out ε rather than fire at the earliest permitted instant needs a source that reports its own error, not one that is merely accurate ADR-04
Control, timing and dispatch planes Cloud Run services, separately deployed and scaled, with distinct service accounts Google Cloud GKE workloads; ECS or Fargate; Azure Container Apps Independent scaling and deployment is the mechanism of the plane split, and leases make long-lived instances unnecessary ADR-02
Dispatch lanes Three Cloud Tasks queues per region with independent rate and concurrency Google Cloud SQS with separate queues; RabbitMQ with per-queue limits; Azure Service Bus Per-queue dispatch-rate control is exactly the throttle the catch-up cap needs, and it is enforced by the queue rather than by application code ADR-09
Outcome events Pub/Sub for correlated outcome events from the platform executor Google Cloud Kafka; SNS and SQS; Event Hubs Decouples the executor's reporting from the dispatch path, so a slow reporter is not a slow dispatcher ADR-12
Fire history Bigtable projection keyed tenant#trigger#instant with a 90-day TTL Google Cloud DynamoDB with TTL; Cassandra; Azure Table Storage 500 million rows a day of append-mostly evidence with a prefix-scan read pattern and native expiry ADR-13
History archive Cloud Storage, archive class, 13 months, dual-region Google Cloud S3 Glacier Instant Retrieval; Azure Blob archive tier Retention without query cost; the archive is for obligations, not for answering questions ADR-13
Lateness and cost reporting BigQuery for lateness facts, shadow divergence and cost per million fires Google Cloud Snowflake; Redshift; ClickHouse self-hosted Keeps punctuality analysis entirely off the fire path, including during the incident it is being used to diagnose ADR-11
Dispatch identity Workload identity tokens minted per attempt, scoped to the trigger, verified by the target's issuer check Google Cloud SPIFFE/SPIRE with short-lived SVIDs; AWS IAM Roles Anywhere; Azure managed identities A credential whose lifetime equals an attempt reduces a standing grant to a seconds-long one ADR-14
Payload encryption Cloud KMS per-tenant keys, decrypted only in the dispatch path Google Cloud HashiCorp Vault transit; AWS KMS; Azure Key Vault with an HSM-backed key Per-tenant keys make deletion enforceable by key destruction rather than by trusting every copy to be found ADR-16
Perimeter and egress VPC Service Controls around the data stores; egress only from the attempt runner Google Cloud AWS PrivateLink and SCPs; Azure Private Endpoints with a firewall One component with egress is an egress policy small enough to audit ADR-14
Observability Cloud Monitoring for due-age and lateness, Cloud Logging with fire keys and no payloads Google Cloud Prometheus and Grafana; Datadog; Azure Monitor The alert is a derived measurement (oldest undispatched due age), so it needs a metric pipeline rather than a log search ADR-13
Audit Append-only audit table with a Cloud Storage Object Lock export, 7 years Google Cloud S3 Object Lock; Azure immutable blob storage; an on-prem WORM appliance The acts worth auditing are the ones that multiply work, and their record has to survive the actor who performed them ADR-15

The decisions, and the alternatives that lost

Decision and delivery

The boundary that defines the architecture, and the two decisions that follow from it directly.

ADR-01 · The scheduler's durable output is a committed fire record, written before any dispatch attempt

Status: Accepted · Shown on views: 02, 07, 12

When a dispatcher dies holding a fire it has decided is due, what has been lost?

Context. The convenient implementation of a scheduler dispatches first and records afterwards: the timer fires, the HTTP call goes out, and a row is written when the call returns. It is one less write on the hot path and it reads naturally. It also has no answer to the question above. A dispatcher that dies after the call and before the write leaves a fire that happened and is not recorded — which means the next scheduler to look at that trigger will fire it again, and nothing in the system can tell the difference between that and the first fire. The inverse ordering has the opposite failure: a fire recorded and not dispatched, which is a retry. One of those is unrecoverable and the other is routine, and the ordering of two writes decides which one the system has.

Decision. An immutable fire record — tenant, trigger, definition version, scheduled instant, deterministic idempotency key, attempt counter — is committed to a strongly consistent store before any dispatch attempt is made. The record, not the delivery, is the scheduler's durable product. Delivery is a retryable consequence of a record that already exists.

How it is realised on AWS. The fire ledger is a Spanner table whose primary key is (tenant_id, trigger_id, scheduled_instant, sequence). A partition owner that has decided an instant is due performs the insert inside the same transaction that advances the due index row, then enqueues onto a Cloud Tasks lane. The attempt runner reads the record, mints a per-attempt credential, signs the payload and dispatches; every attempt writes an attempt row against the same fire key. Nothing in the dispatch plane can create a fire.

Option Verdict Reasoning
Commit the fire record, then dispatch Chosen Makes the unrecoverable failure impossible and the recoverable one routine; costs one transactional write per fire on the critical path
Dispatch, then record the outcome Rejected Cheaper and simpler; a crash between the two produces a fire that happened and cannot be proved, which is the one state the design cannot recover from
Write an intent record best-effort, reconcile later Rejected Reconciliation needs a source of truth to reconcile against, and the intent record was supposed to be it
Rely on the queue's own durability as the record Rejected A queue entry is not addressable by scheduled instant, cannot be queried by a tenant asking "did it run", and its retention is not the history retention

What it buys

  • A dispatcher can die at any point in an attempt and the worst outcome is a duplicate attempt under a key the executor already knows
  • "Did it run" becomes a primary-key lookup rather than a log search, which is what makes fire history a product surface instead of a support tool
  • Retry, catch-up, backfill and manual run are all the same operation — attempts against a record — rather than four code paths

What it costs

  • One strongly consistent write per fire sits on the critical path, which at a 45,000/s design peak is the single largest capacity commitment in the system
  • A ledger outage stops fires entirely rather than degrading them, which is deliberate (ADR-02) and is the hardest consequence to explain to a tenant
  • The ledger grows at 500 million rows a day and needs a retention and partitioning story from day one rather than later

Choose differently when. If the work being triggered were idempotent by nature and cheap to repeat — a cache warm, a metrics scrape, a health probe — the unrecoverable failure would not matter and dispatch-then-record would be the right trade. The decision is justified by schedules that send money, email and statements, where a fire that cannot be proved is worse than a fire that is late.

Why it holds up over time. "The decision is the product" is a statement about where authority lives, not about which store holds it. Spanner, a successor nobody has shipped, or a different consistent store can each satisfy it, and everything downstream — the key as the deduplication mechanism, lanes as retries against records, history as a projection — survives the substitution unchanged. What would not survive is a move to an eventually consistent store for the ledger, because the uniqueness constraint is the whole mechanism.

Lesson. In any system that acts on the world on a timer, decide which of your two unavoidable failures is the recoverable one, and then order your writes so you only ever get that one.

ADR-02 · The control plane and the timing plane share only the state of record, and there is one write region

Status: Accepted · Shown on views: 06, 07, 15

If the API that creates schedules is down, should schedules still fire?

Context. A scheduler has two obviously different jobs. One is interactive, bursty, request-shaped and has a human waiting: create, amend, pause, dry-run, read history. The other is autonomous, continuous, deadline-driven and has nobody waiting: decide which instants are due and hand them over. Building them as one service is the default, and it means a deploy, a bad query, a traffic spike on the console or an authorisation dependency outage takes the fire path with it — the plane with a deadline is held hostage by the plane with a user. The same question repeats across regions: a second region that can also decide instants doubles write availability and introduces two authorities deciding the same overlap question, which the fire key deduplicates only after both have already decided.

Decision. The control plane and the timing plane are separately deployed, separately scaled and separately available, sharing nothing but the strongly consistent state of record. A control-plane outage stops new and amended definitions; it does not stop scheduled fires. Scheduling authority lives in exactly one write region, with a standby region that replicates state and holds no partition leases.

How it is realised on AWS. Both planes run as independent Cloud Run services with separate service accounts, deployment pipelines and scaling policies, against one Spanner instance holding registry, due index and ledger. Partition leases are rows in that instance, so only a process with write access to the primary region can hold one. The standby region runs the same images scaled to zero with a read replica; promotion is a declared operation that moves the lease table's authority, not an automatic failover.

Option Verdict Reasoning
Separate planes, one write region, cold standby Chosen Fire path survives control-plane failure; region loss is a declared promotion against a stated RPO
One service for both planes Rejected Simplest to operate and deploy; makes every console incident a punctuality incident
Active-active scheduling across two regions Rejected Best availability; two authorities decide overlap and concurrency independently, and the fire key cannot undo a decision already taken
Separate planes with the control plane in a second region Rejected Attractive for console availability; a definition written in region B that region A has not yet seen is a fire that silently did not happen

What it buys

  • Console traffic, a bad management deploy and an identity-provider outage are all latency incidents rather than missed fires
  • The timing plane's dependency list is short enough to reason about: the state of record and the time authority
  • Overlap and concurrency decisions have exactly one authority, which is what makes the skip and queue policies in ADR-12 meaningful

What it costs

  • Two deployments, two scaling policies and two on-call surfaces for one product
  • Region loss costs up to the replication lag in decisions and a declared promotion rather than a transparent failover
  • Control-plane availability is lower than the dispatch plane's, and tenants will experience that as the service being down

Choose differently when. If the scheduler served one tenant with a hundred triggers, the operational cost of two planes would dominate the availability it buys and one service would be right. At 50,000 tenants, the console is busy enough that treating its incidents as fire-path incidents is indefensible.

Why it holds up over time. The plane split is an availability-domain statement and outlives any runtime. Active-active scheduling stays rejected for a structural reason rather than a technical one: the fire key makes a duplicate fire harmless, but two regions independently evaluating "is the previous execution still running" produce two different answers, and no key reconciles those. That reasoning does not change if replication gets faster.

Lesson. Split a system where its deadlines differ, not where its nouns differ. A plane with a user and a plane with a deadline should never be able to take each other down.

ADR-03 · Idempotency lives in the data model: the fire's primary key is the idempotency key, and delivery is at-least-once

Status: Accepted · Shown on views: 11, 12, 20

Two schedulers both believe the same instant is due. What stops two fires?

Context. Every distributed scheduler eventually has two processes that believe they own the same trigger: a lease handover, a network partition, a paused process resuming, a clock disagreement. The usual responses are to make that impossible — a stronger lease, a consensus round, a global leader — or to detect it afterwards with a deduplication store keyed by some generated identifier. The first is expensive and still probabilistic, because a lease is a promise about time held by a process that may not know the time. The second adds a component on the critical path whose own availability becomes the fire path's availability, and whose retention window becomes the deduplication guarantee. Both are machinery built to compensate for an identifier that was chosen badly.

Decision. The fire's identity is derived deterministically from (tenant, trigger, scheduled instant, sequence within instant) and is the primary key of the fire ledger. Two processes that both believe an instant is due compute the same key, so the second insert fails on a uniqueness constraint. Delivery to the executor is at-least-once, the key travels with every dispatch, and the residual duplicate rate is published as a contract the executor must tolerate rather than concealed behind an exactly-once claim.

How it is realised on AWS. The ledger's Spanner primary key is the four-part tuple; the insert is a plain INSERT whose ALREADY_EXISTS is handled as a normal, expected return meaning "another owner decided this instant". The same key is sent in the dispatch payload and in a header, so an HTTP target, a queue consumer and the internal executor all deduplicate on the same value. Attempts are child rows under the fire key, so a retry can never create a second fire.

Option Verdict Reasoning
Deterministic key as the ledger primary key Chosen Turns split ownership into a constraint violation; needs no extra component and no retention window
Generated UUID per fire with a deduplication store Rejected Works for transport retries; cannot deduplicate two independent decisions, because they generate different identifiers
Consensus round before every fire Rejected Strongest agreement; adds a round trip to a 45,000/s peak and still leaves the clock question open
Claim exactly-once delivery via transactional outbox to the executor Rejected Honest only where the platform owns the executor, which is one of four target forms

What it buys

  • Leader election is allowed to be imperfect, which means leases can be short and takeover fast (15 s) without risking correctness
  • No deduplication component sits on the fire path, so there is one less thing whose outage is an outage
  • The duplicate rate is a published number a tenant can design against rather than a property they discover

What it costs

  • Every executor must deduplicate. A tenant target that cannot is a correctness gap the platform cannot close for them
  • The sequence component of the key has to be decided by whoever decides an instant produces more than one fire, which is a subtlety in the catch-up path
  • A monotonic key means hot-spotting on the ledger's key range at the top of the hour, mitigated by the tenant prefix leading the key

Choose differently when. If the platform owned every executor, a transactional handoff would make exactly-once honest and the published duplicate rate unnecessary. With tenant HTTP targets in the mix it is not achievable, and claiming it would be the more dangerous choice.

Why it holds up over time. Deterministic identity from the semantics of the event is the oldest reliable trick in distributed systems and does not depend on a store, a broker or a protocol. What could change is the willingness of executors to deduplicate; if that becomes unrealistic, the decision does not need replacing so much as supplementing with a platform-side suppression store — an addition, not a rewrite.

Lesson. Before building a deduplication service, ask whether the thing being deduplicated could have been given a name that makes duplicates impossible to express.

Time

What time it is, how wrong that might be, and what a wall-clock expression actually means.

ADR-04 · The node clock is untrusted; dispatch waits out a bounded uncertainty interval, so lateness is a budget and earliness is a bug

Status: Accepted · Shown on views: 06, 12, 20

A node believes it is 09:00:00. How wrong could it be, and which direction of wrongness is acceptable?

Context. A scheduler is a program whose entire output depends on the one value that distributed systems are worst at agreeing on. NTP-synchronised clocks in a datacentre are usually within a few milliseconds and occasionally wrong by seconds — a lost sync, a step correction, a virtualised clock after a live migration. The default implementation reads the system clock and fires when it passes the scheduled instant, which means a node whose clock runs fast fires early. Early is the dangerous direction: scheduled work overwhelmingly reads a window of data that is defined by the scheduled instant, so a job that fires before its instant reads a window that has not closed, produces a short report, and looks like it succeeded. Late work reads the right window and arrives inconveniently. These two failures are not symmetrical and should not share a budget.

Decision. Dispatch decisions are taken only against a time source that reports a bounded uncertainty interval, and the scheduler waits out that interval rather than firing at the earliest instant it permits — scanning for instants due at or before t−ε. Lateness is measured, bounded and published; earliness has no acceptable rate at any percentile. A node whose reported uncertainty exceeds a declared ceiling sheds its partition leases and continues serving reads rather than firing on a clock it cannot justify.

How it is realised on AWS. The timing plane takes time from Spanner commit timestamps and the platform's bounded-uncertainty time service rather than from clock_gettime. The due scan selects instants <= now_lower_bound, so the scheduler is deliberately ε late. Each partition owner records the ε it observed at lease grant on the lease row; above 100 ms it releases the lease and the lease manager reassigns the partition. Clock ε is a per-node metric, not an aggregate, because a fleet-wide drift is invisible in an average.

Option Verdict Reasoning
Bounded-uncertainty source, wait out ε, shed leases above a ceiling Chosen Guarantees no earliness for about 10 ms of deliberate lateness; needs a time service that reports its own error
Trust the NTP-synchronised system clock Rejected Free and usually fine; produces early fires exactly when a node is unhealthy, and no node can tell that it is the wrong one
Take time only from the store's commit timestamp Rejected Close to chosen and simpler; a read per scan cycle is acceptable but leaves no per-node signal with which to shed a bad clock
Fire at the midpoint of the uncertainty interval Rejected Halves the deliberate lateness and makes earliness a rate rather than an impossibility

What it buys

  • No fire is ever early, at any percentile, which removes an entire class of silent wrong answers from the work the scheduler triggers
  • A bad clock is a detectable local condition with a local response, rather than a fleet-wide correctness problem
  • Short leases become safe, which is what makes a 15 s takeover budget achievable (ADR-11)

What it costs

  • Every fire pays ε of deliberate lateness — cheap at 10 ms, and the decision would not survive a bound of 500 ms
  • The design depends on a platform capability (a clock that reports its own error) that is not available everywhere, which constrains portability
  • A node sheds leases on a clock problem, so a correlated clock event reduces scheduling capacity exactly when nothing is wrong with the schedulers

Choose differently when. If the scheduler only triggered work whose correctness did not depend on its instant — notifications, cache warms, retries — earliness would be harmless and the simplest clock would do. The decision is justified by work that reads a window the scheduled instant defines.

Why it holds up over time. The rule survives every change of clock technology, because it is a statement about which direction of error is acceptable rather than about how accurate the clock is. If clocks get better, ε shrinks and the deliberate lateness disappears; if they get worse, the ceiling catches it. What would break the decision is a platform with no bounded-error time source at all, which would force the weaker commit-timestamp variant above.

Lesson. When a system's output depends on a measurement, ask which direction of measurement error is survivable, and spend the error budget entirely in that direction.

ADR-05 · Recurrences are recomputed each cycle from a versioned definition, not materialised ahead into the due index

Status: Accepted · Shown on views: 09, 11, 17

Should the next thousand fire instants be computed now and stored, or computed each time they are needed?

Context. A calendar recurrence is a function of an expression, a time zone, a time-zone database and a validity window. Materialising its instants weeks ahead makes the due index trivially cheap to scan: it is already a sorted list of absolute times. It also creates a cache that four separate events invalidate — an amendment, a pause, a time-zone database update, and a change to the validity window — and the invalidation has to find every materialised instant for one trigger among many millions. Recomputing instead costs CPU on every scan cycle, proportional to the trigger population rather than to the instants actually due, and it makes the correct answer the only answer the system can give.

Decision. The due index holds exactly one row per live trigger, carrying its single next instant, recomputed and advanced when that instant is claimed. The recurrence is evaluated against the trigger's named zone at the moment of computation. A definition is immutable once fired against; an amendment creates a new version and the index row is recomputed under it, so the instant the scheduler uses is always derived from the rules currently in force.

How it is realised on AWS. The recurrence compiler is a pure function in the control plane and the timing plane, shared as one library and exercised by dry-run on the production path, so the next three instants shown to a developer are the instants the scheduler will use. The due index row carries trigger_id, version_id and next_instant; claiming an instant and advancing the row happen in the transaction that inserts the fire (ADR-01). An amendment writes a new trigger_version and recomputes the row within the 5-second propagation budget.

Option Verdict Reasoning
One next instant per trigger, recomputed on advance Chosen Always correct under the current rules, with no invalidation to get wrong; costs a compiler call per fire and per amendment
Materialise instants weeks ahead Rejected Cheapest possible scan; four separate invalidation paths, each of which silently produces a fire at an instant the current rules disagree with
Materialise a short horizon, say one hour Rejected A real middle ground and the most likely future revision; still needs invalidation, and buys little while the scan is per-partition
In-memory timer wheels per partition Rejected Lowest dispatch latency; makes the authoritative structure non-durable, so a restart re-derives it anyway

What it buys

  • A time-zone database update, an amendment and a pause are all picked up with no invalidation sweep, because nothing was stored to invalidate
  • Dry-run is exact rather than indicative, which is what makes it useful at the moment a developer commits a schedule
  • The due index stays small — one row per trigger — so a partition scan is proportional to what is due rather than to history

What it costs

  • Recurrence evaluation sits on the scan path, so a slow or pathological expression is a scheduling cost rather than a one-off authoring cost
  • A trigger population of 20 million means 20 million index rows kept current, which is a write rate proportional to the fire rate
  • There is no cheap way to answer "show me every fire in the estate next Tuesday", which a materialised index would give for free

Choose differently when. If recurrence evaluation were expensive — tenant-supplied predicates against business calendars, say, as the ask defers to Phase 3 — the balance would move towards a short materialised horizon with explicit invalidation. For three fixed recurrence forms it does not.

Why it holds up over time. The decision is really about where correctness lives: in a recomputation, or in a cache plus its invalidation. That trade does not change with hardware. If evaluation cost ever dominates, the short-horizon variant is a localised change to the due index's contract and leaves the ledger, the key and the lanes untouched.

Lesson. A cache of derived values is a commitment to invalidate it correctly on every input that feeds it. Count those inputs before you build the cache.

ADR-07 · Wall-clock recurrences resolve against a named zone with declared DST rules, and the time-zone database version is recorded on every fire

Status: Accepted · Shown on views: 09, 11, 20

What does "every day at 02:30" mean on the day 02:30 happens twice, and on the day it does not happen at all?

Context. A calendar recurrence in a named zone has two days a year where it is ambiguous or impossible, and a scheduler must have an answer for both. Worse, the mapping from wall-clock to absolute instant is not fixed: time-zone databases are updated several times a year, and a government changing its rules moves a future instant that the scheduler may already have told a tenant about. Storing instants in UTC and forgetting the zone makes the system immune to this and also wrong — a tenant who asked for 09:00 local means 09:00 local after the rules change, not 08:00. Any system that does not record which rules it used cannot explain, six months later, why a fire landed where it did.

Decision. Calendar recurrences store the wall-clock expression plus the IANA zone and resolve against the zone at the moment of computation. A wall-clock time that does not exist on a given day fires at the instant the clock jumps to; a wall-clock time that occurs twice fires on the first occurrence only. Both outcomes are visible in history as such. The time-zone database version used is recorded as a column on every fire record.

How it is realised on AWS. The recurrence compiler pins a tzdata version per deployment and writes it to the trigger_version row and to every fire. The two DST rules are implemented once in the shared compiler, so dry-run shows the same substitute instant production will use. A tzdata upgrade is a deliberate deployment that re-resolves future instants; the version on past fires explains any instant a tenant disputes.

Option Verdict Reasoning
Expression + zone, resolved at computation, version recorded Chosen Means what the tenant asked; the recorded version is what makes a disputed instant explainable
Convert to UTC at authoring time Rejected Immune to rule changes and simple; silently wrong for the tenant twice a year and permanently wrong after a rule change
Fire on both occurrences of an ambiguous time Rejected Defensible for monitoring work; produces a double billing run once a year
Skip a non-existent wall-clock time entirely Rejected Honest, and means a daily job silently does not run one day a year — the failure nobody notices until the month-end total is short

What it buys

  • The tenant's expression means what they think it means, including after their government changes the rules
  • The two hard days a year have a declared, testable, documented behaviour rather than whatever the library did
  • A disputed instant is explainable from data, because the rules in force are recorded alongside the fire

What it costs

  • A tzdata upgrade is a change to future behaviour and needs treating as a deploy with a review, not a dependency bump
  • Zone resolution on every computation adds cost to the scan path (ADR-05 already pays for recomputation)
  • The "first occurrence only" rule will surprise someone whose job genuinely wanted both

Choose differently when. If every tenant were in one zone with no daylight saving, the whole decision would collapse into storing UTC. It is justified by a multi-tenant platform whose tenants' customers are the ones whose local morning matters.

Why it holds up over time. Time-zone rules will keep changing and the IANA database will keep being the way to track them. Recording the version used is a general principle — any derived value computed from an external ruleset should carry the ruleset's version — and it does not depend on anything in this stack.

Lesson. If a value is computed from rules that someone else can change, store the version of the rules alongside the value, or you will one day be unable to explain your own output.

Policy after the outage

What a trigger does about the fires it missed — declared before it misses them, not decided during the incident.

ADR-06 · The missed-fire policy is declared per trigger before the outage, and defaults to firing once for the most recent missed instant

Status: Accepted · Shown on views: 04, 13, 17

Who decides what a trigger does about the fires it missed, and when is that decision taken?

Context. After any interruption — a platform incident, a long pause, a trigger resumed after a weekend — there is a set of instants that should have fired and did not, and three defensible things to do with them. Fire all of them in order, fire once for the most recent, or record them as skipped and fire nothing. Which is correct depends entirely on what the work means, and the platform cannot know that: a billing run and a cache warm have opposite correct answers and look identical from the scheduler's side. The alternative to the tenant declaring it is an operator deciding during the recovery, across 50,000 tenants at once, under pressure, with no information about any of them — a decision that is guaranteed to be wrong for most of the estate.

Decision. Every trigger declares its missed-fire policy at definition time from a closed set — fire-all, fire-once-now, skip — and the platform default is fire-once-now. The policy is part of the versioned definition, so the decision exists before the outage that needs it, and the recovery executes a declaration rather than improvising one.

How it is realised on AWS. The policy is a column on trigger_version, surfaced on the authoring screen as the one distributed-systems question a developer is asked, with each option's consequence stated in a sentence. The catch-up path in view 13 applies the horizon filter first (ADR-08) and the policy second. Dry-run shows what a policy change would have done to the last recorded gap, so a tenant can change it with evidence.

Option Verdict Reasoning
Per-trigger declaration, default fire-once-now Chosen Puts the decision with the only party who knows what the work means; costs one question at authoring time and a default that will be wrong for some
Platform-wide policy Rejected Nothing to configure and nothing to get wrong; guarantees the wrong behaviour for a large share of the estate
Operator decision during recovery Rejected Maximum flexibility exactly when there is no capacity to use it, and no information to use it with
Infer from the trigger's recent behaviour Rejected Plausible and seductive; a scheduler guessing whether work is idempotent is a scheduler that will be wrong about money

What it buys

  • Recovery is deterministic and rehearsable: the behaviour after an outage is a property of the definition, not of who was on call
  • The default is the one least likely to duplicate billable work, which is the failure a tenant will escalate
  • A tenant who needs fire-all declares it, and the declaration is visible to reviewers in their infrastructure code

What it costs

  • Every tenant is asked a question most of them would rather not answer, at the moment they least want friction
  • A tenant who accepted the default and wanted fire-all loses work they expected; the cause is recorded but the work is gone
  • Three policies × the horizon × the overlap policy is a behaviour matrix that documentation has to make legible

Choose differently when. If the platform owned the work as well as the timer — a managed job runner where every job declared its own idempotency — the platform could infer the policy and the question would be unnecessary. With opaque tenant payloads behind opaque tenant endpoints, it cannot.

Why it holds up over time. "Declare the recovery behaviour before the failure" is a general property of operable systems and does not depend on this stack. The specific default might move: if telemetry shows most tenants override it, the default is wrong and should change. The closed set of three is the part worth defending, because an open-ended policy language here would be a scripting surface on the fire path.

Lesson. If a system will have to make a decision during an incident that it cannot make correctly without information it does not have, collect that information before the incident and call it configuration.

ADR-08 · A catch-up horizon bounds every missed-fire policy, and the original scheduled instant travels with the fire

Status: Accepted · Shown on views: 05, 13, 14

A trigger is resumed after a month. How much work does it owe?

Context. The missed-fire policy (ADR-06) answers what to do with missed instants but not how many of them there can be. A minute-schedule paused for a month and resumed under fire-all owes 43,200 fires; the tenant who clicked resume expected a job to start running again, not a month of history to arrive. The same unbounded set appears after a long incident and after an authorised backfill. Separately, work triggered late almost always reads a window of data defined by its scheduled instant: a statement run for 1 March that executes on 3 March must still report February. A fire that only knows when it was dispatched cannot do that.

Decision. Every trigger carries a catch-up horizon — default 1 hour, maximum 24 hours — and an instant older than the horizon is recorded as expired with its cause and never dispatched, whatever the missed-fire policy says. This is the one place the platform overrules the tenant's declaration. Every caught-up fire carries its original scheduled instant, distinct from its dispatch instant, and both are visible in history.

How it is realised on AWS. The horizon filter runs before the policy in the catch-up path, writing a fire row in expired state with a cause so the tenant can see the decision rather than a gap. The fire's scheduled_instant is part of its primary key and is what the dispatch payload carries; dispatched_at lives on the attempt row. Horizon extension beyond the default is a privileged operation (ADR-15).

Option Verdict Reasoning
Platform-enforced horizon, original instant preserved Chosen Bounds the worst case absolutely; costs a tenant some work they may have wanted
Unbounded catch-up under the tenant's policy Rejected Most faithful to the declaration; one resume becomes a self-inflicted denial of service against the tenant's own endpoint
Horizon as a soft limit with a warning Rejected Keeps the tenant in control; a warning nobody reads is an unbounded catch-up with extra steps
Dispatch with the dispatch instant only Rejected Simpler payload; makes every late fire produce work about the wrong window

What it buys

  • The worst case after any interruption is bounded by a number the tenant can see and reason about
  • Late work is still correct work, because the window it reads is defined by the instant it was scheduled for
  • An expired instant is a visible decision with a cause rather than a silent absence, which is what makes view 05's journey recoverable

What it costs

  • A tenant who genuinely needed a month of catch-up must request a backfill, which is a privileged, approved operation
  • The horizon is a default and defaults are not read; the first time it matters is during an incident (view 05's trough)
  • Two timestamps on every fire is a payload and documentation cost that tenants will initially get wrong

Choose differently when. If the triggered work were always idempotent and cheap, an unbounded catch-up would be harmless and the horizon would be needless friction. The horizon is justified by work that costs money or sends mail every time it runs.

Why it holds up over time. A bound on the recovery set is a general requirement of any system that accumulates obligations while it is down, and the principle survives any implementation. The specific defaults are the arguable part and should move with observed data — which is why they are stated as assumptions rather than constants.

Lesson. Any queue that fills while a system is down needs a stated maximum age, decided before the outage. Without one, recovery is unbounded by construction.

ADR-09 · Catch-up, backfill and on-time fires use separate queues with a per-tenant catch-up cap and a published shedding order

Status: Accepted · Shown on views: 07, 13, 14

A backlog is released. What stops its drain from being the next outage?

Context. Recovery is the most dangerous moment in a scheduler's operation, and it looks like success while it happens. Fires that have been accumulating are suddenly all due, and the natural behaviour — dispatch them as fast as the pipeline allows — turns a scheduling incident into an incident for every executor downstream, including the ones that had nothing to do with it. A single shared queue with priorities does not help as much as it appears to: priorities decide ordering, but a saturated pipeline is saturated for everyone, and the lane that should have been throttled is competing with the lane that should not have been.

Decision. On-time, catch-up and backfill fires are dispatched through separate queues with independent drain rates. A tenant's catch-up dispatch is capped at a fraction of its steady-state quota — assumed at 10% — so a backlog drains over hours rather than minutes. The shedding order is published: backfill first, then catch-up, then over-quota steady-state traffic, and only then on-time fires. Recovery is deliberately slower than the failure that caused it.

How it is realised on AWS. Three Cloud Tasks queues per region, each with its own dispatch rate and concurrency, fed by the rate shaper. Catch-up drains oldest-first within a trigger and interleaves fairly across a tenant's triggers, so no single trigger's backlog starves the rest. A retry stays in the lane its fire came from, and its attempt budget is truncated at the next scheduled instant. A tenant whose targets are systematically failing has its retry volume charged to its own quota.

Option Verdict Reasoning
Separate queues per class, per-tenant catch-up cap, published shedding order Chosen Makes the throttle structural rather than a tuning value; costs three queues to operate and a long drain the tenant must accept
One queue with priorities Rejected Fewer moving parts; under saturation the priority decides order, not whether the backlog is throttled
Unthrottled drain with autoscaling Rejected Fastest recovery for the platform; moves the incident to every tenant executor at once
Withhold the decision so the backlog never materialises Rejected A real alternative (Question 5 in the ask); keeps the ledger smaller but makes the backlog invisible and un-queryable

What it buys

  • An outage's recovery cannot become a second outage, and the limit is a published number rather than an operator's judgement
  • The shedding order means degradation is a decision taken in advance, which is what lets an SRE predict what will happen under load
  • A tenant with a failing endpoint pays for its own retries instead of consuming shared dispatch capacity

What it costs

  • A one-hour backlog takes roughly ten hours to drain, which is a long time to tell a tenant their work is still coming
  • The oldest-undispatched-due-age alert is loud for the whole drain, because the backlog is deliberately visible
  • Three queues, three sets of capacity and three sets of saturation behaviour to operate and to reason about

Choose differently when. If every executor were elastic and owned by the platform, an unthrottled drain with autoscaling would recover faster at no external cost. With opaque tenant endpoints of unknown capacity, the conservative drain is the only responsible default.

Why it holds up over time. "Recovery slower than failure" is a control-theory statement about a system with a feedback loop into its own dependencies, and holds regardless of queue technology. The cap's value is the arguable part; the existence of a cap and of a published order is not.

Lesson. Size your recovery path for the capacity of whoever has to absorb it, not for the capacity of the system doing the recovering.

ADR-12 · Overlap is enforced against recorded outcome state, with unknown as a first-class terminal state bounded by a window

Status: Accepted · Shown on views: 11, 14, 17

The previous execution has not reported back. Is it still running?

Context. A trigger whose work sometimes takes longer than its interval will eventually be due again while the previous run is still going, and three behaviours are defensible: run anyway, skip, or queue one. Enforcing skip or queue requires the scheduler to know whether the previous execution has finished — which means depending on a signal it does not control, from an executor it does not own. When that signal never arrives, there are exactly two choices and both are bad. Treat the missing callback as finished and risk genuine concurrent execution; treat it as still running and withhold every subsequent fire indefinitely, which is a silent outage that looks like a working trigger with no fires.

Decision. Overlap policy is declared per trigger (allow, skip, queue) and enforced against recorded outcome state, bounded by the trigger's outcome window. A previous fire still in unknown state at the end of its window is treated as finished for overlap purposes, and that decision is recorded. unknown is a first-class terminal state, not an absence. A dispatch attempt that times out is recorded as unknown rather than failed, because the target may have accepted it. Under queue, at most one pending fire is held; a newer instant discards and records the older one.

How it is realised on AWS. work_outcome is a 1:0..1 child of the fire; the absent row is the unknown state. The timing plane's overlap check reads the previous fire's state and its age against the outcome window before claiming the next instant, and writes the resulting skip with a cause when it declines. The outcome window is a column on the trigger version.

Option Verdict Reasoning
Outcome state with a bounded window, unknown treated as finished Chosen Bounds the silent-outage failure, which is the one a tenant cannot see; accepts occasional concurrent execution
Treat unknown as still running Rejected Never allows concurrency; a single lost callback stops a trigger forever and looks healthy while doing it
Delegate overlap entirely to the executor Rejected Architecturally cleanest — the scheduler stays stateless about executions; leaves every tenant to build the same lock
Hold a lease for the execution's duration Rejected Strongest mutual exclusion; makes the scheduler's lease lifetime depend on tenant work it cannot bound

What it buys

  • A lost callback costs at most one window of uncertainty rather than a permanently stopped trigger
  • Every overlap decision is recorded with a cause, so a tenant sees "skipped: previous still running" instead of a gap
  • A timed-out attempt is honestly recorded as unknown, which keeps the duplicate-tolerance contract (ADR-03) coherent

What it costs

  • The scheduler is now stateful about executions it does not run, which is a dependency on a signal it cannot guarantee
  • Under a systematically missing callback, skip degrades towards allow, and the tenant may not notice the degradation
  • The outcome window is another per-trigger value to choose badly

Choose differently when. If every executor reliably reported terminal state — a platform-owned runner, not a tenant endpoint — unknown would be rare enough that treating it as still running would be the safer default. With tenant HTTP targets it is not.

Why it holds up over time. The forced choice between risking concurrency and risking silence is structural: it exists whenever one system's decision depends on another system's unreliable report, and no technology removes it. What a future change could do is shrink the unknown population, which would make the window shorter without changing the rule.

Lesson. When a decision depends on a signal that may never arrive, name the missing signal as a state, bound how long you will wait for it, and write down which of the two bad outcomes you chose.

Scale, cost and recovery

Absorbing the top-of-hour peak, and treating the due index as something rebuilt rather than restored.

ADR-10 · The top-of-hour peak is absorbed by opt-out jitter smearing and the lateness budget, not by provisioned capacity

Status: Accepted · Shown on views: 02, 14, 16

Sixty per cent of all fires land in the first second of a minute. Do we buy capacity for that second?

Context. Scheduled work clusters on human-legible instants: the top of the minute, the top of the hour, midnight, midnight UTC, 09:00. At an assumed 5,800 fires a second mean, the design peak is around 45,000 a second for five seconds at the top of the hour — roughly 8× the mean minute. Provisioning the dispatch tier for that peak means paying for capacity that is idle 99.9% of the time; the alternative is to accept that a fire is a little late. The question is which fires genuinely care about their exact instant, and the honest answer is that most do not: a digest email at 09:00:03 is indistinguishable from one at 09:00:00, while a market-open task at 09:00:03 may be worthless.

Decision. Fires colliding on a popular instant are smeared within a declared jitter window, with per-trigger opt-out where the instant is externally significant. The peak is absorbed by the jitter window, a dispatch queue with a declared drain rate, and the published lateness budget — not by capacity provisioned for the peak second. The peak capacity avoided by smearing is reported, so the trade is visible rather than assumed.

How it is realised on AWS. The rate shaper applies a deterministic per-trigger offset within the window, so a given trigger's smear is stable rather than randomly different each hour — which matters because a tenant watching their own fires should see a consistent pattern. Opt-out is a flag on the trigger version. The avoided-peak figure is a BigQuery report alongside cost per million fires.

Option Verdict Reasoning
Opt-out jitter smearing plus the lateness budget Chosen Removes the peak as a capacity problem; costs instant exactness for triggers that silently cared
Provision the dispatch tier for the peak Rejected Exact instants for everyone; pays for 8× capacity used five seconds an hour
Opt-in jitter Rejected Safest for correctness; almost nobody opts in, so the peak remains and the mechanism is dead weight
Queue at the peak with no smearing Rejected Half of the chosen option, and the one that makes the whole peak arrive as lateness on the unlucky triggers rather than spread across all of them

What it buys

  • Steady-state capacity is sized for the mean rather than the peak, which is the single largest cost lever in the design
  • The smear is deterministic per trigger, so a tenant sees a stable pattern rather than jitter that looks like instability
  • The avoided peak is reported, so the decision can be revisited with a number instead of an argument

What it costs

  • Triggers whose instant mattered and that did not opt out are late by up to the window, and nobody finds out until it matters
  • The opt-out is discoverable mainly by reading the per-trigger lateness distribution, which not every tenant will do
  • "Scheduled for 09:00" is now a statement about a window, which has to be documented honestly

Choose differently when. If the service's tenants were predominantly financial — market opens, settlement windows, regulatory cut-offs — the exactness would be the product and provisioning for the peak would be the right answer. For a general developer platform it is not.

Why it holds up over time. The trade between buying peak capacity and spending a latency budget is permanent and technology-independent; only the prices move. If serverless dispatch capacity became genuinely instantaneous and free at the margin, provisioning for the peak would win, and the jitter window could be set to zero without touching anything else in the design.

Lesson. Before buying capacity for a peak, find out how many of the requests in that peak actually care about arriving in it.

ADR-11 · The due index is a rebuildable projection, verified continuously by a shadow rebuild rather than backed up

Status: Accepted · Shown on views: 07, 10, 15

The due index is corrupt. Do we restore it, or recompute it?

Context. The due index is the structure the scheduler scans, and it is derived entirely from the trigger registry: for each live trigger, the next instant its recurrence produces. Backing it up treats it as data; recomputing it treats it as a function. The second is obviously correct and is also the claim that quietly stops being true — a recompute path that is never exercised is a restore procedure with better marketing, and the first time it runs is during the incident that needed it. The same applies to partition reassignment: a takeover path exercised only during failures is a path nobody has confidence in at the moment they need it most.

Decision. The due index, the partition leases and every cache are rebuildable projections of the registry and are not backed up. The rebuild is exercised continuously: a shadow rebuilder recomputes a partition's instants from the registry and compares them against the live index, and any divergence is treated as an incident. Partition ownership moves routinely in production rather than only during failure, so the takeover path is the common path.

How it is realised on AWS. The shadow rebuilder is a component in the timing plane (view 07), not a script. It walks partitions on a rota, recomputes next instants with the same shared compiler the scan path uses, and reports divergence to BigQuery. Lease reassignment is driven on a schedule as well as by expiry, so the 15-second takeover budget is measured continuously rather than estimated. Only the registry, the ledger and the audit store are backed up.

Option Verdict Reasoning
Rebuildable, verified by continuous shadow rebuild, routine reassignment Chosen Makes recoverability a measured property; costs a permanent component and the compute it consumes
Back up the due index Rejected Familiar and cheap to set up; restores a point in time, which for a derived structure is strictly worse than recomputing the current one
Rebuildable, with rebuild tested in a lower environment Rejected The common compromise; tests the code, not the production data, which is where the divergence will be
Rebuild on demand, no verification Rejected Correct in principle; the claim decays silently and is discovered false during an incident

What it buys

  • A corrupt or lost due index is a recompute rather than a recovery procedure, and the recompute has been running all along
  • Takeover is the common path, so the 15-second reassignment budget is a measurement rather than a hope
  • The backup surface is three stores instead of eight, which makes the backup story small enough to actually verify

What it costs

  • A permanent component and its compute, spent entirely on verifying something that is supposed to be true
  • Shadow divergence is reported rather than alerted in the MVP, so there is a window where the claim is unverified
  • Routine reassignment means the fleet is always mid-handover somewhere, which makes duplicate-key rejections normal traffic rather than an anomaly

Choose differently when. If the due index were expensive to recompute — materialised instants for 20 million triggers weeks ahead, as ADR-05 rejected — rebuild would stop being cheaper than restore and the decision would invert. The two decisions are linked: recomputation is affordable because the index holds one row per trigger.

Why it holds up over time. "Derived data is rebuilt, not restored" is a property of the data's relationship to its source and survives any storage technology. The part that needs continued investment is the verification, because the claim is the kind that decays. If the shadow rebuilder is ever switched off for cost, the decision has quietly become the rejected fourth option.

Lesson. A recovery path you do not run is a recovery path you do not have. If something is rebuildable, rebuild it on a schedule and compare.

Evidence and trust

What the tenant can see, what a schedule is allowed to point at, and who is allowed to multiply work.

ADR-13 · Fire history is a tenant-facing product surface, and the platform alert is oldest undispatched due age

Status: Accepted · Shown on views: 05, 16, 17

How does anyone find out that the scheduler has stopped firing?

Context. A scheduler's characteristic failure is unique among services: it is up, serving its API, passing its health checks, consuming no unusual resources, and not firing. Every dashboard says healthy. CPU, memory, error rate, request latency and queue depth are all normal, because the absence of work produces no signal. The same blindness applies to the tenant: a job that did not run generates nothing — no error, no log line, no alert — so the first notification is usually a customer. Both problems are the same problem viewed from two sides, and both are solved by measuring the gap between what should have happened and what did.

Decision. The platform's paging alert is oldest undispatched due instant age per partition, not process health. Fire history is a tenant-facing product surface recording, per fire: scheduled, decided and dispatched instants, attempt outcomes, terminal state, the cause of any non-dispatch, the definition version and the tzdata version — with each trigger's recent lateness distribution published so a tenant sees punctuality degrade before it becomes an incident.

How it is realised on AWS. Fire history is a Bigtable projection of the ledger keyed by tenant#trigger#instant with a 90-day TTL and a 13-month Cloud Storage archive. Oldest undispatched due age is computed per partition from the due index and alerts above an assumed 60 seconds. The History API is a separate service from the Management API (view 08) so a history query never competes with authoring, and the per-trigger lateness view is derived in BigQuery.

Option Verdict Reasoning
Due-age alerting plus tenant-facing per-fire history and lateness Chosen Catches the invisible failure from both sides; costs a projection, an archive and a product surface
Health and resource alerting Rejected Standard and cheap; blind to the one failure that matters most here
Alert on fire rate dropping below a baseline Rejected Useful as a secondary signal; a legitimate quiet period and a stalled scheduler look identical
Platform-internal history only, support answers tenant questions Rejected Less to build; makes every "did it run" a support ticket and leaves the tenant's on-call blind (view 05)

What it buys

  • A stalled scheduler pages within a minute rather than being discovered by a tenant's customer
  • "Did my job run last night" is answerable by the tenant, with a cause, which is what makes view 05's journey recoverable
  • Three separately recorded instants per fire turn ambiguous complaints into specific ones

What it costs

  • The due-age alert is loud for the whole of a legitimate catch-up drain, which is the price of a visible backlog (ADR-09)
  • 500 million fires a day of history is a storage and query cost that is pure evidence and produces no fires
  • A per-fire record for work the platform did not run invites tenants to read "dispatched" as "succeeded"

Choose differently when. If the scheduler only served platform-internal callers with their own observability, the tenant-facing surface would be unnecessary and the alert alone would do. The product surface is justified by tenants who have no other way to see the gap.

Why it holds up over time. "Alert on the gap between intended and actual, not on the health of the thing in between" is a general principle for any system whose failure mode is silence, and it outlasts any monitoring stack. The specific alert threshold should move with measured data.

Lesson. For any system whose failure produces no output, make the alert the distance between what should have happened and what did. Health checks cannot see absence.

ADR-14 · A target must be verifiably owned by its tenant, no tenant request reaches the fire zone, and dispatch credentials are minted per attempt

Status: Accepted · Shown on views: 08, 18, 19

What stops a schedule from being a timer-driven request forgery primitive aimed at someone else?

Context. A scheduler that will make an authenticated HTTP request to any URL a tenant names, repeatedly, on a timer, with a retry budget, is a remarkably convenient attack tool. It has outbound network access, it is inside the platform's perimeter, it holds credentials, and it is designed to keep trying. The same properties make the dispatch credential valuable: a long-lived platform identity captured from a dispatch grants the holder whatever that identity can do, for as long as it lives. Both problems come from treating the dispatch target as data rather than as a privilege.

Decision. A target must be a registered endpoint on a domain the tenant has verified, or a queue the tenant owns, checked at definition time and again at dispatch time. No tenant request reaches the timing or dispatch planes: both are reachable only from inside the perimeter, and the only component with egress is the attempt runner. Dispatch credentials are minted per attempt with a lifetime no longer than the attempt timeout, and every dispatch is signed with a timestamped, replay-resistant signature over the fire record and payload. Inbound outcome callbacks are verified by the same signature within a short window.

How it is realised on AWS. Target verification is a separate target entity with its own verification state (view 11), referenced by the trigger version rather than embedded as a URL. The attempt runner mints a workload-identity token scoped to the trigger for each attempt, so a captured dispatch buys one trigger for one timeout. The callback verifier sits at the perimeter (view 18) and rejects a stale signature rather than logging it.

Option Verdict Reasoning
Verified target entity, perimeter isolation, per-attempt credentials, signed both ways Chosen Closes the forgery path and bounds credential capture; costs tenants a domain verification step
Allow any HTTPS URL with an allowlist of blocked ranges Rejected Lowest friction; a denylist of internal ranges is a game the defender loses eventually
Long-lived per-tenant dispatch credential Rejected Simpler to operate and to debug; a captured credential is a standing grant
Egress through a tenant-supplied proxy Rejected Pushes the problem to the tenant and is right for some enterprises; too much setup for a general platform

What it buys

  • A schedule cannot be pointed at a host the tenant does not demonstrably control, so the service is not an amplifier
  • A captured dispatch credential expires with the attempt, which reduces a standing grant to a seconds-long one
  • One component has egress, so the egress policy is small enough to audit

What it costs

  • Domain verification is friction at exactly the moment a developer wants to ship (view 04)
  • Verification is point-in-time: a target legitimately owned at definition and later reassigned is a residual risk mitigated only by re-checking at dispatch
  • Per-attempt minting adds a call to the dispatch path at 45,000/s peak, which has to be cached carefully without defeating the point

Choose differently when. If the scheduler only dispatched to queues and executors inside the platform, target verification would be unnecessary and the perimeter would do the whole job. The tenant HTTP target is what makes this decision load-bearing.

Why it holds up over time. "A dispatch target is a privilege, not a parameter" survives any identity technology, and per-attempt credentials get cheaper as token minting does. What may change is the verification mechanism — domain verification is a convention, not a law — but the requirement that ownership be proved does not.

Lesson. Any feature that lets a user name an outbound destination is a request-forgery feature until ownership of that destination is proved. Treat the destination as a privilege with a lifecycle.

ADR-15 · Backfill, horizon extension and quota change are grants separate from trigger authorship

Status: Accepted · Shown on views: 08, 18, 19

Should the person who can create a schedule also be able to make it emit a month of work?

Context. Most operations on a scheduler are small: create a trigger, amend it, pause it, read its history. Three are not. Backfill generates fires for a declared past window. Extending the catch-up horizon turns a bounded recovery into a larger one. Raising a quota lifts the cap on dispatch rate and in-flight concurrency. Each of them multiplies real work — mail sent, money moved, endpoints called — and each of them is most attractive at the moment judgement is worst, which is during or just after an incident. Bundling them with authorship means the blast radius of a routine role is the blast radius of the largest operation in the system.

Decision. Trigger authorship, policy change (catch-up horizon, overlap window, quota) and read-only history access are three separately grantable roles, and backfill is its own privileged operation with its own API surface. Every privileged act is written to an append-only, tamper-evident audit record with actor and prior value. An author who can create a trigger cannot widen its blast radius.

How it is realised on AWS. Three inbound APIs (view 08) rather than one, so the authorisation story is visible in the surface rather than buried in a handler. View 19 draws the horizon extension being denied to an author, because the split is only real if the refusal happens. Backfill is additionally rate-capped in its own dispatch lane (ADR-09) and bounded by a declared maximum volume per operation.

Option Verdict Reasoning
Separate grants for authorship, policy and privilege, with audit Chosen Bounds the routine role to routine damage; costs an approval step during incidents
One role for everything a tenant can do Rejected Simplest to grant and to explain; makes every trigger author able to emit a month of billing runs
Approval workflow for every change Rejected Safest; makes ordinary authoring slow enough that tenants route around the platform
Rate limits instead of privileges Rejected Bounds the volume but not the intent, and a limit high enough to be useful is high enough to hurt

What it buys

  • The routine role's worst case is a badly scheduled trigger, not a month of duplicated work
  • An audit trail with prior values makes a bad backfill explainable and attributable afterwards
  • The separation is visible in the API surface, so a security reviewer can check it without reading handler code

What it costs

  • Recovery needs an approver, which adds minutes to the moment a tenant most wants to act (view 05)
  • A tenant that grants the policy role to everyone with the author role recreates the problem; audit detects it and nothing prevents it
  • Three roles and one privileged surface is more to document and more to get wrong in a tenant's own access model

Choose differently when. If backfill were cheap and harmless — regenerating a derived view, say — the separation would be needless ceremony. It is justified by backfill emitting the same work as a real fire, indistinguishable at the executor.

Why it holds up over time. Separating the privilege to do a thing from the privilege to multiply it is a general access-control principle and does not depend on this identity provider. The specific split of three roles may be refined; the principle that recovery operations are privileged should not be.

Lesson. Find the operations in your system that multiply work rather than performing it, and make each one a separate grant. They are the ones that will be used under pressure.

ADR-16 · Tenant payloads live only in the registry and the ledger, encrypted under a per-tenant key, and never in evidence or logs

Status: Accepted · Shown on views: 10, 11, 18

The scheduler carries an opaque blob it never interprets. Where is that blob allowed to exist?

Context. The tenant payload is the one piece of data in the system the platform cannot reason about. It is opaque by design — the scheduler does not interpret it — which also means the platform cannot know whether it contains a customer identifier, an account number, a webhook secret or a medical record. Every piece of machinery the service has wants to copy it: a dispatch log for debugging, a fire history row for the tenant's own view, an error message when a dispatch fails, a metric label for cardinality analysis, a reporting table for cost attribution. Each copy is a place that now carries unknown sensitive data under a different retention and a different access policy.

Decision. A tenant payload exists in exactly two places — the trigger registry and the fire ledger — encrypted at rest under a per-tenant key, and is never written to fire history, reporting, metrics, traces or logs. It is passed to the executor and nowhere else. On tenancy termination, definitions, payloads and history are deleted within 30 days while non-identifying fire counts are retained for billing.

How it is realised on AWS. The payload is stored by reference (payload_ref) on the trigger version and resolved only in the dispatch path, where the attempt runner decrypts it under the tenant's key immediately before signing and sending. History rows (view 10) carry the fire key, timestamps, states and causes, and never the payload. Error messages carry the fire key, so a debugging path starts from an identifier rather than from content.

Option Verdict Reasoning
Two stores, per-tenant key, never in evidence Chosen Keeps the unknown-sensitivity blob inside one small boundary; makes debugging start from a key rather than from content
Payload in history for tenant convenience Rejected Genuinely useful for a tenant diagnosing a bad fire; puts unknown sensitive data under a 90-day analytics retention
One platform key for all payloads Rejected Much simpler key management; removes crypto-shredding as a deletion mechanism and makes one key the whole estate's boundary
Reject payloads entirely; tenants fetch their own parameters Rejected Cleanest privacy position; makes every trigger need a config lookup the tenant must build

What it buys

  • The sensitive-data surface is two stores with one retention story, which is small enough to actually review
  • A per-tenant key makes deletion enforceable by key destruction rather than by trusting every copy to be found
  • Debugging conventions start from the fire key, which is a habit that keeps content out of logs by default

What it costs

  • A tenant cannot see the payload a fire carried, which makes some diagnosis harder than it needs to be
  • A decrypt per attempt on the dispatch path at peak, which is a real cost and a dependency on the key service
  • Per-tenant keys at 50,000 tenants is key lifecycle work the platform now owns

Choose differently when. If payloads were constrained to a declared schema the platform validated — identifiers only, no free text — the sensitivity would be known and history could safely carry them. Opacity is what forces the strict boundary.

Why it holds up over time. "Data whose sensitivity you cannot assess gets the strictest boundary you have" is a principle rather than a mechanism, and survives any change of key management. The arguable part is the opacity itself: a future schema-validated payload would legitimately reopen the decision.

Lesson. If your system carries data it does not interpret, it also cannot classify it — so give it the handling you would give the most sensitive thing it could be.

Every package used, in one table

These terms are used precisely in this package. Several are used loosely in the wider literature on schedulers and cron services, and the difference matters when reading the decision records.

Package What it is What it does here Considered instead
Scheduled instant The absolute UTC instant a recurrence resolves to, derived from the expression, the named zone and a recorded time-zone database version. The thing the fire is identified by, ordered by, and measured against. "Fire time", which is usually used for the moment a dispatch actually left — the quantity this package calls the dispatch instant.
Fire record An immutable row committed before any dispatch attempt, carrying the tenant, trigger, definition version, scheduled instant and a deterministic key. The scheduler's durable product and the system of record for "was this instant decided". A "job run", which implies work happened; a fire record says only that a decision was taken.
Idempotency key The fire's four-part primary key (tenant, trigger, scheduled instant, sequence), computed rather than generated. Makes a duplicate decision a uniqueness violation instead of something to detect, and gives the executor something to deduplicate on. A generated request identifier, which deduplicates transport retries but cannot deduplicate two independent decisions.
ε (clock uncertainty) The half-width of the interval a bounded time source reports around the current instant. The quantity the scheduler waits out, which is what converts clock skew into lateness rather than earliness. "Clock skew", which usually names the error itself rather than a reported bound on it.
Lateness Dispatch instant minus scheduled instant, measured per fire and published per trigger as a distribution. The budget the design spends; it is the measure of availability for the dispatch plane. "Latency", which in a request-shaped service means time to respond rather than distance from a deadline.
Missed-fire policy A per-trigger declaration — fire-all, fire-once-now or skip — of what to do with instants that should have fired and did not. Moves the post-outage decision from the incident to the definition. "Misfire instruction" in some scheduler libraries, usually with the same three options and no horizon.
Catch-up horizon The maximum age of a missed instant that may still be dispatched; beyond it the instant is recorded as expired. The platform's bound on how much work any interruption can accumulate, overriding the tenant's policy. A "grace period", which usually means how long a late fire is still considered on time.
Catch-up storm A backlog of missed instants released at once by a recovery, a resume or a backfill. The expected failure mode of a scheduler, and the thing the lane split and the rate cap exist for. A "thundering herd", which names simultaneous arrival generally rather than the self-inflicted recovery case.
Overlap policy A per-trigger declaration — allow, skip or queue — of what to do when a fire comes due while the previous execution has not reported terminal. The only control the scheduler has over concurrency in work it does not run. "Concurrency policy", which in some systems bounds parallel runs numerically rather than deciding what to do about one.
Unknown (terminal state) The recorded state of a fire whose outcome never arrived inside its outcome window, including a dispatch attempt that timed out. Prevents the scheduler withholding every later fire on a missing callback, at the price of occasional concurrent execution. "Failed", which asserts the work did not happen — a claim the platform cannot make about a timeout.
Backfill An authorised generation of fires for a declared past window, labelled as such and rate-capped in its own lane. The recovery path for work lost beyond the horizon, and a privilege separate from authorship. A "replay", which in event systems means re-delivering records that already exist rather than creating new ones.
Partition lease A short-TTL claim on a range of the trigger keyspace, recording the ε observed at grant. The unit of mutual exclusion in the timing plane — deliberately imperfect, because the fire key makes imperfection survivable. "Leader election", which implies one authority per service rather than one per partition.
Due index One row per live trigger holding its single next instant, partitioned and scanned in time order. The structure the scheduler reads; a rebuildable projection of the registry, not a system of record. A "timer wheel", which is an in-memory structure with the same purpose and no durability.
Oldest undispatched due age The age of the oldest instant that is due and has not yet been dispatched, measured per partition. The paging alert, because it is the only signal that distinguishes a healthy scheduler from a silent one. "Queue depth", which counts work waiting without saying how long the oldest item has been waiting.
The package

Everything as it was delivered.

These files are served exactly as they were produced — the diagram pages keep their own house style because that is the artifact, not a rendering of it.