Distributed Job Scheduler

Architecture Views

20 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

The scheduled-trigger service behind a developer platform's "every weekday at 09:00" box. Read it in seven acts: what sits inside the boundary, who it is for and what they get to do, how it is put together, what it stores, what happens when an instant comes due, how it is operated, and why it is safe. Every rate, latency and threshold here is a stated assumption, chosen to be arguable rather than measured — the package is a design, not a report on a running system.

Context and scope

What the scheduler is responsible for — deciding that an instant is due — and the much larger thing it is deliberately not responsible for, which is running the work.

People and journeys

Who declares a schedule, who gets paged when one does not fire, and the two moments where the architecture becomes visible to the person living it.
03 Tenant engineering Tenant developer 4,200 across 50 k tenants Goal — I want a digest to go out every weekday at 09:00 in my customers' time zone, and I do not want to learn anything about distributed timers to get it. Core journeys Ship a scheduled job see view 04 Amend a schedule safely Dry-run before committing Tenant on-call pages at 07:40 Goal — Something did not land overnight. Tell me whether it fired, whether you decided not to, or whether my endpoint refused it — before I start reading my own logs. Core journeys The morning after an outage see view 05 Pause a misbehaving trigger Platform Platform SRE 6 on rotation Goal — I need one number that tells me the scheduler is not firing, because a scheduler that is up and silent looks perfectly healthy on every other dashboard. Core journeys Drain a zone without losing a fire Throttle a tenant's catch-up Rebuild the due index in shadow Product owner sets defaults Goal — I want the default behaviour after an outage to be the one that is least likely to bill a customer twice. Core journeys Set a tenant's quota Approve a backfill Assurance Security reviewer quarterly Goal — Show me that a schedule cannot be pointed at a host the tenant does not own, and that the person who can backfill is not simply anyone who can edit a trigger. Core journeys Trace a dispatch credential Audit a horizon extension Machines in the cast Tenant target HTTP, queue, executor Goal — Hand me a fire with a key I can deduplicate on, and do not assume my answer means the work finished. Core journeys Accept a signed dispatch Report an outcome Time authority ε ≤ 10 ms Goal — Tell the truth about how wrong I might be, and be believed rather than averaged. Core journeys Serve a bounded interval Who the Scheduler Is For, and What They Get to Do Person or role Journey / task External / third party Security / platform Goals are in each actor's own voice. The two journeys with their own view are the ones where the architecture is visible to the person living it. v 1.0 · owner Platform Architecture · date 2026-10 Actors and Their Core Journeys Six parties, their goals in their own voice, and what each of them gets to do. HTML page SVG draw.io

Structure

The planes the service is built from, and the one thing they share: a strongly consistent state of record.

Data

Which stores are truth and which are rebuildable projections — the distinction that makes recovery a routine rather than a procedure.
11 tenant tenant_id PK quota_triggers quota_dispatch_per_s quota_inflight catchup_rate_pct trigger trigger_id PK tenant_id FK name UQ(tenant,name) state live|paused|deleted current_version FK trigger_version version_id PK trigger_id FK recurrence zone IANA target_id FK payload_ref encrypted missed_policy overlap_policy catchup_horizon_s attempt_budget target target_id PK tenant_id FK kind http|queue|exec endpoint verified_at verification_state audit_entry audit_id PK tenant_id FK actor action prior_value at due_index_row partition_id PK1 next_instant PK2 trigger_id FK version_id FK claimed_by nullable fire tenant_id PK1 trigger_id PK2 scheduled_instant PK3 sequence PK4 version_id FK origin sched|manual|backfill state non_dispatch_cause decided_at tzdata_version attempt attempt_id PK fire_key FK n dispatched_at transport_result http_status token_jti partition_lease partition_id PK owner_id expires_at epsilon_ms_at_grant work_outcome fire_key PK/FK terminal_state reported_at reported_by cause history_row row_key tenant#trigger#instant fire + attempts, projected ttl 90 d 1 : N 1 : N N : 1 1 : N 1 : 1 1 : N 1 : N N : 1 1 : 0..1 N : 1 Data Model — Trigger, Instant, Fire, Attempt The fire's four-part primary key is the idempotency key: a duplicate decision is a constraint violation, not a detection problem. Non-dispatch is a field on the fire, not a separate record. v 1.0 · owner Platform Architecture · date 2026-10 Data Model Eleven entities, and one four-part primary key that does the work of a deduplication service. HTML page SVG draw.io

Runtime

One instant from due to delivered, then the two paths that are hardest to get right: catch-up after an outage, and five classes of fire sharing one dispatch pipeline.

Operations

Where it runs, the one signal that catches a scheduler which is up and silent, and the loop that closes when a tenant can see its own lateness.

Assurance

The trust boundaries, the privilege split, and every failure class with its structural answer and its accepted residual risk.
20 Assumed to fail How it shows Structural answer Residual risk Clock Skew, step, lost sync ε above bound Shed leases, wait out ε Fleet-wide ε drift Owner loss Death between claim and send Lease expiry Re-derive from registry One duplicate attempt Split ownership Two owners mid-handover Duplicate-key rate Fire key uniqueness Executor must dedupe State unavailable Registry or index down Rising due age Fail closed on dispatch Stall, then catch-up Target Transient, 4xx, systematic Success by target class Budget, stop, isolate Tenant unaware of 4xx Poison definition No future instant Error on the trigger Quarantine one trigger Slow compiler path Catch-up storm Backlog released at once Backlog depth per tenant Horizon, cap, own lane Alert noisy while draining Slow executor Run outlives the interval Unknown-state fraction Overlap policy + window Concurrent run, or silence Zone or region Zone loss, region loss Degraded lateness p99 Reassign; declared promotion 15 s of decisions at RPO tzdata change Future instant moves Version on each fire Recomputed instant governs Surprise at the boundary Assurance — Failure Classes and Their Answers Every row's residual risk is accepted knowingly. The two that would change the design are fleet-wide clock drift and an executor that cannot deduplicate. v 1.0 · owner Platform Architecture · date 2026-10 Failure Classes and Their Answers Ten classes, each with what it looks like, what answers it structurally, and what is left over. HTML page SVG draw.io

Architecture One-Pager

The problem, the shape of the answer, the decisions that carry it, and what a prototype should prove.

The timing plane decides; the execution plane works; and the only thing crossing between them is an immutable, uniquely-keyed fire record that is committed before it is delivered.

Every developer platform ships a four-word scheduling box, and behind it is a problem that looks solved and is not. A multi-tenant timer fleet has to fire tens of millions of independent triggers at the right instant while a partition owner is being replaced, a node's clock is wrong, a zone is draining, and a tenant has just resumed four thousand triggers that slept through a day. The convenient implementation — fire from a node's system clock and record the result afterwards — fails in three ways that each look like success at the time. It fires early when a clock runs fast, producing work that reads a window of data which has not closed yet and reports a short answer. It loses fires that happened but were never recorded, which no amount of later reconciliation can recover because the record that was supposed to prove it is the thing that is missing. And when it comes back from an outage it releases everything it owes at once, turning a scheduling incident into an incident for every executor downstream. On top of those, it cannot answer the one question tenants actually ask — "did my job run last night?" — because a job that did not run produces no signal at all, in a service that is up, healthy on every dashboard, and silent.

Separate the plane that decides from the plane that works, and make the scheduler's durable output a decision rather than a result. A partition owner asks a time source for a bounded uncertainty interval, scans its share of a time-ordered due index for instants due at or before the lower bound of that interval — so it is deliberately ε late and never early — applies the trigger's declared overlap policy, and commits an immutable fire record whose primary key is derived deterministically from the tenant, the trigger and the scheduled instant. That key is the idempotency key, so two owners mid-handover compute the same one and the second insert loses to a uniqueness constraint instead of producing a second fire. Only then is anything dispatched, through one of three separate queues — on-time, catch-up, backfill — each with its own drain rate, so a backlog released by a recovery is rate-shaped rather than flooded and recovery is deliberately slower than the failure that caused it. Delivery is at-least-once with the key attached and a published residual duplicate rate, not an exactly-once claim nobody can verify. Everything derived — the due index, the caches, the fire history — is a rebuildable projection of a registry and a ledger that are the only systems of record, and the rebuild runs continuously in a shadow partition rather than waiting to be needed. What the tenant gets is a per-fire record with its scheduled, decided and dispatched instants and the cause of any non-dispatch; what the platform watches is the age of the oldest undispatched due instant, which is the only signal that can tell a working scheduler from a silent one.

What it is, and what it is not

A durable decision, committed before deliveryA timer that fires and then writes down what happened
Deliberately ε late, never early, at any percentileAs punctual as the node's clock allows, in both directions
At-least-once with a key and a published duplicate rateExactly-once delivery
Idempotency as a uniqueness constraint in the data modelA deduplication service on the dispatch path
Imperfect leases made survivable by the fire keyLeader election strong enough to be trusted alone
Post-outage behaviour declared per trigger beforehandAn operator decision taken during the recovery
Recovery rate-shaped slower than the failureDrain the backlog as fast as the pipeline allows
A scheduler that hands over workA workflow or DAG orchestrator
Fire history as a tenant-facing product surfaceA platform log that support can search

The decisions that are the architecture

01Commit before dispatch, always

The fire record is written before any attempt is made, so a dispatcher that dies mid-attempt loses an attempt and never a fire. Of the two orderings available, one failure is unrecoverable and the other is a retry; the ordering of two writes decides which one the system can have.

ADR-01

02The control plane cannot take the fire path down

Two planes, separately deployed and scaled, sharing only the state of record. Console traffic, a bad management deploy and an identity outage become latency incidents instead of missed fires. One write region, because two regions evaluating "is the previous run still going" produce two answers no key reconciles.

ADR-02

03Idempotency lives in the data model

The fire's primary key is (tenant, trigger, scheduled instant, sequence) — computed, not generated. Split ownership becomes a constraint violation rather than a detection problem, which is what lets leases be short, takeover fast, and leader election imperfect.

ADR-03

04Late is a budget; early is a bug

The node clock is untrusted. The scheduler takes a bounded uncertainty interval, waits it out, and scans for instants due at or before its lower bound — paying about 10 ms of deliberate lateness to make earliness impossible. A node whose uncertainty exceeds the ceiling sheds its leases rather than firing on a clock it cannot justify.

ADR-04

05Recurrences are recomputed, not materialised

One due-index row per trigger, holding one next instant, recomputed on advance against the named zone. An amendment, a pause, a validity change and a time-zone database update are all picked up with no invalidation sweep, because nothing was cached to invalidate — and dry-run is exact rather than indicative.

ADR-05

06The post-outage decision is declared before the outage

Every trigger declares its missed-fire policy from a closed set of three, defaulting to firing once for the most recent missed instant. A billing run and a cache warm have opposite correct answers and look identical to the platform; the only party who knows which is which is the tenant, and the only time they can say so is in advance.

ADR-06

07Wall-clock time means what the tenant meant

Expression plus IANA zone, resolved at computation, with declared rules for the day a time does not exist and the day it happens twice, and the time-zone database version recorded on every fire — so a disputed instant six months later is explainable from data rather than from memory.

ADR-07

08Every policy is bounded by a horizon

An instant older than the catch-up horizon is recorded as expired with its cause and never dispatched, whatever the tenant declared — the one place the platform overrules them. Caught-up fires carry their original scheduled instant, because the work reads a window defined by it.

ADR-08

09Recovery is slower than failure, by design

Three queues with independent drain rates and a per-tenant catch-up cap, plus a published shedding order: backfill, catch-up, over-quota, then on-time. A one-hour backlog takes about ten hours to drain, which is the price of not making the recovery into the next outage for every executor downstream.

ADR-09

10The peak is absorbed, not purchased

Sixty per cent of fires land in the first second of a minute. Deterministic opt-out jitter smearing and the lateness budget absorb a 45,000/s peak that lasts five seconds an hour, instead of provisioning eight times the steady-state capacity to be idle for the rest of it.

ADR-10

11Rebuildable means rebuilt on a schedule

The due index, the leases and every cache are projections of the registry and are not backed up. A shadow rebuilder recomputes and compares continuously, and partition ownership moves routinely rather than only during failure — because a recovery path that is only exercised in an incident is a claim, not a capability.

ADR-11

12`unknown` is a state, not an absence

Overlap is enforced against recorded outcome state bounded by a window. A fire still unknown at the end of its window counts as finished, and the decision is recorded. The alternative — treating it as still running — turns one lost callback into a permanently stopped trigger that looks perfectly healthy.

ADR-12

13Alert on the gap, not on health

The paging signal is oldest undispatched due instant age, because a scheduler that is up and not firing passes every other check. The same measurement, per fire and per trigger, is a tenant-facing product surface — so "did my job run" is answerable without reading platform logs or opening a ticket.

ADR-13

14A dispatch target is a privilege, not a parameter

Targets must be verifiably owned, checked at definition and again at dispatch; no tenant request reaches the fire zone; and credentials are minted per attempt with the attempt's lifetime. A service that will make authenticated requests to any URL on a timer, with a retry budget, is an attack tool until ownership is proved.

ADR-14

15The operations that multiply work are separate grants

Backfill, horizon extension and quota change are grantable apart from trigger authorship, and audited with prior values. Each of them is most attractive at the moment judgement is worst, which is during the incident — so the routine role's worst case stays routine.

ADR-15

16Opaque data gets the strictest boundary

The tenant payload exists in two stores, encrypted under a per-tenant key, and never in history, reporting, metrics or logs. The platform cannot classify what it does not interpret, so it handles the payload as the most sensitive thing it could be and debugging starts from a fire key.

ADR-16

Why this should still be right in ten years

Clouds, queue services, container runtimes and consistency models will all be replaced inside this platform's life. These are the properties that should outlast them.

The boundary is not a technology

"The timing plane's product is a committed decision" is a statement about where authority lives, not about which store holds it. Spanner, a successor nobody has shipped, or a different strongly consistent store can each satisfy it, and every decision downstream — the key as the deduplication mechanism, lanes as retries against records, history as a projection — survives the substitution unchanged.

Deterministic identity outlives every protocol

Naming a fire from its own semantics rather than generating an identifier for it is the oldest reliable trick in distributed systems. It depends on no store, no broker and no transport, and it is why imperfect leader election is affordable here. What could change is executors' willingness to deduplicate — which would call for an added suppression store, not a redesign.

The asymmetry between late and early is permanent

Scheduled work reads windows defined by its scheduled instant. That is a property of the work, not of the scheduler, and no improvement in clocks makes an early fire safe. If clocks get better, ε shrinks and the deliberate lateness disappears on its own; if they get worse, the lease-shedding ceiling catches it.

Recovery will always need to be slower than failure

A system that accumulates obligations while it is down and then discharges them into dependencies it does not control has a feedback loop into its own failure. That is control theory, not engineering fashion. The cap's value should move with measured executor capacity; the existence of a cap should not.

Silence will always need a gap measurement

Any system whose failure produces no output is invisible to health checks, forever. Alerting on the distance between what should have happened and what did is the only class of signal that sees it, and that remains true whatever the monitoring stack looks like.

The declared-policy habit transfers

Collecting, in advance, the information a system will need to make a decision it cannot make correctly during an incident is a general property of operable systems. Here it is the missed-fire policy; elsewhere it is a failover preference or a data-loss tolerance. The mechanism is configuration; the principle is foresight.

Non-functional targets

Every figure below is a stated assumption. They are listed with the mechanism that is supposed to deliver them and the view where that mechanism is drawn, so a reviewer can disagree with the number and the means separately.

QualityTargetHow it is metView
Fire punctuality p50 ≤ 1 s, p99 ≤ 5 s, p99.9 ≤ 30 s of lateness Due scan at 1 s tick on leased partitions, with the dispatch lane's drain rate sized for the mean rather than the peak 12
Earliness Zero at every percentile Scan for instants due at or before the lower bound of the clock's uncertainty interval; ε is waited out, not averaged 12
Degraded punctuality p99 ≤ 60 s for up to 10 minutes during zone loss or reassignment Partitions spread across three zones; reassignment exercised routinely rather than only on failure 15
Dispatch plane availability ≥ 99.95% monthly, measured as due instants dispatched inside their p99.9 budget Punctuality as the availability definition; stateless dispatch tier scaled independently of the timing tier 06
Control plane availability ≥ 99.9% monthly, failing independently of the fire path Separate deployment, scaling and service accounts, sharing only the state of record 06
Trigger population 20,000,000 active triggers across 50,000 tenants 256-partition keyspace with one due-index row per trigger; partition count changeable online 07
Steady-state throughput 500,000,000 fires/day — 5,800/s mean One transactional fire insert per fire; lane queues with declared drain rates 09
Peak throughput 45,000 fires/s for 5 s at the top of the hour Deterministic opt-out jitter smearing plus the lateness budget, instead of provisioned peak capacity 14
Clock uncertainty ε ≤ 10 ms; leases shed above 100 ms Bounded-uncertainty time source; per-node ε recorded at lease grant and alerted per node, not averaged 12
Duplicate dispatch ≤ 1 in 10^6 fires delivered more than once Deterministic fire key as the ledger primary key; executors deduplicate on it 11
Durability RPO 0 in-region; cross-region RPO ≤ 15 s Synchronous three-zone replication for registry and ledger; asynchronous standby replica 15
Recovery time RTO 5 min dispatch plane, 30 min control plane; partition reassignment ≤ 15 s p99 Declared region promotion; short leases made safe by the fire key; routine reassignment 15
Catch-up bound Horizon 1 h default / 24 h max; catch-up ≤ 10% of tenant quota Horizon filter ahead of the policy; separate catch-up queue with its own drain rate 13
Configuration propagation Create, amend, pause, resume, delete effective within 5 s p99 Recomputed due-index row written in the same transaction as the version 09
History 90 days queryable, 13 months archived; 30-day single-trigger query p95 ≤ 500 ms Bigtable projection keyed tenant#trigger#instant with TTL; archive-class object storage beyond it 10
Detection Oldest undispatched due instant > 60 s pages immediately Derived per-partition metric from the due index, alerted independently of process health 16
Cost ≤ $0.60 per million fires at steady state Mean-sized capacity, no per-trigger always-on resource, tiered history, retry and catch-up charged to the originating tenant 16

Scope

In scope

  • Trigger definition and lifecycle — calendar, interval and one-shot recurrences in named zones, with versioning, pause, resume, delete, manual run and exact dry-run
  • Time semantics: a bounded-uncertainty clock, declared daylight-saving rules, and a recorded time-zone database version per fire
  • A partitioned, time-ordered due index with lease-held ownership and clock-gated instant claiming
  • An append-only fire ledger whose primary key is the idempotency key, committed before any dispatch
  • At-least-once dispatch to an authenticated HTTP target, a tenant queue or an internal executor, with bounded jittered retry truncated at the next instant
  • Missed-fire policies, a platform-enforced catch-up horizon, and a bounded authorised backfill
  • Overlap policies with `unknown` as a terminal state, and per-trigger and per-tenant concurrency limits
  • Per-tenant quotas, fair dispatch queues, deterministic jitter smearing and a published shedding order
  • Tenant-facing per-fire history with causes, and the oldest-undispatched-due-age alert
  • Target ownership verification, per-attempt dispatch credentials, signed dispatch and verified callbacks, and the privilege split for backfill, horizon and quota

Explicitly out of scope

  • Executing the triggered work — the scheduler hands over a record and a credential and nothing more
  • Workflow and DAG orchestration, retries of business logic, and compensation; the separate distributed-workflow-orchestration-platform package covers that ground
  • Trigger-to-trigger dependencies, deferred to Phase 3 and only if they can be added without becoming an orchestrator
  • Sub-minute recurrences, which need a separately budgeted punctuality contract
  • Tenant-supplied recurrence predicates such as business calendars and market holidays, deferred to Phase 3 behind a sandboxed evaluation budget
  • Active-active multi-region scheduling, deferred because two authorities evaluating overlap produce answers no key reconciles
  • The product surfaces that render a schedule, and the tenant's own alerting on its work's outcome

What a four-week prototype should prove

The prototype's job is to falsify the two central claims — that imperfect leases plus a deterministic key really do make duplicate fires harmless, and that a bounded clock lets the scheduler be reliably late and never early — on one tenant and a few hundred triggers, not to build a platform.

  1. One tenant, 500 triggers across all three recurrence forms, in three time zones including one with daylight saving, with the two awkward days exercised deliberately rather than waited for.
  2. The timing plane on leased partitions against a strongly consistent store, with the fire key as the ledger's primary key and the clock gate reading a bounded uncertainty interval.
  3. Two dispatch lanes — on-time and catch-up — with independent drain rates and a per-tenant cap, against one deliberately slow HTTP target and one that times out without answering.
  4. The shadow rebuilder, running from day one, so the rebuild claim is under measurement for the whole four weeks rather than tested at the end.
  5. Fire history with scheduled, decided and dispatched instants and the cause of every non-dispatch, queryable by the tenant — because the prototype has to answer "did it run" to be worth anything.
  6. Measured cost per million fires on the prototype's shape, compared against the assumed $0.60, which is the assumption with the least evidence behind it.
  • A partition owner is killed mid-cycle, repeatedly, under load: every instant has exactly one fire record and at most one extra dispatch attempt, with the measured duplicate rate compared to the assumed 1 in 10^6
  • A lease handover happens while a scan is in flight: the second insert loses to the uniqueness constraint, and no application code resolves the conflict
  • A node's clock is stepped forward by two seconds: no fire goes out early, at any percentile
  • A node's reported ε is pushed past 100 ms: it sheds its leases and the partition is reassigned inside the 15-second budget, rather than firing on the bad clock
  • A 1-minute trigger is paused for three hours and resumed under each of the three missed-fire policies: the dispatch rate stays inside the 10% catch-up cap, and the horizon expires what it should with a cause the tenant can read
  • A dispatch is held open past its timeout with no callback: the fire is recorded unknown at the end of its window, the overlap decision is recorded, and the next fire is not withheld indefinitely
  • The due index is dropped entirely and rebuilt from the registry: the result matches what the shadow rebuilder had been reporting all along
  • A trigger is pointed at a domain the tenant has not verified: the definition is refused at write time, not at the first fire

Open risks, carried rather than hidden

RiskIf it landsResponse
An executor that cannot deduplicate on the fire key At-least-once stops being a contract and becomes a defect: a duplicate fire produces duplicate work, and the whole leases-may-be-imperfect argument collapses into a correctness problem Make deduplication an onboarding requirement with a conformance test rather than a documentation note; publish the duplicate rate on the target's own page; and keep a platform-side suppression store as a costed Phase 2 option for targets that genuinely cannot.
Fleet-wide clock drift, where ε is wrong everywhere at once rather than on one node Lease-shedding has nothing to compare against, the clock gate approves a bad interval, and the design's central guarantee — no early fires — fails silently and uniformly Alert on per-node ε rather than an aggregate, cross-check the time service against the store's commit timestamps as an independent source, and treat a correlated ε excursion as a platform incident with scheduling suspended rather than degraded.
The missed-fire default is wrong for most tenants, and they accepted it without reading it Either duplicated billable work or silently skipped work, discovered by a customer, with the platform's correct behaviour as the explanation — the trough in view 05 Instrument override rates by tenant segment and treat a high override rate as evidence the default is wrong; show the policy's consequence in a sentence at the point of choice; and make dry-run replay what each policy would have done to the last recorded gap.
The oldest-undispatched-due-age alert is loud during every legitimate catch-up drain The one signal that detects the system's characteristic failure becomes the one the on-call rotation learns to ignore, which is worse than not having it Compute the metric separately per lane so a catch-up drain does not mask an on-time stall, suppress it against a declared drain with an expected completion time, and page on the on-time lane alone.
Recomputing instants for 20 million triggers every cycle becomes the scaling wall Scan cost grows with the trigger population rather than with the fires actually due, and punctuality degrades as the population grows rather than as load does Measure compiler cost per trigger per cycle from day one as a first-class capacity metric; keep the short-horizon materialisation variant of ADR-05 costed and ready, since it is a localised change to the due index's contract and leaves the ledger, the key and the lanes untouched.

Distributed Job Scheduler — Architecture One-Pager and Decision Record

Why every component and every technology on these 20 views is what it is, and what each choice costs.

The architecture behind one small box in a developer platform's settings screen: "every weekday at 09:00". The one-pager is the argument in fifteen minutes; the decision record is the same argument with its working shown, sixteen decisions deep.

Status of this document. This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure is a stated assumption, chosen to be defensible and arguable rather than measured. The operating context assumed throughout is a developer platform with 50,000 tenants and 20 million active triggers, firing 500 million times a day at a mean of 5,800 fires a second, with 60% of all fires landing in the first second of a minute and the top of the hour carrying eight times the mean minute — a design peak of 45,000 fires a second for five seconds. Punctuality is assumed at p50 ≤ 1 s, p99 ≤ 5 s and p99.9 ≤ 30 s of lateness, with zero tolerance for earliness at any percentile; the clock uncertainty bound is assumed at ε ≤ 10 ms with leases shed above 100 ms; the residual duplicate dispatch rate is assumed at 1 in 10^6. Where a number came from nowhere, it is marked as an assumption in the view cards and in the records, and a reviewer should treat each of them as an invitation to disagree.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it is realised on AWSThe concrete mechanism: which service or package, configured how, in which subscription.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
Why it holds up over timeWhat keeps the decision right as scale, staff and technology change.
LessonThe principle that transfers beyond this platform.

Decision map

Decision and delivery 3

The boundary that defines the architecture, and the two decisions that follow from it directly.

ADR-01The scheduler's durable output is a committed fire record, written before any dispatch attempt ADR-02The control plane and the timing plane share only the state of record, and there is one write region ADR-03Idempotency lives in the data model: the fire's primary key is the idempotency key, and delivery is at-least-once

Time 3

What time it is, how wrong that might be, and what a wall-clock expression actually means.

ADR-04The node clock is untrusted; dispatch waits out a bounded uncertainty interval, so lateness is a budget and earliness is a bug ADR-05Recurrences are recomputed each cycle from a versioned definition, not materialised ahead into the due index ADR-07Wall-clock recurrences resolve against a named zone with declared DST rules, and the time-zone database version is recorded on every fire

Policy after the outage 4

What a trigger does about the fires it missed — declared before it misses them, not decided during the incident.

ADR-06The missed-fire policy is declared per trigger before the outage, and defaults to firing once for the most recent missed instant ADR-08A catch-up horizon bounds every missed-fire policy, and the original scheduled instant travels with the fire ADR-09Catch-up, backfill and on-time fires use separate queues with a per-tenant catch-up cap and a published shedding order ADR-12Overlap is enforced against recorded outcome state, with `unknown` as a first-class terminal state bounded by a window

Scale, cost and recovery 2

Absorbing the top-of-hour peak, and treating the due index as something rebuilt rather than restored.

ADR-10The top-of-hour peak is absorbed by opt-out jitter smearing and the lateness budget, not by provisioned capacity ADR-11The due index is a rebuildable projection, verified continuously by a shadow rebuild rather than backed up

Evidence and trust 4

What the tenant can see, what a schedule is allowed to point at, and who is allowed to multiply work.

ADR-13Fire history is a tenant-facing product surface, and the platform alert is oldest undispatched due age ADR-14A target must be verifiably owned by its tenant, no tenant request reaches the fire zone, and dispatch credentials are minted per attempt ADR-15Backfill, horizon extension and quota change are grants separate from trigger authorship ADR-16Tenant payloads live only in the registry and the ledger, encrypted under a per-tenant key, and never in evidence or logs

Technology by capability

Google Cloud was chosen for this exercise for two reasons that point the same way. The first is rotation: across this repository's use cases Microsoft Azure and self-hosted open source dominate, with Amazon Web Services next, and Google Cloud carries the smallest share — six of forty-four packages before this one. Reaching for the same cloud each time teaches a service catalogue rather than architecture. The second reason is that this particular topic genuinely belongs here. A scheduler's hardest requirement is a defensible answer to "what time is it, and how wrong could that be", and this is the platform that exposes a bounded clock-uncertainty interval on a transactional store rather than asking the design to assume one. ADR-04 is the decision that depends on it, and it would read as hand-waving on a stack that cannot report its own clock error. Everything else below is replaceable; the requirement document deliberately stays vendor-neutral so that it survives a change of cloud.

Open source This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
Trigger registry, due index, fire ledger Spanner regional instance, synchronous across three zones, commit timestamps as time authority Google Cloud CockroachDB or YugabyteDB self-hosted; DynamoDB with a separate time service; Azure Cosmos DB with strong consistency One store that enforces the fire key's uniqueness constraint and reports bounded clock uncertainty — the two properties ADR-01, ADR-03 and ADR-04 all depend on ADR-01
Bounded-uncertainty time Platform time service plus Spanner commit timestamps, with per-node ε recorded at lease grant Google Cloud AWS Time Sync with clock bound; a self-run PTP fabric; chrony with a measured error estimate The decision to wait out ε rather than fire at the earliest permitted instant needs a source that reports its own error, not one that is merely accurate ADR-04
Control, timing and dispatch planes Cloud Run services, separately deployed and scaled, with distinct service accounts Google Cloud GKE workloads; ECS or Fargate; Azure Container Apps Independent scaling and deployment is the mechanism of the plane split, and leases make long-lived instances unnecessary ADR-02
Dispatch lanes Three Cloud Tasks queues per region with independent rate and concurrency Google Cloud SQS with separate queues; RabbitMQ with per-queue limits; Azure Service Bus Per-queue dispatch-rate control is exactly the throttle the catch-up cap needs, and it is enforced by the queue rather than by application code ADR-09
Outcome events Pub/Sub for correlated outcome events from the platform executor Google Cloud Kafka; SNS and SQS; Event Hubs Decouples the executor's reporting from the dispatch path, so a slow reporter is not a slow dispatcher ADR-12
Fire history Bigtable projection keyed tenant#trigger#instant with a 90-day TTL Google Cloud DynamoDB with TTL; Cassandra; Azure Table Storage 500 million rows a day of append-mostly evidence with a prefix-scan read pattern and native expiry ADR-13
History archive Cloud Storage, archive class, 13 months, dual-region Google Cloud S3 Glacier Instant Retrieval; Azure Blob archive tier Retention without query cost; the archive is for obligations, not for answering questions ADR-13
Lateness and cost reporting BigQuery for lateness facts, shadow divergence and cost per million fires Google Cloud Snowflake; Redshift; ClickHouse self-hosted Keeps punctuality analysis entirely off the fire path, including during the incident it is being used to diagnose ADR-11
Dispatch identity Workload identity tokens minted per attempt, scoped to the trigger, verified by the target's issuer check Google Cloud SPIFFE/SPIRE with short-lived SVIDs; AWS IAM Roles Anywhere; Azure managed identities A credential whose lifetime equals an attempt reduces a standing grant to a seconds-long one ADR-14
Payload encryption Cloud KMS per-tenant keys, decrypted only in the dispatch path Google Cloud HashiCorp Vault transit; AWS KMS; Azure Key Vault with an HSM-backed key Per-tenant keys make deletion enforceable by key destruction rather than by trusting every copy to be found ADR-16
Perimeter and egress VPC Service Controls around the data stores; egress only from the attempt runner Google Cloud AWS PrivateLink and SCPs; Azure Private Endpoints with a firewall One component with egress is an egress policy small enough to audit ADR-14
Observability Cloud Monitoring for due-age and lateness, Cloud Logging with fire keys and no payloads Google Cloud Prometheus and Grafana; Datadog; Azure Monitor The alert is a derived measurement (oldest undispatched due age), so it needs a metric pipeline rather than a log search ADR-13
Audit Append-only audit table with a Cloud Storage Object Lock export, 7 years Google Cloud S3 Object Lock; Azure immutable blob storage; an on-prem WORM appliance The acts worth auditing are the ones that multiply work, and their record has to survive the actor who performed them ADR-15

The decisions, and the alternatives that lost

Decision and deliveryThe boundary that defines the architecture, and the two decisions that follow from it directly.

ADR-01

The scheduler's durable output is a committed fire record, written before any dispatch attempt

Accepted

When a dispatcher dies holding a fire it has decided is due, what has been lost?

Context
The convenient implementation of a scheduler dispatches first and records afterwards: the timer fires, the HTTP call goes out, and a row is written when the call returns. It is one less write on the hot path and it reads naturally. It also has no answer to the question above. A dispatcher that dies after the call and before the write leaves a fire that happened and is not recorded — which means the next scheduler to look at that trigger will fire it again, and nothing in the system can tell the difference between that and the first fire. The inverse ordering has the opposite failure: a fire recorded and not dispatched, which is a retry. One of those is unrecoverable and the other is routine, and the ordering of two writes decides which one the system has.
Decision
An immutable fire record — tenant, trigger, definition version, scheduled instant, deterministic idempotency key, attempt counter — is committed to a strongly consistent store before any dispatch attempt is made. The record, not the delivery, is the scheduler's durable product. Delivery is a retryable consequence of a record that already exists.
How it is realised on AWS
The fire ledger is a Spanner table whose primary key is (tenant_id, trigger_id, scheduled_instant, sequence). A partition owner that has decided an instant is due performs the insert inside the same transaction that advances the due index row, then enqueues onto a Cloud Tasks lane. The attempt runner reads the record, mints a per-attempt credential, signs the payload and dispatches; every attempt writes an attempt row against the same fire key. Nothing in the dispatch plane can create a fire.
Options weighed
  • ChosenCommit the fire record, then dispatch: Makes the unrecoverable failure impossible and the recoverable one routine; costs one transactional write per fire on the critical path
  • RejectedDispatch, then record the outcome: Cheaper and simpler; a crash between the two produces a fire that happened and cannot be proved, which is the one state the design cannot recover from
  • RejectedWrite an intent record best-effort, reconcile later: Reconciliation needs a source of truth to reconcile against, and the intent record was supposed to be it
  • RejectedRely on the queue's own durability as the record: A queue entry is not addressable by scheduled instant, cannot be queried by a tenant asking "did it run", and its retention is not the history retention
Consequences
What it buys
  • A dispatcher can die at any point in an attempt and the worst outcome is a duplicate attempt under a key the executor already knows
  • "Did it run" becomes a primary-key lookup rather than a log search, which is what makes fire history a product surface instead of a support tool
  • Retry, catch-up, backfill and manual run are all the same operation — attempts against a record — rather than four code paths
What it costs
  • One strongly consistent write per fire sits on the critical path, which at a 45,000/s design peak is the single largest capacity commitment in the system
  • A ledger outage stops fires entirely rather than degrading them, which is deliberate (ADR-02) and is the hardest consequence to explain to a tenant
  • The ledger grows at 500 million rows a day and needs a retention and partitioning story from day one rather than later
Choose differently when
If the work being triggered were idempotent by nature and cheap to repeat — a cache warm, a metrics scrape, a health probe — the unrecoverable failure would not matter and dispatch-then-record would be the right trade. The decision is justified by schedules that send money, email and statements, where a fire that cannot be proved is worse than a fire that is late.
Why it holds up over time
"The decision is the product" is a statement about where authority lives, not about which store holds it. Spanner, a successor nobody has shipped, or a different consistent store can each satisfy it, and everything downstream — the key as the deduplication mechanism, lanes as retries against records, history as a projection — survives the substitution unchanged. What would not survive is a move to an eventually consistent store for the ledger, because the uniqueness constraint is the whole mechanism.
LessonIn any system that acts on the world on a timer, decide which of your two unavoidable failures is the recoverable one, and then order your writes so you only ever get that one.
Shown on views02 07 12
ADR-02

The control plane and the timing plane share only the state of record, and there is one write region

Accepted

If the API that creates schedules is down, should schedules still fire?

Context
A scheduler has two obviously different jobs. One is interactive, bursty, request-shaped and has a human waiting: create, amend, pause, dry-run, read history. The other is autonomous, continuous, deadline-driven and has nobody waiting: decide which instants are due and hand them over. Building them as one service is the default, and it means a deploy, a bad query, a traffic spike on the console or an authorisation dependency outage takes the fire path with it — the plane with a deadline is held hostage by the plane with a user. The same question repeats across regions: a second region that can also decide instants doubles write availability and introduces two authorities deciding the same overlap question, which the fire key deduplicates only after both have already decided.
Decision
The control plane and the timing plane are separately deployed, separately scaled and separately available, sharing nothing but the strongly consistent state of record. A control-plane outage stops new and amended definitions; it does not stop scheduled fires. Scheduling authority lives in exactly one write region, with a standby region that replicates state and holds no partition leases.
How it is realised on AWS
Both planes run as independent Cloud Run services with separate service accounts, deployment pipelines and scaling policies, against one Spanner instance holding registry, due index and ledger. Partition leases are rows in that instance, so only a process with write access to the primary region can hold one. The standby region runs the same images scaled to zero with a read replica; promotion is a declared operation that moves the lease table's authority, not an automatic failover.
Options weighed
  • ChosenSeparate planes, one write region, cold standby: Fire path survives control-plane failure; region loss is a declared promotion against a stated RPO
  • RejectedOne service for both planes: Simplest to operate and deploy; makes every console incident a punctuality incident
  • RejectedActive-active scheduling across two regions: Best availability; two authorities decide overlap and concurrency independently, and the fire key cannot undo a decision already taken
  • RejectedSeparate planes with the control plane in a second region: Attractive for console availability; a definition written in region B that region A has not yet seen is a fire that silently did not happen
Consequences
What it buys
  • Console traffic, a bad management deploy and an identity-provider outage are all latency incidents rather than missed fires
  • The timing plane's dependency list is short enough to reason about: the state of record and the time authority
  • Overlap and concurrency decisions have exactly one authority, which is what makes the `skip` and `queue` policies in ADR-12 meaningful
What it costs
  • Two deployments, two scaling policies and two on-call surfaces for one product
  • Region loss costs up to the replication lag in decisions and a declared promotion rather than a transparent failover
  • Control-plane availability is lower than the dispatch plane's, and tenants will experience that as the service being down
Choose differently when
If the scheduler served one tenant with a hundred triggers, the operational cost of two planes would dominate the availability it buys and one service would be right. At 50,000 tenants, the console is busy enough that treating its incidents as fire-path incidents is indefensible.
Why it holds up over time
The plane split is an availability-domain statement and outlives any runtime. Active-active scheduling stays rejected for a structural reason rather than a technical one: the fire key makes a duplicate *fire* harmless, but two regions independently evaluating "is the previous execution still running" produce two different answers, and no key reconciles those. That reasoning does not change if replication gets faster.
LessonSplit a system where its deadlines differ, not where its nouns differ. A plane with a user and a plane with a deadline should never be able to take each other down.
Shown on views06 07 15
ADR-03

Idempotency lives in the data model: the fire's primary key is the idempotency key, and delivery is at-least-once

Accepted

Two schedulers both believe the same instant is due. What stops two fires?

Context
Every distributed scheduler eventually has two processes that believe they own the same trigger: a lease handover, a network partition, a paused process resuming, a clock disagreement. The usual responses are to make that impossible — a stronger lease, a consensus round, a global leader — or to detect it afterwards with a deduplication store keyed by some generated identifier. The first is expensive and still probabilistic, because a lease is a promise about time held by a process that may not know the time. The second adds a component on the critical path whose own availability becomes the fire path's availability, and whose retention window becomes the deduplication guarantee. Both are machinery built to compensate for an identifier that was chosen badly.
Decision
The fire's identity is derived deterministically from (tenant, trigger, scheduled instant, sequence within instant) and is the primary key of the fire ledger. Two processes that both believe an instant is due compute the same key, so the second insert fails on a uniqueness constraint. Delivery to the executor is at-least-once, the key travels with every dispatch, and the residual duplicate rate is published as a contract the executor must tolerate rather than concealed behind an exactly-once claim.
How it is realised on AWS
The ledger's Spanner primary key is the four-part tuple; the insert is a plain `INSERT` whose `ALREADY_EXISTS` is handled as a normal, expected return meaning "another owner decided this instant". The same key is sent in the dispatch payload and in a header, so an HTTP target, a queue consumer and the internal executor all deduplicate on the same value. Attempts are child rows under the fire key, so a retry can never create a second fire.
Options weighed
  • ChosenDeterministic key as the ledger primary key: Turns split ownership into a constraint violation; needs no extra component and no retention window
  • RejectedGenerated UUID per fire with a deduplication store: Works for transport retries; cannot deduplicate two independent decisions, because they generate different identifiers
  • RejectedConsensus round before every fire: Strongest agreement; adds a round trip to a 45,000/s peak and still leaves the clock question open
  • RejectedClaim exactly-once delivery via transactional outbox to the executor: Honest only where the platform owns the executor, which is one of four target forms
Consequences
What it buys
  • Leader election is allowed to be imperfect, which means leases can be short and takeover fast (15 s) without risking correctness
  • No deduplication component sits on the fire path, so there is one less thing whose outage is an outage
  • The duplicate rate is a published number a tenant can design against rather than a property they discover
What it costs
  • Every executor must deduplicate. A tenant target that cannot is a correctness gap the platform cannot close for them
  • The sequence component of the key has to be decided by whoever decides an instant produces more than one fire, which is a subtlety in the catch-up path
  • A monotonic key means hot-spotting on the ledger's key range at the top of the hour, mitigated by the tenant prefix leading the key
Choose differently when
If the platform owned every executor, a transactional handoff would make exactly-once honest and the published duplicate rate unnecessary. With tenant HTTP targets in the mix it is not achievable, and claiming it would be the more dangerous choice.
Why it holds up over time
Deterministic identity from the semantics of the event is the oldest reliable trick in distributed systems and does not depend on a store, a broker or a protocol. What could change is the willingness of executors to deduplicate; if that becomes unrealistic, the decision does not need replacing so much as supplementing with a platform-side suppression store — an addition, not a rewrite.
LessonBefore building a deduplication service, ask whether the thing being deduplicated could have been given a name that makes duplicates impossible to express.
Shown on views11 12 20

TimeWhat time it is, how wrong that might be, and what a wall-clock expression actually means.

ADR-04

The node clock is untrusted; dispatch waits out a bounded uncertainty interval, so lateness is a budget and earliness is a bug

Accepted

A node believes it is 09:00:00. How wrong could it be, and which direction of wrongness is acceptable?

Context
A scheduler is a program whose entire output depends on the one value that distributed systems are worst at agreeing on. NTP-synchronised clocks in a datacentre are usually within a few milliseconds and occasionally wrong by seconds — a lost sync, a step correction, a virtualised clock after a live migration. The default implementation reads the system clock and fires when it passes the scheduled instant, which means a node whose clock runs fast fires early. Early is the dangerous direction: scheduled work overwhelmingly reads a window of data that is defined by the scheduled instant, so a job that fires before its instant reads a window that has not closed, produces a short report, and looks like it succeeded. Late work reads the right window and arrives inconveniently. These two failures are not symmetrical and should not share a budget.
Decision
Dispatch decisions are taken only against a time source that reports a bounded uncertainty interval, and the scheduler waits out that interval rather than firing at the earliest instant it permits — scanning for instants due at or before t−ε. Lateness is measured, bounded and published; earliness has no acceptable rate at any percentile. A node whose reported uncertainty exceeds a declared ceiling sheds its partition leases and continues serving reads rather than firing on a clock it cannot justify.
How it is realised on AWS
The timing plane takes time from Spanner commit timestamps and the platform's bounded-uncertainty time service rather than from `clock_gettime`. The due scan selects instants `<= now_lower_bound`, so the scheduler is deliberately ε late. Each partition owner records the ε it observed at lease grant on the lease row; above 100 ms it releases the lease and the lease manager reassigns the partition. Clock ε is a per-node metric, not an aggregate, because a fleet-wide drift is invisible in an average.
Options weighed
  • ChosenBounded-uncertainty source, wait out ε, shed leases above a ceiling: Guarantees no earliness for about 10 ms of deliberate lateness; needs a time service that reports its own error
  • RejectedTrust the NTP-synchronised system clock: Free and usually fine; produces early fires exactly when a node is unhealthy, and no node can tell that it is the wrong one
  • RejectedTake time only from the store's commit timestamp: Close to chosen and simpler; a read per scan cycle is acceptable but leaves no per-node signal with which to shed a bad clock
  • RejectedFire at the midpoint of the uncertainty interval: Halves the deliberate lateness and makes earliness a rate rather than an impossibility
Consequences
What it buys
  • No fire is ever early, at any percentile, which removes an entire class of silent wrong answers from the work the scheduler triggers
  • A bad clock is a detectable local condition with a local response, rather than a fleet-wide correctness problem
  • Short leases become safe, which is what makes a 15 s takeover budget achievable (ADR-11)
What it costs
  • Every fire pays ε of deliberate lateness — cheap at 10 ms, and the decision would not survive a bound of 500 ms
  • The design depends on a platform capability (a clock that reports its own error) that is not available everywhere, which constrains portability
  • A node sheds leases on a clock problem, so a correlated clock event reduces scheduling capacity exactly when nothing is wrong with the schedulers
Choose differently when
If the scheduler only triggered work whose correctness did not depend on its instant — notifications, cache warms, retries — earliness would be harmless and the simplest clock would do. The decision is justified by work that reads a window the scheduled instant defines.
Why it holds up over time
The rule survives every change of clock technology, because it is a statement about which direction of error is acceptable rather than about how accurate the clock is. If clocks get better, ε shrinks and the deliberate lateness disappears; if they get worse, the ceiling catches it. What would break the decision is a platform with no bounded-error time source at all, which would force the weaker commit-timestamp variant above.
LessonWhen a system's output depends on a measurement, ask which direction of measurement error is survivable, and spend the error budget entirely in that direction.
Shown on views06 12 20
ADR-05

Recurrences are recomputed each cycle from a versioned definition, not materialised ahead into the due index

Accepted

Should the next thousand fire instants be computed now and stored, or computed each time they are needed?

Context
A calendar recurrence is a function of an expression, a time zone, a time-zone database and a validity window. Materialising its instants weeks ahead makes the due index trivially cheap to scan: it is already a sorted list of absolute times. It also creates a cache that four separate events invalidate — an amendment, a pause, a time-zone database update, and a change to the validity window — and the invalidation has to find every materialised instant for one trigger among many millions. Recomputing instead costs CPU on every scan cycle, proportional to the trigger population rather than to the instants actually due, and it makes the correct answer the only answer the system can give.
Decision
The due index holds exactly one row per live trigger, carrying its single next instant, recomputed and advanced when that instant is claimed. The recurrence is evaluated against the trigger's named zone at the moment of computation. A definition is immutable once fired against; an amendment creates a new version and the index row is recomputed under it, so the instant the scheduler uses is always derived from the rules currently in force.
How it is realised on AWS
The recurrence compiler is a pure function in the control plane and the timing plane, shared as one library and exercised by dry-run on the production path, so the next three instants shown to a developer are the instants the scheduler will use. The due index row carries trigger_id, version_id and next_instant; claiming an instant and advancing the row happen in the transaction that inserts the fire (ADR-01). An amendment writes a new trigger_version and recomputes the row within the 5-second propagation budget.
Options weighed
  • ChosenOne next instant per trigger, recomputed on advance: Always correct under the current rules, with no invalidation to get wrong; costs a compiler call per fire and per amendment
  • RejectedMaterialise instants weeks ahead: Cheapest possible scan; four separate invalidation paths, each of which silently produces a fire at an instant the current rules disagree with
  • RejectedMaterialise a short horizon, say one hour: A real middle ground and the most likely future revision; still needs invalidation, and buys little while the scan is per-partition
  • RejectedIn-memory timer wheels per partition: Lowest dispatch latency; makes the authoritative structure non-durable, so a restart re-derives it anyway
Consequences
What it buys
  • A time-zone database update, an amendment and a pause are all picked up with no invalidation sweep, because nothing was stored to invalidate
  • Dry-run is exact rather than indicative, which is what makes it useful at the moment a developer commits a schedule
  • The due index stays small — one row per trigger — so a partition scan is proportional to what is due rather than to history
What it costs
  • Recurrence evaluation sits on the scan path, so a slow or pathological expression is a scheduling cost rather than a one-off authoring cost
  • A trigger population of 20 million means 20 million index rows kept current, which is a write rate proportional to the fire rate
  • There is no cheap way to answer "show me every fire in the estate next Tuesday", which a materialised index would give for free
Choose differently when
If recurrence evaluation were expensive — tenant-supplied predicates against business calendars, say, as the ask defers to Phase 3 — the balance would move towards a short materialised horizon with explicit invalidation. For three fixed recurrence forms it does not.
Why it holds up over time
The decision is really about where correctness lives: in a recomputation, or in a cache plus its invalidation. That trade does not change with hardware. If evaluation cost ever dominates, the short-horizon variant is a localised change to the due index's contract and leaves the ledger, the key and the lanes untouched.
LessonA cache of derived values is a commitment to invalidate it correctly on every input that feeds it. Count those inputs before you build the cache.
Shown on views09 11 17
ADR-07

Wall-clock recurrences resolve against a named zone with declared DST rules, and the time-zone database version is recorded on every fire

Accepted

What does "every day at 02:30" mean on the day 02:30 happens twice, and on the day it does not happen at all?

Context
A calendar recurrence in a named zone has two days a year where it is ambiguous or impossible, and a scheduler must have an answer for both. Worse, the mapping from wall-clock to absolute instant is not fixed: time-zone databases are updated several times a year, and a government changing its rules moves a future instant that the scheduler may already have told a tenant about. Storing instants in UTC and forgetting the zone makes the system immune to this and also wrong — a tenant who asked for 09:00 local means 09:00 local after the rules change, not 08:00. Any system that does not record which rules it used cannot explain, six months later, why a fire landed where it did.
Decision
Calendar recurrences store the wall-clock expression plus the IANA zone and resolve against the zone at the moment of computation. A wall-clock time that does not exist on a given day fires at the instant the clock jumps to; a wall-clock time that occurs twice fires on the first occurrence only. Both outcomes are visible in history as such. The time-zone database version used is recorded as a column on every fire record.
How it is realised on AWS
The recurrence compiler pins a tzdata version per deployment and writes it to the trigger_version row and to every fire. The two DST rules are implemented once in the shared compiler, so dry-run shows the same substitute instant production will use. A tzdata upgrade is a deliberate deployment that re-resolves future instants; the version on past fires explains any instant a tenant disputes.
Options weighed
  • ChosenExpression + zone, resolved at computation, version recorded: Means what the tenant asked; the recorded version is what makes a disputed instant explainable
  • RejectedConvert to UTC at authoring time: Immune to rule changes and simple; silently wrong for the tenant twice a year and permanently wrong after a rule change
  • RejectedFire on both occurrences of an ambiguous time: Defensible for monitoring work; produces a double billing run once a year
  • RejectedSkip a non-existent wall-clock time entirely: Honest, and means a daily job silently does not run one day a year — the failure nobody notices until the month-end total is short
Consequences
What it buys
  • The tenant's expression means what they think it means, including after their government changes the rules
  • The two hard days a year have a declared, testable, documented behaviour rather than whatever the library did
  • A disputed instant is explainable from data, because the rules in force are recorded alongside the fire
What it costs
  • A tzdata upgrade is a change to future behaviour and needs treating as a deploy with a review, not a dependency bump
  • Zone resolution on every computation adds cost to the scan path (ADR-05 already pays for recomputation)
  • The "first occurrence only" rule will surprise someone whose job genuinely wanted both
Choose differently when
If every tenant were in one zone with no daylight saving, the whole decision would collapse into storing UTC. It is justified by a multi-tenant platform whose tenants' customers are the ones whose local morning matters.
Why it holds up over time
Time-zone rules will keep changing and the IANA database will keep being the way to track them. Recording the version used is a general principle — any derived value computed from an external ruleset should carry the ruleset's version — and it does not depend on anything in this stack.
LessonIf a value is computed from rules that someone else can change, store the version of the rules alongside the value, or you will one day be unable to explain your own output.
Shown on views09 11 20

Policy after the outageWhat a trigger does about the fires it missed — declared before it misses them, not decided during the incident.

ADR-06

The missed-fire policy is declared per trigger before the outage, and defaults to firing once for the most recent missed instant

Accepted

Who decides what a trigger does about the fires it missed, and when is that decision taken?

Context
After any interruption — a platform incident, a long pause, a trigger resumed after a weekend — there is a set of instants that should have fired and did not, and three defensible things to do with them. Fire all of them in order, fire once for the most recent, or record them as skipped and fire nothing. Which is correct depends entirely on what the work means, and the platform cannot know that: a billing run and a cache warm have opposite correct answers and look identical from the scheduler's side. The alternative to the tenant declaring it is an operator deciding during the recovery, across 50,000 tenants at once, under pressure, with no information about any of them — a decision that is guaranteed to be wrong for most of the estate.
Decision
Every trigger declares its missed-fire policy at definition time from a closed set — `fire-all`, `fire-once-now`, `skip` — and the platform default is `fire-once-now`. The policy is part of the versioned definition, so the decision exists before the outage that needs it, and the recovery executes a declaration rather than improvising one.
How it is realised on AWS
The policy is a column on trigger_version, surfaced on the authoring screen as the one distributed-systems question a developer is asked, with each option's consequence stated in a sentence. The catch-up path in view 13 applies the horizon filter first (ADR-08) and the policy second. Dry-run shows what a policy change would have done to the last recorded gap, so a tenant can change it with evidence.
Options weighed
  • ChosenPer-trigger declaration, default fire-once-now: Puts the decision with the only party who knows what the work means; costs one question at authoring time and a default that will be wrong for some
  • RejectedPlatform-wide policy: Nothing to configure and nothing to get wrong; guarantees the wrong behaviour for a large share of the estate
  • RejectedOperator decision during recovery: Maximum flexibility exactly when there is no capacity to use it, and no information to use it with
  • RejectedInfer from the trigger's recent behaviour: Plausible and seductive; a scheduler guessing whether work is idempotent is a scheduler that will be wrong about money
Consequences
What it buys
  • Recovery is deterministic and rehearsable: the behaviour after an outage is a property of the definition, not of who was on call
  • The default is the one least likely to duplicate billable work, which is the failure a tenant will escalate
  • A tenant who needs `fire-all` declares it, and the declaration is visible to reviewers in their infrastructure code
What it costs
  • Every tenant is asked a question most of them would rather not answer, at the moment they least want friction
  • A tenant who accepted the default and wanted `fire-all` loses work they expected; the cause is recorded but the work is gone
  • Three policies × the horizon × the overlap policy is a behaviour matrix that documentation has to make legible
Choose differently when
If the platform owned the work as well as the timer — a managed job runner where every job declared its own idempotency — the platform could infer the policy and the question would be unnecessary. With opaque tenant payloads behind opaque tenant endpoints, it cannot.
Why it holds up over time
"Declare the recovery behaviour before the failure" is a general property of operable systems and does not depend on this stack. The specific default might move: if telemetry shows most tenants override it, the default is wrong and should change. The closed set of three is the part worth defending, because an open-ended policy language here would be a scripting surface on the fire path.
LessonIf a system will have to make a decision during an incident that it cannot make correctly without information it does not have, collect that information before the incident and call it configuration.
Shown on views04 13 17
ADR-08

A catch-up horizon bounds every missed-fire policy, and the original scheduled instant travels with the fire

Accepted

A trigger is resumed after a month. How much work does it owe?

Context
The missed-fire policy (ADR-06) answers what to do with missed instants but not how many of them there can be. A minute-schedule paused for a month and resumed under `fire-all` owes 43,200 fires; the tenant who clicked resume expected a job to start running again, not a month of history to arrive. The same unbounded set appears after a long incident and after an authorised backfill. Separately, work triggered late almost always reads a window of data defined by its scheduled instant: a statement run for 1 March that executes on 3 March must still report February. A fire that only knows when it was dispatched cannot do that.
Decision
Every trigger carries a catch-up horizon — default 1 hour, maximum 24 hours — and an instant older than the horizon is recorded as expired with its cause and never dispatched, whatever the missed-fire policy says. This is the one place the platform overrules the tenant's declaration. Every caught-up fire carries its original scheduled instant, distinct from its dispatch instant, and both are visible in history.
How it is realised on AWS
The horizon filter runs before the policy in the catch-up path, writing a fire row in `expired` state with a cause so the tenant can see the decision rather than a gap. The fire's `scheduled_instant` is part of its primary key and is what the dispatch payload carries; `dispatched_at` lives on the attempt row. Horizon extension beyond the default is a privileged operation (ADR-15).
Options weighed
  • ChosenPlatform-enforced horizon, original instant preserved: Bounds the worst case absolutely; costs a tenant some work they may have wanted
  • RejectedUnbounded catch-up under the tenant's policy: Most faithful to the declaration; one resume becomes a self-inflicted denial of service against the tenant's own endpoint
  • RejectedHorizon as a soft limit with a warning: Keeps the tenant in control; a warning nobody reads is an unbounded catch-up with extra steps
  • RejectedDispatch with the dispatch instant only: Simpler payload; makes every late fire produce work about the wrong window
Consequences
What it buys
  • The worst case after any interruption is bounded by a number the tenant can see and reason about
  • Late work is still correct work, because the window it reads is defined by the instant it was scheduled for
  • An expired instant is a visible decision with a cause rather than a silent absence, which is what makes view 05's journey recoverable
What it costs
  • A tenant who genuinely needed a month of catch-up must request a backfill, which is a privileged, approved operation
  • The horizon is a default and defaults are not read; the first time it matters is during an incident (view 05's trough)
  • Two timestamps on every fire is a payload and documentation cost that tenants will initially get wrong
Choose differently when
If the triggered work were always idempotent and cheap, an unbounded catch-up would be harmless and the horizon would be needless friction. The horizon is justified by work that costs money or sends mail every time it runs.
Why it holds up over time
A bound on the recovery set is a general requirement of any system that accumulates obligations while it is down, and the principle survives any implementation. The specific defaults are the arguable part and should move with observed data — which is why they are stated as assumptions rather than constants.
LessonAny queue that fills while a system is down needs a stated maximum age, decided before the outage. Without one, recovery is unbounded by construction.
Shown on views05 13 14
ADR-09

Catch-up, backfill and on-time fires use separate queues with a per-tenant catch-up cap and a published shedding order

Accepted

A backlog is released. What stops its drain from being the next outage?

Context
Recovery is the most dangerous moment in a scheduler's operation, and it looks like success while it happens. Fires that have been accumulating are suddenly all due, and the natural behaviour — dispatch them as fast as the pipeline allows — turns a scheduling incident into an incident for every executor downstream, including the ones that had nothing to do with it. A single shared queue with priorities does not help as much as it appears to: priorities decide ordering, but a saturated pipeline is saturated for everyone, and the lane that should have been throttled is competing with the lane that should not have been.
Decision
On-time, catch-up and backfill fires are dispatched through separate queues with independent drain rates. A tenant's catch-up dispatch is capped at a fraction of its steady-state quota — assumed at 10% — so a backlog drains over hours rather than minutes. The shedding order is published: backfill first, then catch-up, then over-quota steady-state traffic, and only then on-time fires. Recovery is deliberately slower than the failure that caused it.
How it is realised on AWS
Three Cloud Tasks queues per region, each with its own dispatch rate and concurrency, fed by the rate shaper. Catch-up drains oldest-first within a trigger and interleaves fairly across a tenant's triggers, so no single trigger's backlog starves the rest. A retry stays in the lane its fire came from, and its attempt budget is truncated at the next scheduled instant. A tenant whose targets are systematically failing has its retry volume charged to its own quota.
Options weighed
  • ChosenSeparate queues per class, per-tenant catch-up cap, published shedding order: Makes the throttle structural rather than a tuning value; costs three queues to operate and a long drain the tenant must accept
  • RejectedOne queue with priorities: Fewer moving parts; under saturation the priority decides order, not whether the backlog is throttled
  • RejectedUnthrottled drain with autoscaling: Fastest recovery for the platform; moves the incident to every tenant executor at once
  • RejectedWithhold the decision so the backlog never materialises: A real alternative (Question 5 in the ask); keeps the ledger smaller but makes the backlog invisible and un-queryable
Consequences
What it buys
  • An outage's recovery cannot become a second outage, and the limit is a published number rather than an operator's judgement
  • The shedding order means degradation is a decision taken in advance, which is what lets an SRE predict what will happen under load
  • A tenant with a failing endpoint pays for its own retries instead of consuming shared dispatch capacity
What it costs
  • A one-hour backlog takes roughly ten hours to drain, which is a long time to tell a tenant their work is still coming
  • The oldest-undispatched-due-age alert is loud for the whole drain, because the backlog is deliberately visible
  • Three queues, three sets of capacity and three sets of saturation behaviour to operate and to reason about
Choose differently when
If every executor were elastic and owned by the platform, an unthrottled drain with autoscaling would recover faster at no external cost. With opaque tenant endpoints of unknown capacity, the conservative drain is the only responsible default.
Why it holds up over time
"Recovery slower than failure" is a control-theory statement about a system with a feedback loop into its own dependencies, and holds regardless of queue technology. The cap's value is the arguable part; the existence of a cap and of a published order is not.
LessonSize your recovery path for the capacity of whoever has to absorb it, not for the capacity of the system doing the recovering.
Shown on views07 13 14
ADR-12

Overlap is enforced against recorded outcome state, with `unknown` as a first-class terminal state bounded by a window

Accepted

The previous execution has not reported back. Is it still running?

Context
A trigger whose work sometimes takes longer than its interval will eventually be due again while the previous run is still going, and three behaviours are defensible: run anyway, skip, or queue one. Enforcing `skip` or `queue` requires the scheduler to know whether the previous execution has finished — which means depending on a signal it does not control, from an executor it does not own. When that signal never arrives, there are exactly two choices and both are bad. Treat the missing callback as finished and risk genuine concurrent execution; treat it as still running and withhold every subsequent fire indefinitely, which is a silent outage that looks like a working trigger with no fires.
Decision
Overlap policy is declared per trigger (`allow`, `skip`, `queue`) and enforced against recorded outcome state, bounded by the trigger's outcome window. A previous fire still in `unknown` state at the end of its window is treated as finished for overlap purposes, and that decision is recorded. `unknown` is a first-class terminal state, not an absence. A dispatch attempt that times out is recorded as unknown rather than failed, because the target may have accepted it. Under `queue`, at most one pending fire is held; a newer instant discards and records the older one.
How it is realised on AWS
`work_outcome` is a 1:0..1 child of the fire; the absent row *is* the unknown state. The timing plane's overlap check reads the previous fire's state and its age against the outcome window before claiming the next instant, and writes the resulting skip with a cause when it declines. The outcome window is a column on the trigger version.
Options weighed
  • ChosenOutcome state with a bounded window, unknown treated as finished: Bounds the silent-outage failure, which is the one a tenant cannot see; accepts occasional concurrent execution
  • RejectedTreat unknown as still running: Never allows concurrency; a single lost callback stops a trigger forever and looks healthy while doing it
  • RejectedDelegate overlap entirely to the executor: Architecturally cleanest — the scheduler stays stateless about executions; leaves every tenant to build the same lock
  • RejectedHold a lease for the execution's duration: Strongest mutual exclusion; makes the scheduler's lease lifetime depend on tenant work it cannot bound
Consequences
What it buys
  • A lost callback costs at most one window of uncertainty rather than a permanently stopped trigger
  • Every overlap decision is recorded with a cause, so a tenant sees "skipped: previous still running" instead of a gap
  • A timed-out attempt is honestly recorded as unknown, which keeps the duplicate-tolerance contract (ADR-03) coherent
What it costs
  • The scheduler is now stateful about executions it does not run, which is a dependency on a signal it cannot guarantee
  • Under a systematically missing callback, `skip` degrades towards `allow`, and the tenant may not notice the degradation
  • The outcome window is another per-trigger value to choose badly
Choose differently when
If every executor reliably reported terminal state — a platform-owned runner, not a tenant endpoint — `unknown` would be rare enough that treating it as still running would be the safer default. With tenant HTTP targets it is not.
Why it holds up over time
The forced choice between risking concurrency and risking silence is structural: it exists whenever one system's decision depends on another system's unreliable report, and no technology removes it. What a future change could do is shrink the unknown population, which would make the window shorter without changing the rule.
LessonWhen a decision depends on a signal that may never arrive, name the missing signal as a state, bound how long you will wait for it, and write down which of the two bad outcomes you chose.
Shown on views11 14 17

Scale, cost and recoveryAbsorbing the top-of-hour peak, and treating the due index as something rebuilt rather than restored.

ADR-10

The top-of-hour peak is absorbed by opt-out jitter smearing and the lateness budget, not by provisioned capacity

Accepted

Sixty per cent of all fires land in the first second of a minute. Do we buy capacity for that second?

Context
Scheduled work clusters on human-legible instants: the top of the minute, the top of the hour, midnight, midnight UTC, 09:00. At an assumed 5,800 fires a second mean, the design peak is around 45,000 a second for five seconds at the top of the hour — roughly 8× the mean minute. Provisioning the dispatch tier for that peak means paying for capacity that is idle 99.9% of the time; the alternative is to accept that a fire is a little late. The question is which fires genuinely care about their exact instant, and the honest answer is that most do not: a digest email at 09:00:03 is indistinguishable from one at 09:00:00, while a market-open task at 09:00:03 may be worthless.
Decision
Fires colliding on a popular instant are smeared within a declared jitter window, with per-trigger opt-out where the instant is externally significant. The peak is absorbed by the jitter window, a dispatch queue with a declared drain rate, and the published lateness budget — not by capacity provisioned for the peak second. The peak capacity avoided by smearing is reported, so the trade is visible rather than assumed.
How it is realised on AWS
The rate shaper applies a deterministic per-trigger offset within the window, so a given trigger's smear is stable rather than randomly different each hour — which matters because a tenant watching their own fires should see a consistent pattern. Opt-out is a flag on the trigger version. The avoided-peak figure is a BigQuery report alongside cost per million fires.
Options weighed
  • ChosenOpt-out jitter smearing plus the lateness budget: Removes the peak as a capacity problem; costs instant exactness for triggers that silently cared
  • RejectedProvision the dispatch tier for the peak: Exact instants for everyone; pays for 8× capacity used five seconds an hour
  • RejectedOpt-in jitter: Safest for correctness; almost nobody opts in, so the peak remains and the mechanism is dead weight
  • RejectedQueue at the peak with no smearing: Half of the chosen option, and the one that makes the whole peak arrive as lateness on the unlucky triggers rather than spread across all of them
Consequences
What it buys
  • Steady-state capacity is sized for the mean rather than the peak, which is the single largest cost lever in the design
  • The smear is deterministic per trigger, so a tenant sees a stable pattern rather than jitter that looks like instability
  • The avoided peak is reported, so the decision can be revisited with a number instead of an argument
What it costs
  • Triggers whose instant mattered and that did not opt out are late by up to the window, and nobody finds out until it matters
  • The opt-out is discoverable mainly by reading the per-trigger lateness distribution, which not every tenant will do
  • "Scheduled for 09:00" is now a statement about a window, which has to be documented honestly
Choose differently when
If the service's tenants were predominantly financial — market opens, settlement windows, regulatory cut-offs — the exactness would be the product and provisioning for the peak would be the right answer. For a general developer platform it is not.
Why it holds up over time
The trade between buying peak capacity and spending a latency budget is permanent and technology-independent; only the prices move. If serverless dispatch capacity became genuinely instantaneous and free at the margin, provisioning for the peak would win, and the jitter window could be set to zero without touching anything else in the design.
LessonBefore buying capacity for a peak, find out how many of the requests in that peak actually care about arriving in it.
Shown on views02 14 16
ADR-11

The due index is a rebuildable projection, verified continuously by a shadow rebuild rather than backed up

Accepted

The due index is corrupt. Do we restore it, or recompute it?

Context
The due index is the structure the scheduler scans, and it is derived entirely from the trigger registry: for each live trigger, the next instant its recurrence produces. Backing it up treats it as data; recomputing it treats it as a function. The second is obviously correct and is also the claim that quietly stops being true — a recompute path that is never exercised is a restore procedure with better marketing, and the first time it runs is during the incident that needed it. The same applies to partition reassignment: a takeover path exercised only during failures is a path nobody has confidence in at the moment they need it most.
Decision
The due index, the partition leases and every cache are rebuildable projections of the registry and are not backed up. The rebuild is exercised continuously: a shadow rebuilder recomputes a partition's instants from the registry and compares them against the live index, and any divergence is treated as an incident. Partition ownership moves routinely in production rather than only during failure, so the takeover path is the common path.
How it is realised on AWS
The shadow rebuilder is a component in the timing plane (view 07), not a script. It walks partitions on a rota, recomputes next instants with the same shared compiler the scan path uses, and reports divergence to BigQuery. Lease reassignment is driven on a schedule as well as by expiry, so the 15-second takeover budget is measured continuously rather than estimated. Only the registry, the ledger and the audit store are backed up.
Options weighed
  • ChosenRebuildable, verified by continuous shadow rebuild, routine reassignment: Makes recoverability a measured property; costs a permanent component and the compute it consumes
  • RejectedBack up the due index: Familiar and cheap to set up; restores a point in time, which for a derived structure is strictly worse than recomputing the current one
  • RejectedRebuildable, with rebuild tested in a lower environment: The common compromise; tests the code, not the production data, which is where the divergence will be
  • RejectedRebuild on demand, no verification: Correct in principle; the claim decays silently and is discovered false during an incident
Consequences
What it buys
  • A corrupt or lost due index is a recompute rather than a recovery procedure, and the recompute has been running all along
  • Takeover is the common path, so the 15-second reassignment budget is a measurement rather than a hope
  • The backup surface is three stores instead of eight, which makes the backup story small enough to actually verify
What it costs
  • A permanent component and its compute, spent entirely on verifying something that is supposed to be true
  • Shadow divergence is reported rather than alerted in the MVP, so there is a window where the claim is unverified
  • Routine reassignment means the fleet is always mid-handover somewhere, which makes duplicate-key rejections normal traffic rather than an anomaly
Choose differently when
If the due index were expensive to recompute — materialised instants for 20 million triggers weeks ahead, as ADR-05 rejected — rebuild would stop being cheaper than restore and the decision would invert. The two decisions are linked: recomputation is affordable because the index holds one row per trigger.
Why it holds up over time
"Derived data is rebuilt, not restored" is a property of the data's relationship to its source and survives any storage technology. The part that needs continued investment is the verification, because the claim is the kind that decays. If the shadow rebuilder is ever switched off for cost, the decision has quietly become the rejected fourth option.
LessonA recovery path you do not run is a recovery path you do not have. If something is rebuildable, rebuild it on a schedule and compare.
Shown on views07 10 15

Evidence and trustWhat the tenant can see, what a schedule is allowed to point at, and who is allowed to multiply work.

ADR-13

Fire history is a tenant-facing product surface, and the platform alert is oldest undispatched due age

Accepted

How does anyone find out that the scheduler has stopped firing?

Context
A scheduler's characteristic failure is unique among services: it is up, serving its API, passing its health checks, consuming no unusual resources, and not firing. Every dashboard says healthy. CPU, memory, error rate, request latency and queue depth are all normal, because the absence of work produces no signal. The same blindness applies to the tenant: a job that did not run generates nothing — no error, no log line, no alert — so the first notification is usually a customer. Both problems are the same problem viewed from two sides, and both are solved by measuring the gap between what should have happened and what did.
Decision
The platform's paging alert is oldest undispatched due instant age per partition, not process health. Fire history is a tenant-facing product surface recording, per fire: scheduled, decided and dispatched instants, attempt outcomes, terminal state, the cause of any non-dispatch, the definition version and the tzdata version — with each trigger's recent lateness distribution published so a tenant sees punctuality degrade before it becomes an incident.
How it is realised on AWS
Fire history is a Bigtable projection of the ledger keyed by tenant#trigger#instant with a 90-day TTL and a 13-month Cloud Storage archive. Oldest undispatched due age is computed per partition from the due index and alerts above an assumed 60 seconds. The History API is a separate service from the Management API (view 08) so a history query never competes with authoring, and the per-trigger lateness view is derived in BigQuery.
Options weighed
  • ChosenDue-age alerting plus tenant-facing per-fire history and lateness: Catches the invisible failure from both sides; costs a projection, an archive and a product surface
  • RejectedHealth and resource alerting: Standard and cheap; blind to the one failure that matters most here
  • RejectedAlert on fire rate dropping below a baseline: Useful as a secondary signal; a legitimate quiet period and a stalled scheduler look identical
  • RejectedPlatform-internal history only, support answers tenant questions: Less to build; makes every "did it run" a support ticket and leaves the tenant's on-call blind (view 05)
Consequences
What it buys
  • A stalled scheduler pages within a minute rather than being discovered by a tenant's customer
  • "Did my job run last night" is answerable by the tenant, with a cause, which is what makes view 05's journey recoverable
  • Three separately recorded instants per fire turn ambiguous complaints into specific ones
What it costs
  • The due-age alert is loud for the whole of a legitimate catch-up drain, which is the price of a visible backlog (ADR-09)
  • 500 million fires a day of history is a storage and query cost that is pure evidence and produces no fires
  • A per-fire record for work the platform did not run invites tenants to read "dispatched" as "succeeded"
Choose differently when
If the scheduler only served platform-internal callers with their own observability, the tenant-facing surface would be unnecessary and the alert alone would do. The product surface is justified by tenants who have no other way to see the gap.
Why it holds up over time
"Alert on the gap between intended and actual, not on the health of the thing in between" is a general principle for any system whose failure mode is silence, and it outlasts any monitoring stack. The specific alert threshold should move with measured data.
LessonFor any system whose failure produces no output, make the alert the distance between what should have happened and what did. Health checks cannot see absence.
Shown on views05 16 17
ADR-14

A target must be verifiably owned by its tenant, no tenant request reaches the fire zone, and dispatch credentials are minted per attempt

Accepted

What stops a schedule from being a timer-driven request forgery primitive aimed at someone else?

Context
A scheduler that will make an authenticated HTTP request to any URL a tenant names, repeatedly, on a timer, with a retry budget, is a remarkably convenient attack tool. It has outbound network access, it is inside the platform's perimeter, it holds credentials, and it is designed to keep trying. The same properties make the dispatch credential valuable: a long-lived platform identity captured from a dispatch grants the holder whatever that identity can do, for as long as it lives. Both problems come from treating the dispatch target as data rather than as a privilege.
Decision
A target must be a registered endpoint on a domain the tenant has verified, or a queue the tenant owns, checked at definition time and again at dispatch time. No tenant request reaches the timing or dispatch planes: both are reachable only from inside the perimeter, and the only component with egress is the attempt runner. Dispatch credentials are minted per attempt with a lifetime no longer than the attempt timeout, and every dispatch is signed with a timestamped, replay-resistant signature over the fire record and payload. Inbound outcome callbacks are verified by the same signature within a short window.
How it is realised on AWS
Target verification is a separate `target` entity with its own verification state (view 11), referenced by the trigger version rather than embedded as a URL. The attempt runner mints a workload-identity token scoped to the trigger for each attempt, so a captured dispatch buys one trigger for one timeout. The callback verifier sits at the perimeter (view 18) and rejects a stale signature rather than logging it.
Options weighed
  • ChosenVerified target entity, perimeter isolation, per-attempt credentials, signed both ways: Closes the forgery path and bounds credential capture; costs tenants a domain verification step
  • RejectedAllow any HTTPS URL with an allowlist of blocked ranges: Lowest friction; a denylist of internal ranges is a game the defender loses eventually
  • RejectedLong-lived per-tenant dispatch credential: Simpler to operate and to debug; a captured credential is a standing grant
  • RejectedEgress through a tenant-supplied proxy: Pushes the problem to the tenant and is right for some enterprises; too much setup for a general platform
Consequences
What it buys
  • A schedule cannot be pointed at a host the tenant does not demonstrably control, so the service is not an amplifier
  • A captured dispatch credential expires with the attempt, which reduces a standing grant to a seconds-long one
  • One component has egress, so the egress policy is small enough to audit
What it costs
  • Domain verification is friction at exactly the moment a developer wants to ship (view 04)
  • Verification is point-in-time: a target legitimately owned at definition and later reassigned is a residual risk mitigated only by re-checking at dispatch
  • Per-attempt minting adds a call to the dispatch path at 45,000/s peak, which has to be cached carefully without defeating the point
Choose differently when
If the scheduler only dispatched to queues and executors inside the platform, target verification would be unnecessary and the perimeter would do the whole job. The tenant HTTP target is what makes this decision load-bearing.
Why it holds up over time
"A dispatch target is a privilege, not a parameter" survives any identity technology, and per-attempt credentials get cheaper as token minting does. What may change is the verification mechanism — domain verification is a convention, not a law — but the requirement that ownership be proved does not.
LessonAny feature that lets a user name an outbound destination is a request-forgery feature until ownership of that destination is proved. Treat the destination as a privilege with a lifecycle.
Shown on views08 18 19
ADR-15

Backfill, horizon extension and quota change are grants separate from trigger authorship

Accepted

Should the person who can create a schedule also be able to make it emit a month of work?

Context
Most operations on a scheduler are small: create a trigger, amend it, pause it, read its history. Three are not. Backfill generates fires for a declared past window. Extending the catch-up horizon turns a bounded recovery into a larger one. Raising a quota lifts the cap on dispatch rate and in-flight concurrency. Each of them multiplies real work — mail sent, money moved, endpoints called — and each of them is most attractive at the moment judgement is worst, which is during or just after an incident. Bundling them with authorship means the blast radius of a routine role is the blast radius of the largest operation in the system.
Decision
Trigger authorship, policy change (catch-up horizon, overlap window, quota) and read-only history access are three separately grantable roles, and backfill is its own privileged operation with its own API surface. Every privileged act is written to an append-only, tamper-evident audit record with actor and prior value. An author who can create a trigger cannot widen its blast radius.
How it is realised on AWS
Three inbound APIs (view 08) rather than one, so the authorisation story is visible in the surface rather than buried in a handler. View 19 draws the horizon extension being *denied* to an author, because the split is only real if the refusal happens. Backfill is additionally rate-capped in its own dispatch lane (ADR-09) and bounded by a declared maximum volume per operation.
Options weighed
  • ChosenSeparate grants for authorship, policy and privilege, with audit: Bounds the routine role to routine damage; costs an approval step during incidents
  • RejectedOne role for everything a tenant can do: Simplest to grant and to explain; makes every trigger author able to emit a month of billing runs
  • RejectedApproval workflow for every change: Safest; makes ordinary authoring slow enough that tenants route around the platform
  • RejectedRate limits instead of privileges: Bounds the volume but not the intent, and a limit high enough to be useful is high enough to hurt
Consequences
What it buys
  • The routine role's worst case is a badly scheduled trigger, not a month of duplicated work
  • An audit trail with prior values makes a bad backfill explainable and attributable afterwards
  • The separation is visible in the API surface, so a security reviewer can check it without reading handler code
What it costs
  • Recovery needs an approver, which adds minutes to the moment a tenant most wants to act (view 05)
  • A tenant that grants the policy role to everyone with the author role recreates the problem; audit detects it and nothing prevents it
  • Three roles and one privileged surface is more to document and more to get wrong in a tenant's own access model
Choose differently when
If backfill were cheap and harmless — regenerating a derived view, say — the separation would be needless ceremony. It is justified by backfill emitting the same work as a real fire, indistinguishable at the executor.
Why it holds up over time
Separating the privilege to *do* a thing from the privilege to *multiply* it is a general access-control principle and does not depend on this identity provider. The specific split of three roles may be refined; the principle that recovery operations are privileged should not be.
LessonFind the operations in your system that multiply work rather than performing it, and make each one a separate grant. They are the ones that will be used under pressure.
Shown on views08 18 19
ADR-16

Tenant payloads live only in the registry and the ledger, encrypted under a per-tenant key, and never in evidence or logs

Accepted

The scheduler carries an opaque blob it never interprets. Where is that blob allowed to exist?

Context
The tenant payload is the one piece of data in the system the platform cannot reason about. It is opaque by design — the scheduler does not interpret it — which also means the platform cannot know whether it contains a customer identifier, an account number, a webhook secret or a medical record. Every piece of machinery the service has wants to copy it: a dispatch log for debugging, a fire history row for the tenant's own view, an error message when a dispatch fails, a metric label for cardinality analysis, a reporting table for cost attribution. Each copy is a place that now carries unknown sensitive data under a different retention and a different access policy.
Decision
A tenant payload exists in exactly two places — the trigger registry and the fire ledger — encrypted at rest under a per-tenant key, and is never written to fire history, reporting, metrics, traces or logs. It is passed to the executor and nowhere else. On tenancy termination, definitions, payloads and history are deleted within 30 days while non-identifying fire counts are retained for billing.
How it is realised on AWS
The payload is stored by reference (`payload_ref`) on the trigger version and resolved only in the dispatch path, where the attempt runner decrypts it under the tenant's key immediately before signing and sending. History rows (view 10) carry the fire key, timestamps, states and causes, and never the payload. Error messages carry the fire key, so a debugging path starts from an identifier rather than from content.
Options weighed
  • ChosenTwo stores, per-tenant key, never in evidence: Keeps the unknown-sensitivity blob inside one small boundary; makes debugging start from a key rather than from content
  • RejectedPayload in history for tenant convenience: Genuinely useful for a tenant diagnosing a bad fire; puts unknown sensitive data under a 90-day analytics retention
  • RejectedOne platform key for all payloads: Much simpler key management; removes crypto-shredding as a deletion mechanism and makes one key the whole estate's boundary
  • RejectedReject payloads entirely; tenants fetch their own parameters: Cleanest privacy position; makes every trigger need a config lookup the tenant must build
Consequences
What it buys
  • The sensitive-data surface is two stores with one retention story, which is small enough to actually review
  • A per-tenant key makes deletion enforceable by key destruction rather than by trusting every copy to be found
  • Debugging conventions start from the fire key, which is a habit that keeps content out of logs by default
What it costs
  • A tenant cannot see the payload a fire carried, which makes some diagnosis harder than it needs to be
  • A decrypt per attempt on the dispatch path at peak, which is a real cost and a dependency on the key service
  • Per-tenant keys at 50,000 tenants is key lifecycle work the platform now owns
Choose differently when
If payloads were constrained to a declared schema the platform validated — identifiers only, no free text — the sensitivity would be known and history could safely carry them. Opacity is what forces the strict boundary.
Why it holds up over time
"Data whose sensitivity you cannot assess gets the strictest boundary you have" is a principle rather than a mechanism, and survives any change of key management. The arguable part is the opacity itself: a future schema-validated payload would legitimately reopen the decision.
LessonIf your system carries data it does not interpret, it also cannot classify it — so give it the handling you would give the most sensitive thing it could be.
Shown on views10 11 18

Every package used, in one table

These terms are used precisely in this package. Several are used loosely in the wider literature on schedulers and cron services, and the difference matters when reading the decision records.

PackageWhat it isWhat it does hereConsidered instead
Scheduled instant The absolute UTC instant a recurrence resolves to, derived from the expression, the named zone and a recorded time-zone database version. The thing the fire is identified by, ordered by, and measured against. "Fire time", which is usually used for the moment a dispatch actually left — the quantity this package calls the dispatch instant.
Fire record An immutable row committed before any dispatch attempt, carrying the tenant, trigger, definition version, scheduled instant and a deterministic key. The scheduler's durable product and the system of record for "was this instant decided". A "job run", which implies work happened; a fire record says only that a decision was taken.
Idempotency key The fire's four-part primary key (tenant, trigger, scheduled instant, sequence), computed rather than generated. Makes a duplicate decision a uniqueness violation instead of something to detect, and gives the executor something to deduplicate on. A generated request identifier, which deduplicates transport retries but cannot deduplicate two independent decisions.
ε (clock uncertainty) The half-width of the interval a bounded time source reports around the current instant. The quantity the scheduler waits out, which is what converts clock skew into lateness rather than earliness. "Clock skew", which usually names the error itself rather than a reported bound on it.
Lateness Dispatch instant minus scheduled instant, measured per fire and published per trigger as a distribution. The budget the design spends; it is the measure of availability for the dispatch plane. "Latency", which in a request-shaped service means time to respond rather than distance from a deadline.
Missed-fire policy A per-trigger declaration — fire-all, fire-once-now or skip — of what to do with instants that should have fired and did not. Moves the post-outage decision from the incident to the definition. "Misfire instruction" in some scheduler libraries, usually with the same three options and no horizon.
Catch-up horizon The maximum age of a missed instant that may still be dispatched; beyond it the instant is recorded as expired. The platform's bound on how much work any interruption can accumulate, overriding the tenant's policy. A "grace period", which usually means how long a late fire is still considered on time.
Catch-up storm A backlog of missed instants released at once by a recovery, a resume or a backfill. The expected failure mode of a scheduler, and the thing the lane split and the rate cap exist for. A "thundering herd", which names simultaneous arrival generally rather than the self-inflicted recovery case.
Overlap policy A per-trigger declaration — allow, skip or queue — of what to do when a fire comes due while the previous execution has not reported terminal. The only control the scheduler has over concurrency in work it does not run. "Concurrency policy", which in some systems bounds parallel runs numerically rather than deciding what to do about one.
Unknown (terminal state) The recorded state of a fire whose outcome never arrived inside its outcome window, including a dispatch attempt that timed out. Prevents the scheduler withholding every later fire on a missing callback, at the price of occasional concurrent execution. "Failed", which asserts the work did not happen — a claim the platform cannot make about a timeout.
Backfill An authorised generation of fires for a declared past window, labelled as such and rate-capped in its own lane. The recovery path for work lost beyond the horizon, and a privilege separate from authorship. A "replay", which in event systems means re-delivering records that already exist rather than creating new ones.
Partition lease A short-TTL claim on a range of the trigger keyspace, recording the ε observed at grant. The unit of mutual exclusion in the timing plane — deliberately imperfect, because the fire key makes imperfection survivable. "Leader election", which implies one authority per service rather than one per partition.
Due index One row per live trigger holding its single next instant, partitioned and scanned in time order. The structure the scheduler reads; a rebuildable projection of the registry, not a system of record. A "timer wheel", which is an in-memory structure with the same purpose and no durability.
Oldest undispatched due age The age of the oldest instant that is due and has not yet been dispatched, measured per partition. The paging alert, because it is the only signal that distinguishes a healthy scheduler from a silent one. "Queue depth", which counts work waiting without saying how long the oldest item has been waiting.
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.