Architecture One-Pager
Solution Architecture v1.0 · Google Cloud · Platform Architecture · 2026-10 · 20 views · 16 architecture decision records
Distributed Job Scheduler · Solution Architecture v1.0 · Google Cloud · Platform Architecture · 2026-10 · 20 views · 16 architecture decision records
The timing plane decides; the execution plane works; and the only thing crossing between them is an immutable, uniquely-keyed fire record that is committed before it is delivered.
Every developer platform ships a four-word scheduling box, and behind it is a problem that looks solved and is not. A multi-tenant timer fleet has to fire tens of millions of independent triggers at the right instant while a partition owner is being replaced, a node's clock is wrong, a zone is draining, and a tenant has just resumed four thousand triggers that slept through a day. The convenient implementation — fire from a node's system clock and record the result afterwards — fails in three ways that each look like success at the time. It fires early when a clock runs fast, producing work that reads a window of data which has not closed yet and reports a short answer. It loses fires that happened but were never recorded, which no amount of later reconciliation can recover because the record that was supposed to prove it is the thing that is missing. And when it comes back from an outage it releases everything it owes at once, turning a scheduling incident into an incident for every executor downstream. On top of those, it cannot answer the one question tenants actually ask — "did my job run last night?" — because a job that did not run produces no signal at all, in a service that is up, healthy on every dashboard, and silent.
Separate the plane that decides from the plane that works, and make the scheduler's durable output a decision rather than a result. A partition owner asks a time source for a bounded uncertainty interval, scans its share of a time-ordered due index for instants due at or before the lower bound of that interval — so it is deliberately ε late and never early — applies the trigger's declared overlap policy, and commits an immutable fire record whose primary key is derived deterministically from the tenant, the trigger and the scheduled instant. That key is the idempotency key, so two owners mid-handover compute the same one and the second insert loses to a uniqueness constraint instead of producing a second fire. Only then is anything dispatched, through one of three separate queues — on-time, catch-up, backfill — each with its own drain rate, so a backlog released by a recovery is rate-shaped rather than flooded and recovery is deliberately slower than the failure that caused it. Delivery is at-least-once with the key attached and a published residual duplicate rate, not an exactly-once claim nobody can verify. Everything derived — the due index, the caches, the fire history — is a rebuildable projection of a registry and a ledger that are the only systems of record, and the rebuild runs continuously in a shadow partition rather than waiting to be needed. What the tenant gets is a per-fire record with its scheduled, decided and dispatched instants and the cause of any non-dispatch; what the platform watches is the age of the oldest undispatched due instant, which is the only signal that can tell a working scheduler from a silent one.
What it is, and what it is not
- A durable decision, committed before delivery — not A timer that fires and then writes down what happened
- Deliberately ε late, never early, at any percentile — not As punctual as the node's clock allows, in both directions
- At-least-once with a key and a published duplicate rate — not Exactly-once delivery
- Idempotency as a uniqueness constraint in the data model — not A deduplication service on the dispatch path
- Imperfect leases made survivable by the fire key — not Leader election strong enough to be trusted alone
- Post-outage behaviour declared per trigger beforehand — not An operator decision taken during the recovery
- Recovery rate-shaped slower than the failure — not Drain the backlog as fast as the pipeline allows
- A scheduler that hands over work — not A workflow or DAG orchestrator
- Fire history as a tenant-facing product surface — not A platform log that support can search
The decisions that are the architecture
- Commit before dispatch, always (ADR-01) — The fire record is written before any attempt is made, so a dispatcher that dies mid-attempt loses an attempt and never a fire. Of the two orderings available, one failure is unrecoverable and the other is a retry; the ordering of two writes decides which one the system can have.
- The control plane cannot take the fire path down (ADR-02) — Two planes, separately deployed and scaled, sharing only the state of record. Console traffic, a bad management deploy and an identity outage become latency incidents instead of missed fires. One write region, because two regions evaluating "is the previous run still going" produce two answers no key reconciles.
- Idempotency lives in the data model (ADR-03) — The fire's primary key is (tenant, trigger, scheduled instant, sequence) — computed, not generated. Split ownership becomes a constraint violation rather than a detection problem, which is what lets leases be short, takeover fast, and leader election imperfect.
- Late is a budget; early is a bug (ADR-04) — The node clock is untrusted. The scheduler takes a bounded uncertainty interval, waits it out, and scans for instants due at or before its lower bound — paying about 10 ms of deliberate lateness to make earliness impossible. A node whose uncertainty exceeds the ceiling sheds its leases rather than firing on a clock it cannot justify.
- Recurrences are recomputed, not materialised (ADR-05) — One due-index row per trigger, holding one next instant, recomputed on advance against the named zone. An amendment, a pause, a validity change and a time-zone database update are all picked up with no invalidation sweep, because nothing was cached to invalidate — and dry-run is exact rather than indicative.
- The post-outage decision is declared before the outage (ADR-06) — Every trigger declares its missed-fire policy from a closed set of three, defaulting to firing once for the most recent missed instant. A billing run and a cache warm have opposite correct answers and look identical to the platform; the only party who knows which is which is the tenant, and the only time they can say so is in advance.
- Wall-clock time means what the tenant meant (ADR-07) — Expression plus IANA zone, resolved at computation, with declared rules for the day a time does not exist and the day it happens twice, and the time-zone database version recorded on every fire — so a disputed instant six months later is explainable from data rather than from memory.
- Every policy is bounded by a horizon (ADR-08) — An instant older than the catch-up horizon is recorded as expired with its cause and never dispatched, whatever the tenant declared — the one place the platform overrules them. Caught-up fires carry their original scheduled instant, because the work reads a window defined by it.
- Recovery is slower than failure, by design (ADR-09) — Three queues with independent drain rates and a per-tenant catch-up cap, plus a published shedding order: backfill, catch-up, over-quota, then on-time. A one-hour backlog takes about ten hours to drain, which is the price of not making the recovery into the next outage for every executor downstream.
- The peak is absorbed, not purchased (ADR-10) — Sixty per cent of fires land in the first second of a minute. Deterministic opt-out jitter smearing and the lateness budget absorb a 45,000/s peak that lasts five seconds an hour, instead of provisioning eight times the steady-state capacity to be idle for the rest of it.
- Rebuildable means rebuilt on a schedule (ADR-11) — The due index, the leases and every cache are projections of the registry and are not backed up. A shadow rebuilder recomputes and compares continuously, and partition ownership moves routinely rather than only during failure — because a recovery path that is only exercised in an incident is a claim, not a capability.
unknownis a state, not an absence (ADR-12) — Overlap is enforced against recorded outcome state bounded by a window. A fire still unknown at the end of its window counts as finished, and the decision is recorded. The alternative — treating it as still running — turns one lost callback into a permanently stopped trigger that looks perfectly healthy.- Alert on the gap, not on health (ADR-13) — The paging signal is oldest undispatched due instant age, because a scheduler that is up and not firing passes every other check. The same measurement, per fire and per trigger, is a tenant-facing product surface — so "did my job run" is answerable without reading platform logs or opening a ticket.
- A dispatch target is a privilege, not a parameter (ADR-14) — Targets must be verifiably owned, checked at definition and again at dispatch; no tenant request reaches the fire zone; and credentials are minted per attempt with the attempt's lifetime. A service that will make authenticated requests to any URL on a timer, with a retry budget, is an attack tool until ownership is proved.
- The operations that multiply work are separate grants (ADR-15) — Backfill, horizon extension and quota change are grantable apart from trigger authorship, and audited with prior values. Each of them is most attractive at the moment judgement is worst, which is during the incident — so the routine role's worst case stays routine.
- Opaque data gets the strictest boundary (ADR-16) — The tenant payload exists in two stores, encrypted under a per-tenant key, and never in history, reporting, metrics or logs. The platform cannot classify what it does not interpret, so it handles the payload as the most sensitive thing it could be and debugging starts from a fire key.
Why this should still be right in ten years
Clouds, queue services, container runtimes and consistency models will all be replaced inside this platform's life. These are the properties that should outlast them.
- The boundary is not a technology. "The timing plane's product is a committed decision" is a statement about where authority lives, not about which store holds it. Spanner, a successor nobody has shipped, or a different strongly consistent store can each satisfy it, and every decision downstream — the key as the deduplication mechanism, lanes as retries against records, history as a projection — survives the substitution unchanged.
- Deterministic identity outlives every protocol. Naming a fire from its own semantics rather than generating an identifier for it is the oldest reliable trick in distributed systems. It depends on no store, no broker and no transport, and it is why imperfect leader election is affordable here. What could change is executors' willingness to deduplicate — which would call for an added suppression store, not a redesign.
- The asymmetry between late and early is permanent. Scheduled work reads windows defined by its scheduled instant. That is a property of the work, not of the scheduler, and no improvement in clocks makes an early fire safe. If clocks get better, ε shrinks and the deliberate lateness disappears on its own; if they get worse, the lease-shedding ceiling catches it.
- Recovery will always need to be slower than failure. A system that accumulates obligations while it is down and then discharges them into dependencies it does not control has a feedback loop into its own failure. That is control theory, not engineering fashion. The cap's value should move with measured executor capacity; the existence of a cap should not.
- Silence will always need a gap measurement. Any system whose failure produces no output is invisible to health checks, forever. Alerting on the distance between what should have happened and what did is the only class of signal that sees it, and that remains true whatever the monitoring stack looks like.
- The declared-policy habit transfers. Collecting, in advance, the information a system will need to make a decision it cannot make correctly during an incident is a general property of operable systems. Here it is the missed-fire policy; elsewhere it is a failover preference or a data-loss tolerance. The mechanism is configuration; the principle is foresight.
Non-functional targets
Every figure below is a stated assumption. They are listed with the mechanism that is supposed to deliver them and the view where that mechanism is drawn, so a reviewer can disagree with the number and the means separately.
| Quality | Target | How it is met | View |
|---|---|---|---|
| Fire punctuality | p50 ≤ 1 s, p99 ≤ 5 s, p99.9 ≤ 30 s of lateness | Due scan at 1 s tick on leased partitions, with the dispatch lane's drain rate sized for the mean rather than the peak | 12 |
| Earliness | Zero at every percentile | Scan for instants due at or before the lower bound of the clock's uncertainty interval; ε is waited out, not averaged | 12 |
| Degraded punctuality | p99 ≤ 60 s for up to 10 minutes during zone loss or reassignment | Partitions spread across three zones; reassignment exercised routinely rather than only on failure | 15 |
| Dispatch plane availability | ≥ 99.95% monthly, measured as due instants dispatched inside their p99.9 budget | Punctuality as the availability definition; stateless dispatch tier scaled independently of the timing tier | 06 |
| Control plane availability | ≥ 99.9% monthly, failing independently of the fire path | Separate deployment, scaling and service accounts, sharing only the state of record | 06 |
| Trigger population | 20,000,000 active triggers across 50,000 tenants | 256-partition keyspace with one due-index row per trigger; partition count changeable online | 07 |
| Steady-state throughput | 500,000,000 fires/day — 5,800/s mean | One transactional fire insert per fire; lane queues with declared drain rates | 09 |
| Peak throughput | 45,000 fires/s for 5 s at the top of the hour | Deterministic opt-out jitter smearing plus the lateness budget, instead of provisioned peak capacity | 14 |
| Clock uncertainty | ε ≤ 10 ms; leases shed above 100 ms | Bounded-uncertainty time source; per-node ε recorded at lease grant and alerted per node, not averaged | 12 |
| Duplicate dispatch | ≤ 1 in 10^6 fires delivered more than once | Deterministic fire key as the ledger primary key; executors deduplicate on it | 11 |
| Durability | RPO 0 in-region; cross-region RPO ≤ 15 s | Synchronous three-zone replication for registry and ledger; asynchronous standby replica | 15 |
| Recovery time | RTO 5 min dispatch plane, 30 min control plane; partition reassignment ≤ 15 s p99 | Declared region promotion; short leases made safe by the fire key; routine reassignment | 15 |
| Catch-up bound | Horizon 1 h default / 24 h max; catch-up ≤ 10% of tenant quota | Horizon filter ahead of the policy; separate catch-up queue with its own drain rate | 13 |
| Configuration propagation | Create, amend, pause, resume, delete effective within 5 s p99 | Recomputed due-index row written in the same transaction as the version | 09 |
| History | 90 days queryable, 13 months archived; 30-day single-trigger query p95 ≤ 500 ms | Bigtable projection keyed tenant#trigger#instant with TTL; archive-class object storage beyond it | 10 |
| Detection | Oldest undispatched due instant > 60 s pages immediately | Derived per-partition metric from the due index, alerted independently of process health | 16 |
| Cost | ≤ $0.60 per million fires at steady state | Mean-sized capacity, no per-trigger always-on resource, tiered history, retry and catch-up charged to the originating tenant | 16 |
Scope
In scope
- Trigger definition and lifecycle — calendar, interval and one-shot recurrences in named zones, with versioning, pause, resume, delete, manual run and exact dry-run
- Time semantics: a bounded-uncertainty clock, declared daylight-saving rules, and a recorded time-zone database version per fire
- A partitioned, time-ordered due index with lease-held ownership and clock-gated instant claiming
- An append-only fire ledger whose primary key is the idempotency key, committed before any dispatch
- At-least-once dispatch to an authenticated HTTP target, a tenant queue or an internal executor, with bounded jittered retry truncated at the next instant
- Missed-fire policies, a platform-enforced catch-up horizon, and a bounded authorised backfill
- Overlap policies with
unknownas a terminal state, and per-trigger and per-tenant concurrency limits - Per-tenant quotas, fair dispatch queues, deterministic jitter smearing and a published shedding order
- Tenant-facing per-fire history with causes, and the oldest-undispatched-due-age alert
- Target ownership verification, per-attempt dispatch credentials, signed dispatch and verified callbacks, and the privilege split for backfill, horizon and quota
Explicitly out of scope
- Executing the triggered work — the scheduler hands over a record and a credential and nothing more
- Workflow and DAG orchestration, retries of business logic, and compensation; the separate distributed-workflow-orchestration-platform package covers that ground
- Trigger-to-trigger dependencies, deferred to Phase 3 and only if they can be added without becoming an orchestrator
- Sub-minute recurrences, which need a separately budgeted punctuality contract
- Tenant-supplied recurrence predicates such as business calendars and market holidays, deferred to Phase 3 behind a sandboxed evaluation budget
- Active-active multi-region scheduling, deferred because two authorities evaluating overlap produce answers no key reconciles
- The product surfaces that render a schedule, and the tenant's own alerting on its work's outcome
What a four-week prototype should prove
The prototype's job is to falsify the two central claims — that imperfect leases plus a deterministic key really do make duplicate fires harmless, and that a bounded clock lets the scheduler be reliably late and never early — on one tenant and a few hundred triggers, not to build a platform.
- One tenant, 500 triggers across all three recurrence forms, in three time zones including one with daylight saving, with the two awkward days exercised deliberately rather than waited for.
- The timing plane on leased partitions against a strongly consistent store, with the fire key as the ledger's primary key and the clock gate reading a bounded uncertainty interval.
- Two dispatch lanes — on-time and catch-up — with independent drain rates and a per-tenant cap, against one deliberately slow HTTP target and one that times out without answering.
- The shadow rebuilder, running from day one, so the rebuild claim is under measurement for the whole four weeks rather than tested at the end.
- Fire history with scheduled, decided and dispatched instants and the cause of every non-dispatch, queryable by the tenant — because the prototype has to answer "did it run" to be worth anything.
- Measured cost per million fires on the prototype's shape, compared against the assumed $0.60, which is the assumption with the least evidence behind it.
- A partition owner is killed mid-cycle, repeatedly, under load: every instant has exactly one fire record and at most one extra dispatch attempt, with the measured duplicate rate compared to the assumed 1 in 10^6
- A lease handover happens while a scan is in flight: the second insert loses to the uniqueness constraint, and no application code resolves the conflict
- A node's clock is stepped forward by two seconds: no fire goes out early, at any percentile
- A node's reported ε is pushed past 100 ms: it sheds its leases and the partition is reassigned inside the 15-second budget, rather than firing on the bad clock
- A 1-minute trigger is paused for three hours and resumed under each of the three missed-fire policies: the dispatch rate stays inside the 10% catch-up cap, and the horizon expires what it should with a cause the tenant can read
- A dispatch is held open past its timeout with no callback: the fire is recorded unknown at the end of its window, the overlap decision is recorded, and the next fire is not withheld indefinitely
- The due index is dropped entirely and rebuilt from the registry: the result matches what the shadow rebuilder had been reporting all along
- A trigger is pointed at a domain the tenant has not verified: the definition is refused at write time, not at the first fire
Open risks, carried rather than hidden
| Risk | If it lands | Response |
|---|---|---|
| An executor that cannot deduplicate on the fire key | At-least-once stops being a contract and becomes a defect: a duplicate fire produces duplicate work, and the whole leases-may-be-imperfect argument collapses into a correctness problem | Make deduplication an onboarding requirement with a conformance test rather than a documentation note; publish the duplicate rate on the target's own page; and keep a platform-side suppression store as a costed Phase 2 option for targets that genuinely cannot. |
| Fleet-wide clock drift, where ε is wrong everywhere at once rather than on one node | Lease-shedding has nothing to compare against, the clock gate approves a bad interval, and the design's central guarantee — no early fires — fails silently and uniformly | Alert on per-node ε rather than an aggregate, cross-check the time service against the store's commit timestamps as an independent source, and treat a correlated ε excursion as a platform incident with scheduling suspended rather than degraded. |
| The missed-fire default is wrong for most tenants, and they accepted it without reading it | Either duplicated billable work or silently skipped work, discovered by a customer, with the platform's correct behaviour as the explanation — the trough in view 05 | Instrument override rates by tenant segment and treat a high override rate as evidence the default is wrong; show the policy's consequence in a sentence at the point of choice; and make dry-run replay what each policy would have done to the last recorded gap. |
| The oldest-undispatched-due-age alert is loud during every legitimate catch-up drain | The one signal that detects the system's characteristic failure becomes the one the on-call rotation learns to ignore, which is worse than not having it | Compute the metric separately per lane so a catch-up drain does not mask an on-time stall, suppress it against a declared drain with an expected completion time, and page on the on-time lane alone. |
| Recomputing instants for 20 million triggers every cycle becomes the scaling wall | Scan cost grows with the trigger population rather than with the fires actually due, and punctuality degrades as the population grows rather than as load does | Measure compiler cost per trigger per cycle from day one as a first-class capacity metric; keep the short-horizon materialisation variant of ADR-05 costed and ready, since it is a localised change to the due index's contract and leaves the ledger, the key and the lanes untouched. |
The reasoning behind every component and technology choice is in the Architecture Decision Record: 16 records across 5 areas, each with the alternatives that lost and what the choice costs.