Distributed Job Scheduler  ·  View 13 of 20  ·  Runtime

Catch-up After an Outage

The expected failure mode, drawn as a pipeline rather than described as an incident.

Editable source SVG draw.io All views
Gap detected Resume or recovery Missed instant set computed Horizon filter Older than horizon? Record expired with cause Policy fire-all every instant fire-once-now default skip record only Commit Fire records original instants Rate shape Tenant catch-up cap ≤ 10% of quota Catch-up lane separate queue Interleave Fair across triggers oldest-first Deliver Attempt runner Tenant target yes drain rate backlog depth Catch-up — Draining a Backlog Without Causing the Next Outage Application we own Decision point Risk / gap Data store Queue / topic External / third party failure / alternate synchronous The on-time lane is untouched by this path. A backlog drains over hours by design — recovery is deliberately slower than the failure. v 1.0 · owner Platform Architecture · date 2026-10

Decisions

  • The horizon filter runs before the policy, so an instant older than the horizon is expired whatever the tenant declared — the one place the platform overrules the tenant (ADR-08).
  • Catch-up has its own lane and its own rate cap, so recovery is deliberately slower than the failure that caused it (ADR-09).
  • Interleaving is fair across a tenant's triggers and oldest-first within one, so a single trigger's backlog cannot starve the rest.

Assumptions

  • Catch-up horizon default 1 h, maximum 24 h. Catch-up dispatch capped at 10% of a tenant's steady-state quota — so a 1-hour backlog drains over roughly ten hours of added load rather than six minutes of saturation. All assumed.
  • fire-once-now as the platform default, on the grounds that a duplicated billing run is worse than a skipped cache warm.

Risks

  • A tenant who wanted fire-all and accepted the default loses work they expected. Mitigated by the cause on every expired instant and by backfill — not prevented.