Distributed Job Scheduler · View 16 of 20 · Operations
Decisions
- The page is on oldest undispatched due age, not on process health. A scheduler that is up and not firing is healthy on every other dashboard — this is the characteristic failure of the system (ADR-13).
- Clock ε is a per-node metric, because a fleet-wide drift is invisible in an average and is one of the two findings that would change the design.
- Tenant-facing signals are a product surface with their own row, so "did it run" never requires platform access.
Assumptions
- Oldest undispatched due instant older than 60 s in any partition pages immediately. Assumed, and knowingly noisy during a legitimate catch-up drain.
Risks
- Shadow divergence is reported, not alerted on, in the MVP. A divergence that appears between reports is a window in which the rebuild claim is unverified.