advanced 3 min answer

Your policy gate reads each waiver's expiry at evaluation time and blocks when it has passed. Sixty waivers granted against the same standard during one six-week release freeze all share an expiry date next Tuesday. Walk through what happens from Tuesday morning.

gitlabwaiversexpirypolicy-gatescorrelated-failure
Show the full answer Hide the answer

Tuesday, in order

08:00. The first pipelines of the day fail at the policy step. Not one team's: sixty repositories across however many teams, because the waivers were granted for the same standard in the same window and inherited the same term.

08:20. The failures are reported as a gate outage. From inside a single pipeline, a correlated policy change and a broken policy engine look identical. This is the first design defect the morning exposes, and it is cheap to fix: the denial message must name the rule, the waiver id and its expiry date, so one log line distinguishes "your exception ended" from "the engine is down".

09:30. The platform team is asked for a global bypass. Whoever grants it removes the control from the whole estate rather than from sixty repositories, and the bypass will outlive the week.

By midday. Sixty extensions are approved as a batch, with no reassessment of any of them, because no approval path reviews sixty exceptions in a morning. This is the moment a time-boxed waiver becomes permanent while continuing to look governed, and it will not appear in any report, because the register will show sixty live waivers with fresh dates and a named owner.

Where it amplifies

If the week also holds a release train, a quarter end or a compliance deadline, the pressure to bypass becomes irresistible. And the person deciding is whoever is on platform on-call, who has become the de facto risk approver for sixty decisions, at 09:30, with no context on any of them.

What stops it

Mechanisms, not vigilance:

  • Jitter expiry at grant time. Spread the dates across the term, and cap how many may expire in one week at what the approval path can genuinely reassess. If a reviewer handles 4 reassessments a day, 20 a week is the ceiling, and the grant tool should refuse to issue a 21st expiring in that week.
  • Graduate the enforcement. Warn in the pipeline output from 30 days out with a countdown, then block. The warning has to appear in the developer's own build log. A notification to a group alias is how a control fails silently: in GitLab's January 2017 data loss the nightly database dump had been failing for some time because the tool was a major version behind the server, and the failure emails were being rejected outright because the mail configuration did not permit them. Nobody read what nothing delivered.
  • Make extension structurally harder than fixing. An extension needs the risk owner's signature and a new term no longer than half the original, so the second extension is 6 months, the third is 3, and the path converges.

What would have to be true for it to self-heal

That the standard is wrong. Sixty waivers against one standard is the strongest available evidence that the standard does not fit how the estate is built, and the self-healing path is a change to the standard that retires the exceptions rather than a sixty-fold extension of them. The register's aggregate shape is the signal; the individual waiver is noise.

When not to enforce expiry in the gate

Below about ten open exceptions, a spreadsheet and a calendar reminder are adequate and the machinery above is waste. Expiry enforcement in the gate earns its cost at the point where no single person can hold the register in their head, which in practice is a few dozen.