practice

Guardrail Availability Policy

also called Fail-Open Guardrail Decision, Safety Check Degradation Policy

The decision, made in advance and per action class, about what the system does when a safety check cannot run - because the alternative is that a timeout in a classifier decides your safety posture at 03:00.

guardrailsavailabilityfail-opensafetydegradation

The output classifier times out. The request is in flight, the user is waiting, and the code has to do something. Whatever the default path does is now your safety policy, and in most systems that default was chosen by whoever wrote the try/except, not by anyone accountable for the risk.

A guardrail is a dependency on the request path. Adding one converts a safety question into an availability question, and the two have opposite defaults: availability engineering says degrade and serve, safety engineering says refuse. The policy is the written answer to which one wins, for which actions.

Why it matters

Guardrail outages are common precisely because guardrails are cheap to add. A classifier hosted by a third party, a moderation endpoint, a policy service - each one is a network call with its own p99 and its own incident history. At three guardrails per request, an availability of 99.9% each gives the composite about 99.7%, which is a few hours a year of "the check could not run" that somebody has to have decided about.

Making it explicit also makes it auditable. A regulator or an enterprise customer asking "what happens if your safety filter is down" is asking for this document.

Implementation patterns

  • Classify actions, then set the policy per class. Irreversible, public or high-value actions fail closed. Drafts a human will read before sending fail open with a marker. Internal read-only responses fail open silently.
  • Fail degraded, not binary. If the model-based classifier is unavailable, run the deterministic checks that are always available - schema validation, a regex denylist, an output length cap - and record that the response was served under a reduced check set.
  • Timeout below the user's budget, not at the vendor's default. A safety check with a 30 s timeout inside a 5 s request budget is not a guardrail, it is a way to time out the request.
  • Cache verdicts for identical content so a repeated output does not depend on the classifier being up twice.
  • Keep an in-process fallback. A small local classifier or a denylist that ships with the binary is worth more during an outage than a better remote one.
  • Emit the check set applied on every response so the degraded population can be found and reviewed afterwards.

Industry example

The building block most teams start with is a hosted moderation endpoint - OpenAI publishes one free of charge to its API users, which is why it appears in so many first designs. Free and remote is an appealing combination, and it is also a network dependency on somebody else's control plane that sits inside your request path. Teams that adopt it without an availability policy discover the gap during the first provider incident, when the only options available at that moment are to block everything or to check nothing.

Failure scenarios

  • Silent fail-open. An exception handler returns "safe" on error, and during a two-hour classifier outage every response ships unchecked, with nothing in the logs distinguishing them.
  • Fail-closed on the wrong class. A classifier blip blocks all drafting, and the product looks broken for a risk that was never severe.
  • Retry amplification. A timing-out classifier is retried three times per response, tripling load on a service that is already failing.
  • Unbounded queueing. Requests wait for the guardrail rather than degrading, so the assistant's latency becomes the classifier's latency and threads exhaust.
  • No record of the degraded population, so after recovery nobody can review what went out unchecked.

Trade-offs

Fail-closed buys a defensible safety position and pays with availability that is now the minimum of every check's availability. Fail-open buys availability and pays with a window where the product's safety claims are untrue. Fail-degraded is the usual right answer and pays in engineering: two check paths to build, test and keep in step, plus a review process for the degraded population.

The cost that surprises people is the review obligation. Serving under reduced checks means somebody has to look at what was served, which is real recurring human work, not a config flag.

When not to use it

When there is only one guardrail and it is local and deterministic, there is no meaningful availability question - a regex does not have an outage. When the action class is uniformly low-risk, a single global fail-open with a logged marker is honest and sufficient; writing a policy per class is ceremony. And when the guardrail is the product - a moderation service - failing open is not a degradation, it is a defect, and the discussion belongs in capacity planning instead.

Interview question

Q: Your output safety classifier is a third-party call with a p99 of 700 ms and roughly two incidents a year. The product streams responses to users. Design the behaviour for every failure mode you can name.

What a strong answer covers: the timeout set from the user's remaining budget rather than the vendor's default; the streaming problem, since a check that needs the whole response cannot gate a stream - so either buffer, or check incrementally and be able to retract, or accept that the check is post-hoc; per-action-class fail behaviour; a local fallback denylist; a per-response record of which checks actually ran; and an alert on the rate of degraded responses rather than on the classifier's own health, because the degraded rate is the thing with a business meaning.

Quick check

Quiz: Three guardrails at 99.9% availability each, all on the request path. What is the composite availability of the checked path, and what does that imply? About 99.7%, so several hours a year running without at least one check - which must be a decision, not an exception handler.

Flashcard: Your safety classifier times out. Fail open or fail closed? — Neither as a global rule: fail closed for irreversible or public actions, open-with-a-marker for drafts a human reviews, and always fall back to the deterministic checks and record which checks actually ran.