Fail-Open Audit Boundary
also called Audit Fail-Open Rule, Accountability Degradation Policy
A pre-declared rule saying which classes of access may proceed when the audit trail cannot be written and which must be refused, so a logging failure does not become a safety failure or an invisible one.
At 02:40 an access-audit service starts taking 8 s per write instead of 12 ms and returns no errors. If the audit write is synchronous with the record read, request threads block: at 30 chart opens per second, 240 requests are in flight by the eighth second against a 200-thread pool, so the pool saturates in roughly seven seconds and clinicians stop being able to open charts. The audit service's dashboards stay green, because it is returning 200s slowly.
The team is now choosing between two failures under time pressure. Refusing access protects the audit record and creates a patient-safety event. Serving access without a record keeps the hospital running and destroys the evidence of exactly the period an auditor would most want to see.
A fail-open audit boundary is the decision made in advance instead: a classification of every access path into those that proceed unlogged-but-queued and those that stop, written down before the incident.
Why it matters
HIPAA's Security Rule makes audit controls at 45 CFR 164.312(b) a required implementation specification: mechanisms that record and examine activity in systems holding electronic protected health information. It does not require the write to be synchronous with the read. The obligation is that the record exists and can be examined, so durability is the property to preserve and synchrony is the one to give up.
Teams that read "required" as "synchronous" ship a design where the logging tier's availability is the clinical system's availability, which is a strictly worse system than one whose audit entries land a few seconds late.
Implementation patterns
- Local durable write-ahead. Append the audit entry to a file on the serving host with an fsync, return the chart, ship asynchronously. Losing a host loses at most the unshipped tail, which is seconds.
- A bounded queue with a declared overflow policy. Size for the longest plausible degradation: 30 entries per second for 4 hours at about 1 KB is roughly 430 MB, which fits anywhere. Overflow stops the fail-open class rather than dropping silently.
- Access classes, written per endpoint. Clinical read by a clinician with a care relationship: proceed. Break-glass: proceed and notify immediately. Bulk export, research extract, administrative search, data sharing: refuse while degraded.
- A visible degradation state. The system declares "audit degraded" in its own status, so the retrospective review knows which window to examine rather than inferring it from gaps.
- A reconciliation job comparing shipped entries against a per-host sequence number, so a silent loss is detected rather than assumed absent.
Industry example
Emergency departments are where this is settled in practice, and it is why break-glass is in production in essentially every hospital record system: the design accepts a broad access right with a mandatory reason code, immediate notification and retrospective review, on the reasoning that a delayed chart is a worse outcome than a reviewed access. A fail-open audit boundary applies the same logic one layer down, to the logging path itself.
Failure scenarios
- In-memory ring buffer of 10,000 entries. Covers five minutes at 30 per second, then overwrites the oldest evidence of the incident window.
- Fail-open applied to every path, so the bulk export tool runs unlogged during the degradation, which is the one access an auditor would ask about.
- No declared degradation state, so the retrospective review cannot bound the affected window and treats the whole day as suspect.
- Retry amplification. Clinicians retry, break-glass rises, and break-glass events are audited, so the degraded service receives more writes than before.
Trade-offs
Fail-open buys availability of care and pays with a window in which the audit record is delayed and, occasionally, incomplete. Fail-closed buys a perfect record and pays with clinical unavailability. The choice is not global: it is per access class, and the dividing line is whether refusing the access can hurt a patient.
Local write-ahead costs disk on every serving host, an extra shipping path to operate and a reconciliation job. That is genuinely more machinery than a synchronous call, and it is the machinery that stops a logging incident from becoming a clinical one.
When not to use it
Outside care delivery the default usually inverts. A data-sharing gateway, an analytics extract tool or an administrative override console should fail closed: if the access cannot be recorded, it does not happen, and nobody is harmed by the refusal. The same applies to any system where the access itself is the risk rather than an input to a time-critical decision. If refusing costs only inconvenience, refuse. And if the audit service is genuinely more available than the application, the whole pattern is unnecessary complexity: measure both before designing around a failure that does not occur.
Interview question
Q: Your electronic health record writes an audit entry synchronously before returning a patient chart. Someone proposes making it asynchronous. Argue both sides, then tell me what you would actually ship and how you would prove to a regulator that nothing was lost.
What a strong answer covers: thread-pool saturation arithmetic showing the synchronous design converts a slow dependency into an outage; the distinction between durability and synchrony in the obligation; a local fsync write-ahead with a sequence number per host; a reconciliation job that proves completeness rather than asserting it; and a per-access-class rule that keeps bulk and administrative paths fail-closed.
Quick check
Quiz: Why does a slow audit service produce a worse outage than a dead one in a synchronous design? Because a dead service fails fast and the caller can apply a fallback, while a slow one holds request threads until the pool saturates and every path through the application stops.
Flashcard: What decides whether an access class fails open or closed when the audit path is degraded? Whether refusing the access can harm a patient. Clinical reads proceed with a durable local queue; bulk export and administrative consoles refuse.