intermediate 3 min answer Multiple choice

At 02:40 the access-audit service behind an electronic health record starts taking 8 s per write instead of 12 ms. It returns no errors. Audit controls at 45 CFR 164.312(b) are a required implementation specification under the HIPAA Security Rule. What should the record system do while the audit path is degraded?

hipaaaudit-loggingfail-opengraceful-degradationhealthcare
Pick one
Show the full answer Hide the answer

What happens second by second

If the audit write is synchronous with the record read, the request thread blocks for 8 s. At 30 record opens per second, 240 requests are in flight after eight seconds against a 200-thread pool, so the pool saturates in roughly seven seconds and the symptom clinicians see is that charts stop opening. An audit-service slowdown has become a clinical outage, and the dashboards still show the audit service as up because it is returning 200s.

Then it amplifies. Clinicians retry. The emergency department escalates to break-glass access, which is itself an audited event, so the degraded service receives more writes than before.

The mechanism that stops it

Two properties, and neither of them is "monitor it".

  • Durability without synchrony. 164.312(b) requires mechanisms that record and examine activity in systems holding electronic protected health information. It does not require the record to be written before the read returns. Append the audit entry to a local write-ahead file with an fsync, return the chart, and ship asynchronously. The obligation is that the entry exists and can be examined, so durability is the property to preserve and synchrony is the one to give up.
  • A rule per access class, decided in advance. Clinical reads by a clinician with a care relationship proceed. Bulk export, research extracts, administrative search and anything without a treatment justification are refused while the audit path is degraded, because those have no patient-safety argument and are precisely the accesses the audit exists to catch.

Size the queue for the longest plausible degradation rather than the average one: 30 records per second over four hours at about 1 KB an entry is roughly 430 MB, which fits on any host. The common failure is an in-memory ring buffer of 10,000 entries, which covers five minutes and then silently overwrites the oldest evidence.

Why the other options fail

  • Block all record access until audit writes recover. This is the reflex of a team that reads "required" as "synchronous". It trades a compliance risk for a patient-safety event, and an unopenable chart in an emergency department is the worse outcome by a wide margin.
  • Serve every request and drop audit writes that time out. This keeps the hospital running and destroys the evidence, including the evidence of any inappropriate access happening during the window. Dropping is never necessary when a local disk queue is available.
  • Fail the audit service over to its standby and hold all reads until it is healthy. Failover is the right repair and the wrong mitigation. It takes minutes at best, it does not help if the cause is a slow dependency the standby shares, and holding reads meanwhile is the first option with extra steps.

When this is the wrong default

Invert it for systems where the access itself is the risk rather than the care. A bulk extract tool, a data-sharing gateway or an administrative override console should fail closed: if the access cannot be recorded, it does not happen. The dividing line is whether refusing the access can hurt a patient. Write that classification down per endpoint before the incident, because 02:40 is not when it gets decided well.