advanced 3 min answer

Interview prompt. A risk committee asks you to show that the guardrails on your customer-facing assistant work. You have a classifier on inputs and another on outputs, and a 99.7% pass rate in production. Convince me.

guardrailsassurancered-teamfalse-positivesdefence-in-depth
Show the full answer Hide the answer

What the interviewer is testing

Whether you understand that a production pass rate measures your traffic, not your guardrail. 99.7% pass says a classifier fired on 0.3% of responses. It says nothing about how many genuine violations it missed, because the denominator — true violations — is unknown and not in that number. A candidate who leads with the pass rate has not understood what is being asked.

The clarifying questions that change the answer

  • Which harm class, and what is the tolerance for each? A leaked account number and an off-tone sentence are not the same risk and must not share a threshold.
  • Is the control preventive or detective? A check that blocks is an availability dependency; a check that flags is a review workload. The committee is usually asking about the first and funding the second.
  • Who is the adversary? A confused customer and someone paid to break it need different evidence.
  • What has to be produceable afterwards? If the regulated artifact is a record of why an output was allowed, that is a logging requirement, not a classifier requirement.

A strong answer's arc

  1. Separate recall from pass rate. The quantity that answers the committee is: of known violations, what share does the guardrail catch, and at what false-positive cost. That is measurable only against labelled sets.
  2. Three sets, maintained. A red-team set of hand-written attacks per harm class, refreshed quarterly and never pasted into a prompt or a few-shot example. A regression set holding every real incident, which only grows. A benign-but-risky set that measures false positives, because refusing a legitimate claim has its own cost. A few hundred cases per harm class resolves recall to roughly ±5 points, which is enough to argue about and not enough to be precise about.
  3. Report per harm class and per language, never aggregated. Aggregates hide the markets where a classifier fails, and those failures arrive as a regional complaint pattern rather than as a metric.
  4. Show defence in depth, and say which layer carries the risk. The classifier is probabilistic. The controls that actually bound the worst case are deterministic: the permissions on the tools the assistant can call, an output validator that rejects an account-number pattern or an off-allowlist URL, and schema conformance. State plainly that the classifier reduces frequency while the capability limits reduce consequence.
  5. Show the availability policy. What happens when a classifier times out, decided in advance per action class: fail closed for payments, fail open with a flag for a help answer. Otherwise a 03:00 timeout decides your safety posture.
  6. Show the loop. Time from an incident to a deployed rule, and the number of blocked samples a human reads weekly. A guardrail nobody reviews becomes a filter nobody can characterise.

Common weak answers

  • "Our pass rate is 99.7%." Measures traffic mix. It will improve on its own if a marketing campaign brings in tamer users.
  • "We use a frontier model as the judge, so it is accurate." Unmeasured without a labelled set, and it shares failure modes with the generator.
  • "We hardened the system prompt." Instructions are not a boundary when untrusted text reaches the same context.
  • "One global threshold, tuned on a sample." Guarantees the strictest harm class is under-protected and the loosest is over-blocked.

What a strong answer adds

The committee's real question is what is the worst thing this can do, and who finds out first. Answer with blast radius — what the assistant is authorised to do, not what it is able to say — then the detection path and the rollback. Finish with the honest limit: recall on attacks nobody has invented yet is unmeasurable by construction, which is the argument for spending the next quarter on capability limits rather than on a better classifier.