Your output safety classifier blocks 0.3% of responses in English and 6% in Portuguese and Turkish. Complaints arrive only from those two markets. The vendor confirms the same model version serves every region and reports no incident. Where do you look first?
Show the full answer Hide the answer
The first three things to look at, and why in that order
- A labelled sample per locale. 300 to 500 responses per language, reviewed by someone fluent, split into truly unsafe and wrongly blocked. Without it you cannot tell whether Turkish traffic is genuinely more borderline or whether the classifier is miscalibrated, and those two findings lead to opposite actions.
- The score distribution, not the block rate. Plot classifier scores per locale. A shifted distribution with the same threshold is calibration; an overlapping distribution with a genuinely heavier tail is content.
- What the classifier is actually shown. Whether the text was machine-translated before scoring, whether locale-specific formatting or transliteration is applied, and whether a system prompt in English is being concatenated with non-English output.
The diagnosis
Safety classifiers are calibrated on their training mix, which is overwhelmingly English, so the same numeric score means different things in different languages. A single global threshold therefore enforces a different strictness per market. Idioms, direct speech registers and words that are innocuous in one language but adjacent to a flagged concept in another all push scores up without any change in intent.
The magnitude matters for the argument you have to make internally. At 200000 responses a day in those markets, 6 percent is 12000 wrongly blocked answers daily, against roughly 600 at the English rate. That is a product outage in those markets that no availability metric will ever show.
The misleading signal
The vendor's "same model version everywhere, no incident" is true and irrelevant. It steers the investigation towards deployment differences and away from the actual cause, which is that one threshold is being applied to populations the model scores differently. Identical configuration producing different outcomes is the signature of a calibration problem, not a deployment one.
Why the other options fail
Raising the threshold globally trades a false-positive problem in two markets for a false-negative problem in all of them. It would let genuinely unsafe English output through in order to fix Portuguese, which is the wrong direction on the axis that actually carries risk.
A larger safety model may have better multilingual coverage, and it is a months-long procurement decision taken before anyone knows the current model's per-locale error rates. It also does not remove the single-threshold mistake, which will reappear at a different rate. Measure first; this may become the answer, but not as the first move.
An automatic regenerate on block hides the symptom and doubles the cost of the affected traffic. The regenerated response is drawn from the same distribution and scored by the same miscalibrated classifier, so it frequently blocks again, and the user now waits twice as long for the same refusal.
The fix, and what it costs
Calibrate per locale against the labelled sample, and treat the threshold as configuration rather than a constant. Add locale as a first-class dimension in guardrail telemetry, alert when any locale's block rate exceeds twice the global median, and give reviewers an override path that feeds the labelled set. The cost is real: per-locale thresholds must be re-validated whenever the classifier or the generating model changes, and that is now a recurring obligation on a small team.
When not to calibrate per locale
When the product serves one language, or when a false negative is catastrophic and a false positive is merely annoying - a child-safety context, for instance. Over-blocking is the right trade when the asymmetry justifies it; the failure here is that nobody chose the asymmetry, they inherited it from the vendor's default.