Quality Proxy Metric
also called Implicit Quality Signal, Behavioural Quality Indicator
A continuously available production signal that moves when answer quality moves - abstention rate, regeneration rate, escalation rate - used to page someone in the hours before a labelled evaluation could ever run.
Quality fell this morning. The labelled evaluation runs weekly and the next one is Thursday. Thumbs-down feedback arrives from under 1% of users and is dominated by people who were already annoyed. Between "something changed" and "we can prove something changed" there is usually a gap of days, and that gap is where the damage accumulates.
Proxy metrics close it. They are not measures of quality. They are measures of behaviour that correlates with quality, available on every interaction, at a volume that makes a 20% relative move statistically obvious within an hour.
Why it matters
LLM quality regressions do not move the signals that ordinary alerting watches. Latency is flat, the error rate is zero, token usage is normal, and the output is fluent. The system is working perfectly and giving worse answers, which is the failure mode that conventional monitoring is structurally unable to see.
The volume argument is what makes proxies useful. A labelled set of 300 examples cannot resolve a 3-point accuracy change. A proxy measured on 2 million interactions a day resolves a 10% relative change in minutes, because the sample size is the traffic.
Implementation patterns
- Abstention rate - the share of responses where the model declined or said it could not find the information. It rises sharply when retrieval degrades, and it is unambiguous because the model said so itself.
- Regeneration rate - how often users ask again, rephrase or hit retry within the same session. A user who re-asks has told you the first answer failed.
- Escalation rate - transfers to a human, from an assistant that is supposed to avoid them.
- Conversation length for a task type that should be resolved in two turns.
- Citation coverage - the fraction of factual sentences with a retrieved source attached, computable automatically and a direct proxy for grounding.
- Empty-retrieval rate and mean top-1 score from the retrieval stage, which lead the answer-quality signal because retrieval fails first.
Segment every one of them by model snapshot, prompt bundle version, locale and tenant. An aggregate proxy hides exactly the regressions that hit one segment, and single-segment regressions are the common case.
Industry example
The pattern is visible in reverse in AI incidents published during 2024 and 2025 where output quality degraded for a subset of traffic while error rates never moved and detection took weeks. The common structure is that the only signals being watched were availability signals, and the only quality signal was a periodic evaluation that sampled traffic in a way that did not oversample the affected segment. Teams that recover quickly from this class of incident have one thing in common: a behavioural signal, segmented, with an alert on it.
Failure scenarios
- Proxy and quality decouple. A product change adds a retry button, regeneration rate doubles, and the alert fires on a UI change. Proxies need re-validation against labels whenever the interface changes.
- Gaming. A prompt that discourages abstention improves the proxy and worsens the product, which is Goodhart's law arriving on schedule.
- Aggregate blindness. Global abstention is flat while one locale has doubled.
- Alert fatigue from thresholds set on absolutes rather than on relative change against the same hour last week; these metrics have strong daily and weekly shape.
- No labelled anchor. A proxy with no periodic labelled evaluation behind it drifts into being a number nobody can interpret.
Trade-offs
Proxies buy detection latency measured in minutes instead of days, at essentially no marginal cost, since the data is already being logged. They pay in interpretability: a proxy tells you that something moved and never what is wrong, so every alert needs an investigation path into traces and examples.
They also pay in maintenance. Each proxy's correlation with real quality has to be established once against labels and re-checked after significant product or model changes, which is a recurring obligation of a few hours per quarter rather than a one-off.
When not to use it
Below roughly 10000 interactions a day the statistics do not work - daily noise in a behavioural rate will exceed the regression you want to catch, and reading 50 conversations by hand each morning is both cheaper and better. Do not use proxies as the release gate; they are a detector, not evidence, and a change that improves the proxy still has to pass a labelled evaluation. And where the correct answer is verifiable directly - the model's output is executed, compiled or reconciled against a system of record - measure the real outcome and skip the proxy entirely.
Interview question
Q: Your assistant handles 500000 interactions a day. You have a weekly labelled evaluation of 400 cases and thumbs feedback on 0.8% of interactions. Design the signal that pages someone when quality drops.
What a strong answer covers: why neither existing signal can detect a regression within a day; the choice of two or three behavioural proxies with a stated mechanism linking each to quality; segmentation by prompt bundle, model snapshot and locale; alerting on relative change against the same window in the previous week rather than on an absolute threshold; validating each proxy against the labelled set once before trusting it; and the escalation path from "abstention up 30% in one locale" to the traces and examples that explain it.
Quick check
Quiz: Why can a 3-point accuracy regression be invisible to a 400-case weekly evaluation but obvious in an abstention rate the same afternoon? Sample size - 400 cases carry a margin of error of roughly ±5 points, while the abstention rate is measured on every interaction.
Flashcard: Name three production signals that move when LLM answer quality falls, when none of them measure quality. — Abstention rate, regeneration or retry rate, and escalation to a human; plus retrieval-side leading indicators such as empty-retrieval rate and mean top-1 score.