A cascade sends every request to a small model and escalates to the frontier model when the small model's self-reported confidence is low. Escalation ran at 18% in week one. After a prompt edit meant to stop the small model hedging, escalation fell to 2% and cost per request fell 40%. Three weeks later a customer audit finds a run of wrong answers. What failed, and which design decision allowed it?
Show the full answer Hide the answer
The trigger
The prompt edit changed the gate, not the capability. "Stop hedging" is an instruction about expressed certainty, and the escalation decision was reading expressed certainty. The small model answered the same questions equally badly and said it was confident, so 16% of traffic that used to reach the frontier model stopped reaching it overnight.
Why it propagated
The decision that allowed this is structural: the component being gated produced the gating signal. A self-report is not a measurement of correctness; it is another output of the same weights, moved by the same prompt, with no independent reference. Two further choices let it run for three weeks:
- Escalation rate was owned as a cost metric. It fell, which read as a win, and the dashboard that showed it was the finance view.
- Evaluation ran on the wrong population. The suite exercised the full pipeline on a fixed set the small model happened to pass, so the gate change moved no score.
Why detection lagged
There is no error signal in a confidently wrong answer. Latency improved, cost improved, no exception was raised, and ground truth — the customer's reaction — arrived weeks later. The traffic that stopped being escalated is also the traffic nobody looks at, because it was auto-answered and closed.
The structural fix versus the tempting local fix
The tempting fix is to raise the confidence threshold. That restores last month's behaviour and will drift again on the next prompt edit, because the mechanism is untouched.
The structural fix has three parts:
- Gate on something independent of the model being gated. A retrieval-support check (does the retrieved text contain the claim), a deterministic validator (does the SQL parse and run, does the JSON match the schema, does the arithmetic check out), or agreement between two cheap samples. Each is measurable without the model's opinion of itself.
- Treat escalation rate as an SLI with a band, not a cost line. Page when it moves more than roughly 30% relative inside a day, in either direction. A collapse and a spike are both incidents.
- Shadow the auto-handled path. Route a fixed 1% to 5% random sample of auto-answered requests to the frontier model or a reviewer and compare. This is the only measurement that sees the population the cascade decided not to escalate, and it is what turns a three-week discovery into a two-day one.
The general lesson
Any routing decision taken on a component's self-assessment is an open loop, and open loops drift on the first unrelated change. The same failure shows up in guardrails that ask a model whether its own output is safe, and in agents that decide for themselves when a task is complete.
When not to replace the confidence gate
When the downstream check is cheap and deterministic, the self-report costs nothing and helps: let the small model try, validate the output mechanically, and escalate on validation failure. The rule is that the escalation trigger must be the verifier's verdict, not the generator's mood. For tasks with no mechanical check — summarisation, advice, tone — build the shadow sample before the cascade, not after.