Mixed-Initiative Interaction and When to Interrupt
The expected-utility rule that decides whether a system should act on its guess, ask the user, or stay silent, why reversibility moves that threshold further than better accuracy does, and what an interruption actually costs.
An assistant watching a mail thread infers that a meeting is being arranged. It has three options: put the meeting on the calendar, ask whether it should, or do nothing. Most interaction-design arguments about agents are arguments about this choice, conducted without the arithmetic that settles it. The arithmetic has been available since 1999, when Eric Horvitz set out principles for coupling automated services to direct manipulation rather than choosing between them, and demonstrated them in Lookout, a scheduling assistant that read mail and proposed calendar entries (Horvitz, 1999, Principles of Mixed-Initiative User Interfaces, CHI '99).
The threshold is a ratio, not a confidence number
Let \(p\) be the probability that the system's inferred goal is the user's actual goal. Acting correctly is worth \(u_a\). Acting wrongly costs \(c_w\), which is the user's recovery work and not the engineer's estimate of it. Asking costs \(c_d\), a dialog the user must read and answer, after which the right action happens.
Acting beats asking when \(\mathbb{E}[\text{act}] > \mathbb{E}[\text{ask}]\), which rearranges to
The tolerable error rate is set by the ratio of the dialog's cost to the stakes, and nothing else. Put numbers on it. If undoing a wrong calendar entry costs 20 units of the user's attention, a correct one saves 5, and a confirmation prompt costs 0.5, then the system may act only when \(1-p < 0.5/25 = 0.02\): ninety-eight percent confidence. Now make the action cheap to reverse, so undo costs 1 instead of 20. The same prompt cost gives \(1-p < 0.5/6 \approx 0.083\), and the system may act above roughly ninety-two percent.
Four points of accuracy bought what eleven points of model improvement would have. That is the design lever, and it is why Horvitz's list includes minimising the cost of poor guesses and scoping the precision of a service to match the system's uncertainty: a tentative, easily withdrawn action is admissible at a confidence that a committing one is not. Amershi and colleagues later turned the same idea into applied guidance, with separate guidelines for timing services on context, supporting efficient dismissal and correction, and scoping services when in doubt (Amershi et al., 2019, Guidelines for Human-AI Interaction, CHI '19).
What asking actually costs
Designers treat \(c_d\) as the time to read a sentence and click a button. The measured cost is stranger than that. Mark, Gudith and Klocke interrupted 48 subjects doing a simulated office task with telephone and instant messages roughly every two minutes. The interrupted groups finished faster than the uninterrupted control, 20.31 and 20.60 minutes against 22.77, with no measured quality difference. What rose was stress, frustration, time pressure, effort and mental workload (Mark, Gudith & Klocke, 2008, The Cost of Interrupted Work: More Speed and Stress, CHI '08; the publisher does not serve a fetchable copy).
Two consequences follow. First, an interruption's cost lands on affect and effort, which are precisely the quantities a completion-time dashboard cannot see, so a confirmation-heavy agent can look efficient in telemetry and be exhausting to use. Second, the widely repeated figure of "23 minutes and 15 seconds to refocus" does not come from that study and is not a resumption measurement; it entered circulation through a press interview about a different, smaller shadowing study. Designing a confirmation budget around it is building on folklore.
When it breaks
A fixed threshold against a drifting \(p\). The confidence cut-off is tuned once, then the model is updated, the user population widens, or the input distribution shifts. The ratio was right for the old \(p\) and nobody recomputes it.
Dismissal becomes reflex. A system that asks often and is usually wrong teaches users to dismiss without reading. The prompt still appears in the interface and in the compliance document, and has stopped carrying consent. This is the mechanism by which adding confirmations makes a system less safe.
Confirmations do not batch linearly. An agent that asks before each of twelve tool calls pays the context-switch cost twelve times for one task. One review of a proposed plan, or one approval scoped to a class of actions, costs a fraction of that; the arithmetic above applies per decision, so the fix is to make the decision coarser rather than the prompt smaller.
Reversibility is sometimes a lie. The threshold calculation rewards cheap undo, which creates pressure to describe irreversible things as reversible. A sent email, an executed trade and a deleted row have an undo button that restores local state and not the world. When \(c_w\) is unbounded, no \(p\) justifies acting, and the honest design asks.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
6 flashcards for this concept
Click a card to reveal the answer.