Interaction Design for AI intermediate 7 min read 6 flashcards

Setting Expectations Before First Use

Why the direction of a system's errors moves perceived accuracy more than its error rate does, what makes an error boundary learnable, and why onboarding is the cheapest intervention available and the first one to go stale.

Two versions of the same meeting-detection assistant, both exactly 50% accurate, produce different products. Kocielnik, Amershi and Bennett built them: one tuned to avoid false positives, so it misses meeting requests, and one tuned to avoid false negatives, so it over-detects. At identical accuracy, the version that erred toward over-detection was perceived as more accurate and more acceptable, and the authors attribute the gap to the cost of recovering from each error type in that task, a glance at a wrong suggestion against a missed meeting (Kocielnik, Amershi & Bennett, 2019, Will You Accept an Imperfect AI?, CHI '19).

The number on the model card did not move. The product did. This is the starting point for expectation setting: users do not experience an error rate, they experience a particular shape of being let down, and they form that impression in the first few minutes.

What the user is actually learning

The useful way to describe a user's mental model of an AI feature is not "how it works" but a decision rule: given this input, will the system get it right? Bansal and colleagues call the thing being learned the system's error boundary, and identify what makes it learnable: parsimony, whether the boundary can be described with few features, and stochasticity, whether identical-looking inputs sometimes succeed and sometimes fail. Task dimensionality sets how much there is to learn (Bansal et al., 2019, Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance, HCOMP 2019).

This reframes a familiar complaint. A system at 80% accuracy with a parsimonious, near-deterministic boundary ("it fails on handwriting and on tables") is a tool: the user routes around the failure and keeps the gains. A system at 90% whose failures are scattered and irreproducible is a lottery, because there is nothing to learn and therefore no way to calibrate when to check. Raising aggregate accuracy while increasing stochasticity can make a feature worse to use, and no dashboard will show it.

Kocielnik's three interventions all attack the same gap by different routes: an accuracy indicator states the expected rate outright, an example-based explanation shows which inputs the detector responds to, and a control lets the user move the detection threshold. Their effect was largest in the condition users liked least, which is the useful property: expectation setting buys the most where the raw experience is worst. Amershi and colleagues made the two halves of this their first two guidelines, making clear what the system can do and how well it can do it (Amershi et al., 2019, CHI '19).

Designing the first five minutes

State the failure mode, not the accuracy. "Accurate about 85% of the time" is unactionable; "reliable on typed text, unreliable on handwriting and on scanned tables" hands the user the boundary.

Lead with an input the system fails on. Onboarding built from successes teaches the ceiling. One honest failure, shown early and cheaply, teaches the shape of the boundary and costs a user nothing, because they did not depend on that answer yet.

Choose the error direction deliberately, and say which you chose. A system that over-suggests and a system that under-suggests are different promises. Users can work with either once they know which they have.

Make the first correction the easiest action on screen. The first repair is where the mental model forms, and a user who discovers that fixing the output is trivial reads subsequent errors as cheap.

When it breaks

The boundary moves and nobody reships the onboarding. A model update that changes which inputs fail invalidates everything the user learned, silently. The user's model is now wrong in the worst way: confidently wrong, in the direction of trusting what used to work.

Onboarding is read once, at the moment of least interest. Expectation setting placed in a first-run modal is competing with the user's actual task. The claims that survive are the ones restated at the point of use.

Stated accuracy is the wrong number anyway. A single rate averages over subpopulations the user does not belong to. A user whose inputs are all handwriting experiences a 40% system that advertised 85%, and correctly concludes the advertisement was dishonest.

Lowered expectations lower use. The intervention that makes an imperfect system acceptable also tells some users not to bother. That tradeoff is real, it is rarely measured, and leaving expectations inflated is not a fix for it; it postpones the same loss to a worse moment.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Kocielnik, Amershi & Bennett, 2019, Will You Accept an Imperfect AI?, CHI '19 microsoft.com
  2. Bansal et al., 2019, Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance, HCOMP 2019 ojs.aaai.org
  3. Amershi et al., 2019, CHI '19 microsoft.com
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track