Interaction Design for AI advanced 8 min read 6 flashcards

Branching, Comparison and the Linear Transcript

Why a chat transcript makes exploration destructive and comparison expensive, what the research interfaces that break linearity actually provide, and the quadratic comparison cost that stops users exploring long before the model runs out of ideas.

A model can produce a thousand plausible answers to one request. The interface shows one, and the next thing the user types overwrites the state that produced it. Exploration in a chat transcript is destructive: to try a different framing you abandon the current one, and to get back you reconstruct it from memory. Everything about how people use these systems follows from that, including the finding that keeps recurring in the studies.

Arawjo and colleagues built ChainForge to let people compare responses across models and prompt variations in a graph rather than a thread, and in characterising how people prompt and test identified three modes: opportunistic exploration, limited evaluation, and iterative refinement (Arawjo et al., 2024, ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing, CHI '24, arXiv:2309.09128). "Limited evaluation" is the diagnosis. People look at a handful of outputs and stop, which is also what the prompting study with non-experts found when none of its ten participants ran more than one or two conversations before intervening (Zamfirescu-Pereira, Wong & Yang, 2023, Why Johnny Can't Prompt, CHI '23).

The cost that stops exploration

The limit is not generation, which is cheap and parallel. It is reading. Suppose four prompt variants, three models, and five test inputs: \(4 \times 3 \times 5 = 60\) outputs. At twenty seconds to read one and form a judgement, that is twenty minutes of attention for a single comparison, before any of it is written down. Pairwise comparison is worse. Ranking \(n\) candidates on one criterion by pairwise reads is \(\binom{n}{2}\) comparisons, so eight candidates is 28 reads, and twelve is 66.

Users respond rationally by reading three and choosing. The output they keep is then a sample from a distribution they never saw, and they have no idea whether it was near the top of it. This is the real failure of the linear transcript: not that it hides alternatives, but that it makes the cost of looking at them feel like the cost of generating them, which it is not.

Three research directions attack different parts of the cost. Luminate asks the model to produce the design space rather than individual artefacts, extracting dimensions from outputs and giving the user structured handles to navigate them, on the argument that current interaction paradigms push users to converge on a few ideas instead of exploring the latent space (Suh et al., 2024, Luminate, CHI '24, arXiv:2310.12953). Sensecape externalises levels of abstraction so a user can move between an overview and the detail, and organise findings at whichever level the task needs (Suh et al., 2023, Sensecape, UIST '23, arXiv:2305.11483). Graphologue converts a response into an interactive node-link diagram in real time, extracting entities and relations, and reports F-scores of 97.24% for entity annotation and 92.39% for relationship annotation after one round of correction (Jiang, Rayan, Dow & Xia, 2023, Graphologue, UIST '23, arXiv:2305.11473).

What to build into an ordinary product

Make a branch cost one click and make it visible. Editing an earlier message should fork rather than overwrite, and the fork should be reachable. Every system that supports this discovers users were already doing it by hand, badly, in separate tabs.

Show the axis, not the alternatives. Four outputs side by side is a reading task. "These differ mainly in formality and length" is a decision. Extracting the dimension of variation is the move Luminate generalises, and it is the only one that reduces reading rather than relocating it.

Give the comparison a criterion before showing the candidates. A user who has not said what "better" means will pick on surface fluency. Ask for the criterion, then order candidates by it, and the quadratic cost collapses to a scan.

Let a branch be named and closed. Exploration without a way to discard a path turns into an unsearchable archive of near-identical attempts.

When it breaks

Branch proliferation is its own failure mode. A tree with forty nodes and no labels is less navigable than a linear transcript, because at least the transcript has an order. Structure helps only with pruning attached.

The criterion often does not exist yet. Early exploration is how users discover what they want, so demanding an evaluation criterion up front can block the work it was meant to support. The sequencing matters: explore loosely, then compare with a criterion, and do not ask for one in the first minute.

Structured exploration costs tokens and latency. Sixty outputs is sixty generations. For long contexts and reasoning models, the spend on options nobody reads is real, and a comparison interface that is free to the user is not free.

Automated comparison inherits a judge. Scoring candidates with a model to spare the user the reading replaces one unmeasured preference with another, and a user who trusts the ordering is trusting a judge whose error profile they have never seen. See /learn/eval-driven-development for what it takes to make such a judge load-bearing.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Arawjo et al., 2024, ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing, CHI '24, arXiv:2309.09128 arxiv.org
  2. Zamfirescu-Pereira, Wong & Yang, 2023, Why Johnny Can't Prompt, CHI '23 dl.acm.org
  3. Suh et al., 2024, Luminate, CHI '24, arXiv:2310.12953 arxiv.org
  4. Suh et al., 2023, Sensecape, UIST '23, arXiv:2305.11483 arxiv.org
  5. Jiang, Rayan, Dow & Xia, 2023, Graphologue, UIST '23, arXiv:2305.11473 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track