Interaction Design for AI advanced 8 min read 6 flashcards

Reasoning Traces as an Interface Surface

Why the thinking text shown next to an answer is a progress indicator and a steering handle rather than an explanation, how providers split the trace into reasoning and progress updates, and what the faithfulness results mean for anyone rendering it.

Something new appeared in interfaces around 2024: a panel that shows the model thinking. It is read as an explanation, and it is not one. It is also the most honest progress indicator the field has produced, and a surprisingly good place to interrupt. Getting the design right starts with knowing which of those three things is on screen, because the APIs now distinguish them explicitly and the research says only one of the claims holds.

Three artefacts, not one

The raw chain of thought is the model's own working text. The industry does not ship it. Anthropic's API returns a summary of the thinking, produced by a different model from the one serving the request, and states plainly that "No display setting returns the raw chain of thought"; the stated rationale is that summarised thinking "provides the full intelligence benefits of thinking while preventing misuse", and that summarisation adds little latency so it can stream. You are billed for the full thinking tokens, not the summary you see (Anthropic, Thinking, Claude Platform docs). OpenAI took the same position when o1 shipped, showing users a summary and keeping the raw trace internal.

The third artefact is the interesting one for designers. The same documentation describes a progress update: a short note between tool calls about what the model just found and what it will do next, "written for the person watching the agent rather than as reasoning", returned as its own block with its own signature, and exposed through a display: "updates" mode in which reasoning stays empty and only these notes come back. The split is an admission in the API surface that a status line and a reasoning trace are different products with different audiences, and that most agent interfaces want the first.

Why it is not an explanation

The trace does not reliably report the computation that produced the answer. Lanham and colleagues intervened on chains of thought by adding mistakes and paraphrasing, and found wide variation across tasks in how much the answer actually depended on the stated reasoning, with larger and more capable models producing less faithful reasoning on most tasks they studied (Lanham et al., 2023, Measuring Faithfulness in Chain-of-Thought Reasoning, arXiv:2307.13702). The direct test for interfaces is sharper: plant a hint in the prompt, let the model use it, and see whether the trace mentions it. Across six hint types and current reasoning models, Chen and colleagues found reveal rates frequently below 20% (Chen et al., 2025, Reasoning Models Don't Always Say What They Think, arXiv:2505.05410).

A panel that is right about its own causes less than a fifth of the time cannot carry the weight of an explanation. It can still carry weight as evidence of intent, which is the narrower claim a cross-organisation position paper makes for chain-of-thought monitoring, while warning that the property is fragile and not guaranteed to survive changes in training practice or a move to latent reasoning (Korbak et al., 2025, Chain of Thought Monitorability, arXiv:2507.11473). For how explanations affect reliance, which is a separate literature with its own uncomfortable results, see /learn/explanations-and-their-effect-on-reliance.

What the surface is good for

Legibility over long work. A spinner for ninety seconds is indistinguishable from a crash. A named step is not. This is the one job the trace does well, and it is why progress updates exist as a separate block type.

A steering handle. A user who reads "searching for the 2024 filing" at second three can stop the run and say "2023" before the model spends ninety seconds being wrong. Interrupting early is worth more than correcting late, and the trace is the only place the user learns early enough to do it.

Debugging, for the people who build the thing. The verbose preamble that some models emit at the start of thinking is useful for prompt engineering, and the docs say as much. That is a developer surface, and shipping it to end users is a category error.

When it breaks

Users read the trace as a commitment. "I will check the invoice total" reads as a promise. It is a sampled intention, and the final answer may not reflect it. Interfaces that render the trace in the same typography as the answer invite exactly this conflation.

The summary is a paraphrase by a third party. A different model wrote the text on screen, and it did not see its own output. Quoting it back to the user as the model's words, or worse, storing it as an audit record of why a decision was made, misrepresents what it is.

Length becomes a quality signal. Users and dashboards both learn that more visible thinking means more effort means a better answer. It means more billed tokens. The correlation with answer quality is task-dependent and not something to train users to rely on.

Exposure changes the trace. Once a trace is user-facing, there is pressure to make it presentable, and the optimisation that makes it presentable is the one Korbak and colleagues warn erodes its value as a monitoring signal. An interface decision here has a safety cost that lands somewhere else entirely.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Anthropic, Thinking, Claude Platform docs platform.claude.com
  2. Lanham et al., 2023, Measuring Faithfulness in Chain-of-Thought Reasoning, arXiv:2307.13702 arxiv.org
  3. Chen et al., 2025, Reasoning Models Don't Always Say What They Think, arXiv:2505.05410 arxiv.org
  4. Korbak et al., 2025, Chain of Thought Monitorability, arXiv:2507.11473 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track