When Chat Is the Wrong Interface
Why a text box is the most expressive and least specifiable interface ever shipped, the expressiveness-versus-adjustability tradeoff that explains when to replace it with controls, and what generated UI costs in learnability.
A user wants a summary, but shorter, and in the second person, and without the bit about pricing. In a chat interface they type that sentence, wait, read, and type another sentence. In a word processor they would have dragged a slider and unticked a box. The text box accepted a request no graphical interface could have expressed; it then made the three small adjustments that followed more expensive than they are anywhere else in computing.
That asymmetry is the whole design question, and it predates language models. Shneiderman's case for direct manipulation rested on continuous representation of the object, physical action instead of typed syntax, and immediately visible results (Shneiderman, 1983, Direct Manipulation: A Step Beyond Programming Languages, IEEE Computer 16(8)). A prompt has none of the three. Horvitz's 1999 paper opens on exactly this standoff, between researchers enhancing direct manipulation and researchers building interface agents, and argues the answer is a coupling rather than a winner (Horvitz, 1999, CHI '99).
Expressiveness and specifiability trade against each other
Score an interface on two axes. Expressiveness is the size of the set of requests it can state at all. Specifiability is how cheaply a user can state exactly the one they mean, and restate it slightly differently.
A text box maximises the first and is close to worst on the second. The request space is unbounded, and every point in it costs a sentence to reach, with no visible indication of which sentences the system responds to. A toolbar inverts both. The evidence that this is a real cost, and not a preference for widgets, is what happens when non-experts are asked to prompt. Zamfirescu-Pereira and colleagues gave ten participants without prompt-design experience a chatbot-design tool and watched them work. They explored opportunistically rather than systematically, none of them ran more than one or two conversations before jumping in to fix things, they over-generalised from single observations, and they carried expectations from instructing humans into instructing a model (Zamfirescu-Pereira, Wong & Yang, 2023, Why Johnny Can't Prompt, CHI '23). The failures are the ones end-user programming research has catalogued for decades, which is the tell: a prompt is a program in an interface that pretends it is a conversation.
Three ways out, each with a bill
Reify the request into controls. Malleable Prompting turns preference expressions found in a natural-language prompt into sliders, dropdowns and toggles, and highlights the output spans a given control influenced. Participants hit target preferences more precisely and rated it more controllable and transparent than prompting alone (Zhang et al., 2026, From Words to Widgets for Controllable LLM Generation, UIST '26, arXiv:2604.10925). The cost is that the widget set is derived from the prompt, so the controls change as the request does.
Generate the interface per query. Chen and colleagues had the model emit a task-specific UI instead of prose, and report up to a 72% improvement in human preference over chat across their task set (Chen et al., 2025, Generative Interfaces for Language Models, arXiv:2508.19227). Read that as the authors' measurement of their own system on their own framework, which is how such numbers should be read; the direction is corroborated by the control-reification work above.
Keep chat for intent, controls for adjustment. The text box states the goal once. Everything after it is bounded, so it gets handles. This is the coupling Horvitz argued for, and it is what most shipped assistants converge on.
When it breaks
Generated UI has no stable layout. Learnability comes from the control being in the same place tomorrow. An interface composed fresh for each query gives up muscle memory, keyboard paths, and the user's ability to predict what is available before asking. A dashboard that rearranges itself is not more usable for being more apt.
Controls imply a bounded parameter space that may not exist. A length slider is honest. A "formality" slider over an open text space suggests a monotone dimension the model does not actually have, and users discover the gap only after trusting it.
Two code paths, two sets of bugs. A hybrid interface has to keep the conversational state and the control state agreed. Users will set a control, then contradict it in prose, and expect the later instruction to win.
Chat is the right answer more often than toolbar nostalgia admits. When the request space is genuinely open, when the user cannot name the operation they want, or when the task happens once, the sentence is cheaper than learning any interface. The question is never whether chat is good; it is whether the next twenty interactions are adjustments.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Shneiderman, 1983, Direct Manipulation: A Step Beyond Programming Languages, IEEE Computer 16(8) cs.umd.edu
- Horvitz, 1999, CHI '99 microsoft.com
- Zamfirescu-Pereira, Wong & Yang, 2023, Why Johnny Can't Prompt, CHI '23 dl.acm.org
- Zhang et al., 2026, From Words to Widgets for Controllable LLM Generation, UIST '26, arXiv:2604.10925 arxiv.org
- Chen et al., 2025, Generative Interfaces for Language Models, arXiv:2508.19227 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.