Interaction Design for AI intermediate 7 min read 7 flashcards

Accessibility in Streaming and Generated Interfaces

Why token streaming fights the assistive-technology mechanism meant to announce it, what a runtime-generated interface does to the possibility of an accessibility review, and where AI output helps and costs disabled users their own voice.

The two interaction patterns that define AI products both break assumptions that accessibility practice depends on. Token streaming produces continuous, unbounded DOM change where assistive technology expects discrete status messages. Runtime-generated interfaces produce markup that no human reviewed before a user met it. Neither is a rendering detail; both decide whether a feature exists for some of its users.

The baseline is not good. In the 2025 run of the WebAIM Million, 94.8% of one million home pages had at least one detectable WCAG 2 failure, averaging 51 errors per page, down from 56.8 the year before (WebAIM, 2025, The WebAIM Million). These are automatically detectable failures only, so the figure is a floor. A new interaction paradigm arriving on top of that baseline does not start from zero.

Streaming against the status-message mechanism

WCAG 2.2, a W3C Recommendation since October 2023, requires at Level AA that status messages be programmatically determinable so assistive technology can present them without moving focus (W3C, 2023, Web Content Accessibility Guidelines 2.2, SC 4.1.3). In practice that means an ARIA live region: a container marked aria-live="polite" or role="status" whose changes a screen reader announces.

Live regions were specified for bounded, occasional updates, such as "item added to basket". A streamed response is the opposite: text arriving in hundreds of increments over ten to thirty seconds. The two available politeness settings fail in opposite directions. With aria-atomic="true" the region is re-announced whole on every change, so a long response is read back as a lengthening prefix, over and over. With aria-atomic="false" only the addition is announced, and high-frequency changes get coalesced or dropped by the screen reader, which differs by product, so the user hears a fragment of the answer and cannot tell which fragment.

The pattern that works gives up the visual metaphor. Let the text render silently while it streams, expose an accessible busy state so the user can distinguish generating from crashed, and announce the completed message once, when it is coherent. A screen-reader user gains nothing from hearing a sentence assembled a word at a time, and the typing animation that reads as liveness to one user is noise to another.

Generated interfaces and the review that no longer happens

An accessibility review is a pre-release activity performed on artefacts. A UI composed by a model at request time has no pre-release. Label quality, focus order, contrast, keyboard reachability and the correctness of every ARIA attribute become properties of a generation, sampled fresh for each user, and the usual remedy of auditing and fixing the page does not apply because the page does not persist. The honest options are to constrain generation to a vetted component library whose accessibility is a property of the components, or to validate the generated tree before it renders and fall back to plain text when it fails. Neither is free, and "the model usually gets it right" is the same reasoning that produced a 94.8% failure rate with hand-written markup.

Where the output itself helps, and what it costs

Generative models can remove real barriers. GenAssist made text-to-image generation usable for blind and low-vision creators by letting them verify whether a candidate followed the prompt, surfacing detail the prompt never specified, and summarising similarities and differences across candidates, built from a language model generating visual questions and vision-language models answering them (Huh, Peng & Pavel, 2023, GenAssist: Making Image Generation Accessible, UIST '23). The design insight generalises: for a user who cannot inspect the artefact, comparison across candidates is the accessible operation.

The cost shows up where output speaks for the user. Valencia and colleagues studied twelve AAC users with live model suggestions across three scenarios, and found participants did expect suggestions to save time and physical and cognitive effort, while insisting that the phrases reflect their own communication style and preferences (Valencia et al., 2023, "The less I type, the better", CHI '23). A suggestion that is faster and not yours is a trade the designer made on the user's behalf, and the voice being substituted belongs to someone who has fewer alternatives than most users do.

When it breaks

Alt text from the model that produced the image. A system describing its own output to a user who cannot see it has no independent check, and its errors are invisible by construction, exactly where they matter most.

Voice interfaces exclude non-standard speech. Recognition trained on typical speech fails on dysarthria and on accented and atonal speech. A voice-first design decision silently removes the users most likely to need hands-free input.

Latency affordances assume vision. Streaming, skeleton loaders and progress shimmer all communicate through the visual channel. Without an equivalent announced state, a long generation is indistinguishable from failure.

Treating this as a compliance item. Both problems above are interaction-design decisions, taken when streaming and generated UI were chosen. An audit at the end can find the missing label; it cannot undo a pattern whose core mechanism fights the assistive technology.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. WebAIM, 2025, The WebAIM Million webaim.org
  2. W3C, 2023, Web Content Accessibility Guidelines 2.2, SC 4.1.3 w3.org
  3. Huh, Peng & Pavel, 2023, GenAssist: Making Image Generation Accessible, UIST '23 dl.acm.org
  4. Valencia et al., 2023, "The less I type, the better", CHI '23 research.google
Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track