LLM Application Architecture intermediate 7 min read 6 flashcards

Conversation Branching and the Message Graph

Why editing a turn or regenerating a reply turns a conversation into a tree rather than a list, what a branch invalidates downstream, and the storage and accounting shapes that follow.

A user regenerates an answer, edits a question two turns back, switches model and asks again. On screen this is a tidy transcript. Underneath it is a tree with four leaves, and an application whose schema says messages: Message[] has already lost three of them. Branching is not an advanced feature bolted on for power users; it is what "regenerate" and "edit" mean, and almost every chat product ships both on day one.

The data structure is a tree with a cursor

Give every message a parent id. Siblings are alternatives: a regenerated reply is a second child of the same user turn, an edited question is a second child of the preceding assistant turn. The conversation the user sees is the path from the root to a selected leaf, and the "2 / 3" control in the corner of a chat UI is simply sibling navigation. Storage is append-only, with mutation expressed as a new child and deletion as a tombstone, plus one mutable pointer per thread naming the active leaf.

Providers meet this in two different ways, and the difference matters. With a stateless messages endpoint you resend the chosen path on every turn, so branching is free on the provider side and entirely your problem on yours. With a chained endpoint, such as the Responses API's previous_response_id, the provider holds the history and you branch by referencing the same previous_response_id from two different calls, which produces two independent continuations of one point (OpenAI, Conversation state). Convenient, and it means the branch topology now lives partly in someone else's system.

A branch invalidates more than it replaces

The tempting assumption is that editing turn three leaves turns four onward reusable. It does not, for three separate reasons.

Signed model artefacts are bound to their prefix. On recent Claude models a replayed thinking block is accepted only while the system prompt, tools and preceding messages are unchanged; otherwise the request fails with a 400 whose message reads "The block is bound to a different conversation", with an opt-in prefix_mismatch_behavior of drop_block as the alternative to an error (Anthropic, Claude API errors). Anthropic's own advice is to keep conversation history append-only. An editing feature therefore has to decide, per branch, whether to drop reasoning blocks or re-run the turns that produced them.

The prompt cache invalidates from the edit point. Cache hits require byte-identical prefixes, so a branch that changes an early message pays full price for everything after it, and the minimum cacheable segment is 512 to 4,096 tokens depending on model, which means short branches may not re-establish a cache at all (Anthropic, Prompt caching).

Derived state does not fork. Extracted structured state, memory writes and tool side effects were produced on the path the user has now abandoned. The message graph can be rewound; a sent email cannot. Anything irreversible needs either a compensating action or a rule that it only fires on the accepted path, which in practice means an approval step between generation and effect.

Cost and storage take a different shape

Storage grows with exploration rather than with conversation length, and the heavy rows are retrieved documents and tool outputs, not user prose. Store those once, content-addressed, and reference them from each branch that used them; a 40 KB search result duplicated across five regenerations is 200 KB of identical bytes.

Accounting splits in two. Token usage attaches to a request, including requests on branches nobody kept. Cost per conversation, the number a product manager wants, attaches to the path the user accepted plus the cost of everything they discarded getting there. Reporting only the second understates spend; reporting only the first makes every regeneration look like a new conversation.

When it breaks

Two tabs, one thread. Concurrent turns create two children of the same parent, and last-writer-wins on the leaf pointer silently abandons one of them. Make the pointer update a compare-and-set on the expected parent, and surface the losing branch instead of dropping it.

Eval sets oversample rejected output. Traces sampled uniformly are biased toward abandoned branches, because users regenerate when an answer is bad. Sampling from accepted paths, with abandoned siblings kept as negatives, is the version worth training or grading on.

"What did the user actually see?" has no answer without history on the leaf pointer. For audit, support and incident review, record pointer moves, not just messages.

Tool results get recomputed per branch. Key the tool cache on the call, as a hash of tool name and arguments, rather than on the message that triggered it, and a branch reuses the work.

Compaction and branching interact badly. A summarised prefix is a derived artefact of one path; branching behind the summary boundary either invalidates it or quietly attributes one branch's history to another. See context compaction and handoff and where state lives in an LLM application.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. OpenAI, Conversation state developers.openai.com
  2. Anthropic, Claude API errors platform.claude.com
  3. Anthropic, Prompt caching platform.claude.com
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track