Reliability & Consequence 7 October 2026 8 min read 1,770 words

Only the reload was wrong

Seven of the fixes in LangGraph 1.2.13 land in the same room: the code that rebuilds a thread's state out of its own past. One of them explains the whole class of bug in six words, and they are the six words an architect should find least comfortable.

The argument

LangGraph reconstructs a thread's state by walking a history that operators can edit and runs can fork, so an agent's state is a query rather than a record — and because live execution stays correct, every failure of that query arrives only after a resume.

The most consequential sentence published about agent architecture this fortnight is not in a specification or a blog post. It is in the description of a bug fix, and it reads: "Live execution was correct; only the reload was wrong."

That is from pull request #8548, merged into LangGraph on 3 October. It arrived in version 1.2.13 on 5 October alongside six other fixes which, once the dependency bumps are stripped out, are the entire release. They are worth listing by their own titles, because the pattern is the argument. Fork before replaying an update checkpoint the thread moved past. Keep an update_state on an older checkpoint out of its other branches. Keep DeltaChannel counters on every update_state path. Don't walk history for a DeltaChannel that was never written. Don't replay an abandoned branch into a DeltaChannel fork. Hydrate subgraph delta channels with the caller-resolved saver. Stop reporting answered interrupts in get_state.

Seven fixes, one room. Every one of them is in the code that rebuilds a thread's state out of its own past.

This is the claim worth making from it: an agent running on a checkpointer does not keep its state. It keeps a history and computes the state on demand, and the history is one that operators are invited to edit and runs are permitted to fork. The state of a long-running agent is therefore a query result, not a record — and the particular cruelty of that arrangement is that the query is only evaluated when someone comes back, so a wrong answer cannot be observed at the moment it is created.

This section has looked at agent memory before, from the side of what gets thrown away: compaction, and the fact that two major frameworks disagree about whether summarising a transcript destroys it or merely hides it. This is the other half of the same problem and it gets far less attention. Not what the runtime forgets, but what it computes when you ask it to remember.

Start with the mechanism, because it is unusually candid in its own source. Everything that follows is read from the projects' own repositories — the release, the pull requests, the source files and the docstrings in them — which is both the strength and the limit of this reading: it is the record of the people who shipped the code, with no independent account of what it did in anyone's production.

The relevant object is DeltaChannel, and its docstring describes it as a "Reducer channel that stores only a sentinel in checkpoint blobs and reconstructs state by replaying ancestor writes through the reducer." The checkpoint holds a placeholder. The value is produced by folding the writes recorded against every ancestor checkpoint through a function, on read. The saver method that supplies the raw material is get_delta_channel_history, whose own docstring is similarly plain: "Walk the parent chain returning per-channel writes + seed."

Replay depth is bounded by two counters held in checkpoint metadata under counters_since_delta_snapshot. A full snapshot is written when a channel reaches snapshot_frequency updates — the default is 1,000 — or when the thread reaches a system-wide bound of 5,000 supersteps since the last snapshot, overridable by an environment variable. So in the ordinary case, reading one field of an agent's state can mean folding up to a thousand recorded writes, and the project has thought carefully about capping that number, which tells you the walk is expected to be long.

Then there is the question of who supplies the fold. The reducer is application code, and the contract placed on it is explicit: reducers "must be deterministic and batching-invariant (associative across folds)", with the required algebra written out as reducer(reducer(state, xs), ys) == reducer(state, xs + ys). The stated reason is operational rather than theoretical — the property "lets LangGraph replay checkpointed writes in larger batches than they were originally produced without changing reconstructed state." Nothing enforces it. A reducer that appends a timestamp, deduplicates against a set it built in a different order, or is merely non-associative will reconstruct a state that the run never held, and it will do so silently.

What makes this more than a conventional event-sourcing trade-off is the second half: the history is not append-only. update_state exists precisely to change it, and its docstring says what the change looks like from the inside — the values are recorded "as if they came from node as_node." Edit an earlier checkpoint, resume from it, and you have forked the thread. This is not a loophole. It is a selling feature, the basis of time travel and of every human-in-the-loop review flow where a person corrects an agent and lets it continue.

Combine the two and the October fixes stop being seven bugs and become one. In #9170, editing a checkpoint, continuing past it, then replaying that edit caused "branches to receive writes they never executed", because the writes were stored against the original checkpoint before the fork was decided; the fix moves the fork decision ahead of the storage and adds a shared checkpoint_superseded helper to ask whether a checkpoint is still the latest in its thread. In #8548 the shared base of a fork kept the pending writes of the abandoned branch, and "nothing recorded which child consumed which write", so the ancestor walk swept them up on reload. The pull request gives the arithmetic: a fork that returned ['in-1', 'first-out', 'in-3', 'third-out'] reloaded as ['in-1', 'first-out', 'in-2', 'in-3', 'third-out']. The in-2 is from the branch that was thrown away. The agent never produced it, and after a restart it was part of the agent's state.

The sharpest of the three is #9103, merged 2 October, because it is the one that touches governance rather than data. When parallel tasks interrupted and one was resumed, its old interrupt write stayed in the open superstep, and get_state went on treating it as pending — which made, in the maintainers' words, "the answered question to reappear in the state." An interrupt is how LangGraph asks a human to decide something. get_state is what a review console displays and what an audit trail reads. For the duration of that bug, the system's own answer to "is anyone still waiting on a person?" could be yes when the person had already answered. The fix is a tightening of what counts as finished: an output write now marks a task complete, a RESUME control write alone does not, and Command(resume=value) without an interrupt identifier raises an error when several interrupts are pending.

And then #9142, which is the detail I would put in front of anyone who thinks this is all beta-grade trivia. Three update_state paths were saving their checkpoint without counters_since_delta_snapshot, which reset every delta channel's snapshot cadence to zero. The counters decide when a snapshot is written; losing them means the next snapshot arrives later than configured and, as the fix notes, the replay walks "kept growing". The operator's intervention — the correction, the approval, the helpful nudge — was quietly lengthening the history that every subsequent read had to traverse.

The strongest case against reading any of this as a design problem is a good one, and it is roughly this. Storing a full copy of a large state at every superstep is untenable for threads that run for days; folding deltas is the standard answer and snapshots bound the cost. Event sourcing is a mature pattern with a literature. DeltaChannel is labelled beta in its own docstring, which warns that the surrounding contract "is not yet stable"; get_delta_channel_history tells implementers to "override at your own risk". The default MessagesState still uses add_messages rather than a delta channel, and the batch reducer offered for messages is marked experimental and explicitly "not full add_messages parity". A patch release that fixes seven history-walking edge cases is not a project in trouble. It is a project doing the work.

All of that is true, and two of the seven fixes are not delta-specific at all — the fork-leakage bug and the answered-interrupt bug sit in core checkpoint semantics that anyone using a checkpointer with a human in the loop depends on today. More importantly, the defence concedes the premise. Event sourcing earns its guarantees from two things: a log nobody rewrites, and a projection that is a pure function of it. Temporal's documentation is unembarrassed about the first, describing Event History as "a complete and durable log of everything that has happened in the lifecycle of a Workflow Execution", from which a worker can "recreate the state of the Workflow Execution to what it was immediately before the crash" — as if the failure never occurred. That is the deal. You give up the right to edit the past, and you accept versioning discipline on the code that replays it, and in exchange the reconstruction is trustworthy.

An agent framework cannot make that deal, because editing the past is what its users came for. It offers the reconstruction without the immutability, and it places the projection in application code under an algebraic requirement it has no way to check. The result is event sourcing's machinery carrying none of event sourcing's promises, filed in most organisations under the heading of which checkpointer to install.

That is the part worth taking away, and it is not really about one library. Choosing a checkpointer looks like an infrastructure decision — Postgres or SQLite, managed or self-hosted — and is actually a choice of consistency model for the agent's state, plus a decision about where the projection function lives and who is allowed to rewrite the inputs. I have not seen that decision recorded in an architecture decision record. It is recorded in a dependency file.

Which brings the six words back round. "Live execution was correct; only the reload was wrong" is the most reassuring sentence in the release and the most alarming one. Everything the agent did while anyone was watching was right. The tests passed, because tests watch execution. The defect surfaces a week later, on a thread someone edited, in a value nobody can trace to a decision — and correctness at the moment of action turns out to be the one property that cannot tell you whether the record will still agree with it tomorrow.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. langgraph releases feed, versions 1.2.12 to 1.2.14 LangChain · 2026-10-06
  2. fix(langgraph), fork before replaying an update checkpoint the thread moved past (#9170) LangChain · 2026-10-05
  3. fix(langgraph), don't replay an abandoned branch into a DeltaChannel fork (#8548) LangChain · 2026-10-03
  4. fix(langgraph), stop reporting answered interrupts in get_state (#9103) LangChain · 2026-10-02
  5. fix(langgraph), keep DeltaChannel counters on every update_state path (#9142) LangChain · 2026-10-05
  6. DeltaChannel source and contract LangChain · 2026-10-07
  7. BaseCheckpointSaver.get_delta_channel_history and checkpoint metadata LangChain · 2026-10-07
  8. Pregel.update_state LangChain · 2026-10-07
  9. Temporal Event History and Replay Temporal · 2026-10-07

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

agentscheckpointingevent sourcingreplayaudit