Energy-Based & Score Models advanced 9 min read 7 flashcards

Modern Hopfield Networks as Attention

How the log-sum-exp energy over continuous states gives an update rule identical to transformer attention, what the three kinds of fixed point mean for a real network, and where the equivalence stops holding.

In 2020 a group at JKU Linz wrote down an energy function for continuous-valued associative memory, derived its single-step update rule, and observed that the rule was already in production: it is the attention operation of a transformer (Ramsauer et al., 2020, Hopfield Networks is All You Need, ICLR 2021, arXiv:2008.02217). The claim is narrow and exact, which is what makes it useful. It is not that transformers resemble memories; it is that one line of a transformer is one step of retrieval in a dense associative memory with exponential storage capacity.

The energy and its update rule

Store \(M\) patterns as the columns of \(X \in \mathbb{R}^{d \times M}\) and let the state \(\xi\) be a continuous vector in \(\mathbb{R}^d\). Define

\[E(\xi) = -\operatorname{lse}\!\left(\beta, X^{\!\top}\xi\right) + \tfrac{1}{2}\xi^{\!\top}\xi + \beta^{-1}\log M + \tfrac{1}{2}M_{\max}^2\]

where \(\operatorname{lse}(\beta, z) = \beta^{-1}\log \sum_i e^{\beta z_i}\) and \(M_{\max}\) is the largest pattern norm. The first term is the exponential interaction function of dense associative memory in its numerically safe form; the quadratic term keeps the state bounded. Minimising this energy by the concave-convex procedure gives an update that converges in one step for well-separated patterns:

\[\xi^{\text{new}} = X\,\operatorname{softmax}\!\left(\beta X^{\!\top}\xi\right)\]

Put \(Q\) in place of \(\xi\), \(K\) in place of \(X\) on the similarity side and \(V\) on the projection side, and set \(\beta = 1/\sqrt{d_k}\), and this is \(\operatorname{softmax}(QK^{\!\top}/\sqrt{d_k})V\). The network retrieves patterns with one update and exponentially small retrieval error, and its capacity is exponential in the dimension of the associative space. The same identification was published in parallel as the mechanism behind a working method for immune repertoire classification, where transformer attention is used explicitly as the update rule of a modern Hopfield network with exponential storage capacity (Widrich et al., 2020, Modern Hopfield Networks and Attention for Immune Repertoire Classification, NeurIPS).

Three kinds of fixed point

The energy has three types of stationary point, and the distinction is the practically useful part of the theory. A fixed point can store a single pattern, which is retrieval in the ordinary sense. It can be a metastable state that averages over a subset of similar patterns. Or it can be the global fixed point that averages over all of them, which is what a uniform softmax produces.

Which regime you are in is set by \(\beta\) and by how well separated the patterns are, and Ramsauer et al. report that trained models use both: transformer and BERT models operate in their first layers preferentially in the global averaging regime and in higher layers in metastable states. That is a testable statement about real networks derived from an energy function, and it reframes a low-entropy attention head as a memory read and a high-entropy one as a pooling operation rather than a broken one.

The same framing explains why \(\beta\) is load-bearing rather than cosmetic. Too small, and every query returns the mean of the values, a failure that has been given its own name in the long-context literature: the first-token attention sink has been argued to be a mechanism that lets language models avoid over-mixing, with context length, depth and data packing all influencing how strongly it appears (Barbero et al., 2025, Why do LLMs attend to the first token?, arXiv:2504.02732). Too large, and retrieval becomes a hard nearest-neighbour lookup that a slightly corrupted query gets wrong.

Variants that fix the softmax's weaknesses

Softmax never returns exactly zero, so every retrieval mixes in a little of every pattern. Replacing it with a sparse transformation derived from Fenchel-Young losses gives a family of Hopfield energies whose updates are sparse, with an explicit link between sparsity and exact memory retrieval, and a structured extension that retrieves associations of patterns rather than single patterns (Santos, Niculae, McNamee and Martins, 2024, Sparse and Structured Hopfield Networks, ICML, PMLR 235). A different direction keeps the energy and rebuilds the architecture around it: the Energy Transformer stacks layers that each lower a designed energy over tokens, which makes its attention deliberately different from conventional attention (Hoover et al., 2023, Energy Transformer, NeurIPS, arXiv:2302.07253).

When it breaks

A transformer layer is not energy descent. The equivalence requires the similarity to be symmetric in the pattern matrix. Real attention uses separate learned projections, so \(W_Q \neq W_K\) in general, and the asymmetry removes the Lyapunov argument that made Hopfield dynamics converge. One attention step is a retrieval step; a stack of attention layers is not minimising anything you can write down.

One step is not convergence. The theory's guarantee of one-step retrieval holds for well-separated patterns. When patterns are correlated, iterating the update moves the state towards a blend rather than towards the nearest pattern, so the fact that transformers apply the rule exactly once is doing useful work. Reading the single step as "an approximation to the fixed point" has it backwards.

Exponential capacity has a separation condition attached. The capacity result assumes patterns that are separated well enough for a single pattern to dominate the softmax. Realistic key matrices have duplicates and near-duplicates, and that is precisely where metastable averaging takes over.

The values are not the keys. In the memory, the thing retrieved is the stored pattern itself. In attention, \(V\) is a separate projection, so the output is a weighted sum of value vectors rather than of the keys that were matched. Interpretations that treat an attention output as "the retrieved key" skip a learned linear map.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Ramsauer et al., 2020, Hopfield Networks is All You Need, ICLR 2021, arXiv:2008.02217 arxiv.org
  2. Widrich et al., 2020, Modern Hopfield Networks and Attention for Immune Repertoire Classification, NeurIPS papers.nips.cc
  3. Barbero et al., 2025, Why do LLMs attend to the first token?, arXiv:2504.02732 arxiv.org
  4. Santos, Niculae, McNamee and Martins, 2024, Sparse and Structured Hopfield Networks, ICML, PMLR 235 proceedings.mlr.press
  5. Hoover et al., 2023, Energy Transformer, NeurIPS, arXiv:2302.07253 arxiv.org
Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track