Model Architecture

Attention Is Memory Retrieval: What the Hopfield View Predicts About Transformers

A 1982 network of 100 binary units could store about 14 memories. In 2020 the same energy idea, with one function changed, turned out to produce transformer attention exactly. The identity is narrow, it is provable, and it predicts things about real networks that the matrix-multiplication view does not.

A Hopfield network of 100 binary units stores about fourteen patterns. Push it to twenty and it does not get noisier; it returns confident blends of things that were never stored, and the originals are gone too. That cliff, quantified as a critical load of \(\alpha_c \approx 0.138\) patterns per unit (Amit, Gutfreund and Sompolinsky, 1985, PRL 55(14)), is why associative memory spent thirty years as a physics curiosity rather than a component.

Then in 2020 a group at JKU Linz changed one function inside the energy, derived the update rule that minimises it, and found the rule already running in production hardware: it is \(\operatorname{softmax}(QK^{\!\top}/\sqrt{d_k})V\) (Ramsauer et al., 2020, Hopfield Networks is All You Need, arXiv:2008.02217). Not an analogy, not a family resemblance. One attention operation is one retrieval step in a dense associative memory whose capacity is exponential in the dimension of the space it stores patterns in. Four years later the line of work that started it won the Nobel Prize in Physics (Royal Swedish Academy of Sciences, 2024).

The reason to care is not historical tidiness. The memory view makes predictions the linear-algebra view does not: that a high-entropy attention head is doing a different job from a low-entropy one rather than failing at the same job, that the temperature \(\beta\) is a capacity knob and not a normalisation detail, that iterating attention to convergence would be worse than applying it once, and that architectures which replace attention lose a specific capability with a measurable price tag.

Why this matters: The memory framing tells you which parts of attention are forced by theory (similarity, sharpening, projection) and which are free choices (the similarity function, the sharpness, the number of steps, whether the output is the matched key or a separate value). Every sub-quadratic architecture proposal is a bet on one of those free choices, and the bets that failed, failed on recall.

TL;DR

  • The capacity of an associative memory is set by its energy function, not by its size. A quadratic energy gives \(0.138N\) patterns; a degree-\(n\) polynomial gives \(\alpha_n N^{n-1}\); an exponential gives capacity exponential in \(N\). Same patterns, same network, different interaction function.
  • The log-sum-exp energy's update rule is attention. \(\xi^{\text{new}} = X\operatorname{softmax}(\beta X^{\!\top}\xi)\) with \(\beta = 1/\sqrt{d_k}\) is \(\operatorname{softmax}(QK^{\!\top}/\sqrt{d_k})V\), and it retrieves in one step with exponentially small error when patterns are well separated.
  • The energy has exactly three kinds of fixed point, storing one pattern, averaging a cluster, or averaging everything. Ramsauer et al. report that BERT's early layers sit in the global-averaging regime and its higher layers in metastable states, which makes attention entropy a read on which regime a head is in.
  • Attention is not energy descent. The derivation needs a symmetric similarity; real attention learns separate \(W_Q\) and \(W_K\). One step is a retrieval; a stack of them minimises nothing you can write down.
  • One step is the design, not a truncation. Iterating the update on correlated patterns drifts towards a blend. In the worked example below, one step lands at overlap 0.856 with the right pattern and four steps drift to 0.788 while a wrong pattern climbs to 0.800.
  • Recall is what sub-quadratic architectures gave up. Across 17 pretrained models, the best gated-convolution architectures trailed attention by up to 2.1 perplexity points on the Pile, and about 82% of that gap was attributable to in-context recall; on associative recall a 70M-parameter attention model beat a 1.4B-parameter gated convolution (Arora et al., 2023, Zoology, arXiv:2312.04927).
  • Explicit memory is coming back as a layer. Meta's memory layers are trainable key-value lookups that add parameters without adding FLOPs, scaled to 128B memory parameters over 1T tokens, beating dense models with more than twice the compute on factual tasks (Berges et al., 2024, arXiv:2412.09764).

At a Glance

flowchart LR
    Q["Query vector"] --> S["Similarity<br/>dot product with every pattern"]
    M["Stored patterns X"] --> S
    S --> SEP["Separation<br/>softmax at sharpness beta"]
    SEP --> P["Projection<br/>weighted sum of patterns"]
    P --> F{"Fixed point type"}
    F --> A["Single pattern<br/>retrieval"]
    F --> B["Metastable<br/>cluster average"]
    F --> C["Global average<br/>over-mixing"]

    classDef blue fill:#1e40af,stroke:#3b82f6,stroke-width:1px,color:#fff
    classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
    classDef teal fill:#0e7490,stroke:#22d3ee,stroke-width:1px,color:#fff
    classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
    classDef slate fill:#334155,stroke:#64748b,stroke-width:1px,color:#e2e8f0

    class Q,M blue
    class S,SEP,P purple
    class F slate
    class A teal
    class B amber
    class C amber

Every single-shot associative memory ever published factors into those three stages, and the models differ only in the similarity and separation functions they choose (Millidge et al., 2022, Universal Hopfield Networks, ICML, arXiv:2202.04557). Attention picks the dot product and softmax. That is the entire design decision.

Before Softmax: Forty Years of Attractors

Hopfield's 1982 construction is three lines. Store patterns \(\xi^\mu \in \{-1,+1\}^N\) as Hebbian outer products in a symmetric weight matrix with zero diagonal, define \(E(s) = -\frac{1}{2}\sum_{i\neq j}w_{ij}s_is_j\), and flip units one at a time in the direction of their local field. Because the weights are symmetric, every accepted flip lowers the energy, so the network cannot cycle: it falls into a local minimum and halts. Stored patterns sit at minima, so running the dynamics from a corrupted copy is recall, and the network "correctly yields an entire memory from any subpart of sufficient size" (Hopfield, 1982, PNAS 79(8)).

What made this a physics paper rather than an engineering one is that the model is an Ising system, which means the capacity question has an exact answer. Hopfield's own Monte Carlo runs at \(N = 30\) and \(N = 100\) put the usable load near \(0.15N\). The replica calculation pinned a first-order transition at \(\alpha_c \approx 0.138\), and the rigorous information-theoretic treatment showed that insisting on exact fixed points costs a logarithm: at most about \(N/(4\log N)\) memories, or \(N/(2\log N)\) if most of them suffice (McEliece, Posner, Rodemich and Venkatesh, 1987, IEEE Trans. Inf. Theory 33(4)). At \(N = 4096\) that is 565 patterns under the generous criterion and 123 under the strict one.

[IMAGE: Line chart with patterns stored on the x-axis and retrieval accuracy on the y-axis for N=100, showing a flat line near 1.0 until about 14 patterns and then a near-vertical drop to chance. A second, dashed line shows graceful degradation for comparison. Caption: "Classical Hopfield capacity is a first-order transition, not a slope. The network does not become unreliable; it stops."]

So the model was elegant, provably limited, and roughly a thousand times too small to be a component in anything. Hinton's Boltzmann machine made the units stochastic and added hidden ones, which bought a generative model and an intractable partition function; that pairing is what the 2024 Nobel citation covers, "for foundational discoveries and inventions that enable machine learning with artificial neural networks". Sparse distributed memory went through high-dimensional address spaces instead. Neither lifted the ceiling by orders of magnitude.

The unlock, when it came, was embarrassingly local. Krotov and Hopfield rewrote the energy as an explicit sum over patterns, \(E(s) = -\sum_\mu F(\xi^\mu\cdot s)\), and asked what happens if \(F\) is not quadratic. With \(F(x) = x^n\) the network stores "many more patterns than the number of neurons" (Krotov and Hopfield, 2016, NeurIPS, arXiv:1606.01164), specifically \(M = \alpha_n N^{n-1}\) with small retrieval errors tolerated, and \(M = N^{n-1}/(c_n\log N)\) for exact fixed points with \(c_n > 2(2n-3)!!\). Push the degree to infinity, so \(F(x) = e^x\), and capacity becomes exponential in the number of neurons while the basins of attraction stay almost as large as the classical model's (Demircigil et al., 2017, arXiv:1702.01929).

timeline
    title From attractor networks to attention
    1982 : Hopfield turns content-addressable memory into energy minimisation
    1985 : Amit, Gutfreund and Sompolinsky fix the critical load near 0.138 patterns per unit
    1987 : McEliece et al. prove exact recall costs a log N factor
    2016 : Krotov and Hopfield lift capacity to N to the power n minus 1 with polynomial energies
    2017 : Demircigil et al. take the degree to infinity for exponential capacity
    2020 : Ramsauer et al. show the log-sum-exp update rule is transformer attention
         : Widrich et al. deploy it as a Hopfield layer for immune repertoire classification
    2021 : Bricken and Pehlevan relate attention to Kanerva sparse distributed memory
         : Geva et al. find feed-forward blocks behaving as key-value memories
    2022 : Millidge et al. factor every associative memory into similarity, separation and projection
    2023 : Bietti et al. read transformer weight matrices as associative memories
         : Arora et al. measure what gated convolutions lose on recall
    2024 : Hopfield and Hinton share the Nobel Prize in Physics
         : Meta scales explicit key-value memory layers to 128B parameters
    2025 : Barbero et al. argue the first-token sink exists to prevent over-mixing

[IMAGE: Side-by-side energy landscapes over the same two-dimensional slice, one per interaction function: quadratic, x cubed, and exponential. Each shows three stored patterns as markers; the quadratic panel has one merged basin, the cubic has three shallow ones, the exponential has three deep narrow ones. Caption: "Same patterns, same network, three interaction functions. Capacity is a property of the curve, not of the size."]

How Retrieval Became Attention

The energy that produces a softmax

The exponential interaction function cannot be computed directly, because \(e^{\xi\cdot s}\) overflows at any realistic dimension. Written in its numerically stable form it becomes log-sum-exp, and that rewrite is where the whole result hides. Take continuous states \(\xi\in\mathbb{R}^d\), store \(M\) patterns as the columns of \(X\in\mathbb{R}^{d\times M}\), and define

\[E(\xi) = -\operatorname{lse}\!\left(\beta, X^{\!\top}\xi\right) + \tfrac{1}{2}\xi^{\!\top}\xi + \beta^{-1}\log M + \tfrac{1}{2}M_{\max}^2\]

with \(\operatorname{lse}(\beta,z) = \beta^{-1}\log\sum_i e^{\beta z_i}\) and \(M_{\max}\) the largest pattern norm. The first term rewards alignment with any stored pattern and is concave; the quadratic term keeps the state from running off to infinity and is convex. Minimising a sum of a concave and a convex term is exactly what the concave-convex procedure does, and applying it yields

\[\xi^{\text{new}} = X\,\operatorname{softmax}\!\left(\beta X^{\!\top}\xi\right)\]

Read the right-hand side from the inside out. \(X^{\!\top}\xi\) is the vector of similarities between the query state and every stored pattern. \(\operatorname{softmax}(\beta\,\cdot)\) sharpens those similarities into weights, with \(\beta\) controlling how sharply. Multiplying by \(X\) projects the weights back into the space of patterns, giving a convex combination of the stored vectors.

[IMAGE: Annotated equation diagram. The expression X softmax(beta X-transpose xi) is written large, with three labelled brackets underneath reading "similarity", "separation" and "projection", and a parallel line below showing softmax(QK-transpose over root d_k)V with the same three brackets aligned under the matching factors. Caption: "The substitution, drawn once. Queries for the state, keys for the similarity term, values for the projection."]

Now substitute. Let the query play the role of \(\xi\), the keys play \(X\) in the similarity term, the values play \(X\) in the projection term, and set \(\beta = 1/\sqrt{d_k}\). The result is \(\operatorname{softmax}(QK^{\!\top}/\sqrt{d_k})V\), which is the attention operation as published in 2017, with nothing added and nothing dropped. Ramsauer et al.'s theoretical claims attach to it directly: retrieval of a well-separated pattern completes in a single update, with exponentially small retrieval error, from a memory whose capacity is exponential in the dimension of the associative space.

Why one step, and not iterate to convergence

A reasonable reaction is that one step must be an approximation to the real fixed point, and that iterating would be better. The theory says the opposite, and it is worth being precise about why.

One-step convergence is guaranteed for well-separated patterns. When patterns are correlated, the softmax at a moderate \(\beta\) spreads weight across the whole cluster, the projection returns their average, and feeding that average back in moves the state further towards the cluster centre rather than towards the pattern you wanted. The energy has a minimum there, so iteration converges, and it converges to the wrong thing. The single step keeps the retrieval anchored to the original query.

This is not a corner case; it is the regime trained transformers spend most of their time in. Ramsauer et al. classify the stationary points into three types, a fixed point storing a single pattern, a metastable state averaging over a subset of similar patterns, and a global fixed point averaging over all of them, and report that transformer and BERT models operate in their first layers preferentially in the global-averaging regime and in higher layers in metastable states.

[IMAGE: Three small 2-D energy landscape panels sharing an axis. Panel 1 at low beta shows one wide basin with all patterns inside it. Panel 2 at medium beta shows two basins, each holding a cluster of patterns, with a marker at the cluster centre. Panel 3 at high beta shows one narrow basin per pattern. Caption: "Beta is not a normalisation constant. It selects whether a head pools, clusters, or retrieves."]

Sharpness is the capacity knob

Separating the three stages explains where capacity actually comes from. The crosstalk that kills the classical model is the sum of contributions from patterns the query does not match. A quadratic interaction barely distinguishes a match from a near-match, so all \(M-1\) wrong patterns contribute and their noise grows with \(M\) until it swamps the signal. A degree-\(n\) polynomial raises the gap between best and rest to the \(n\)th power. The exponential, and its softmax descendant, suppresses non-matching patterns geometrically. The wrong patterns never disappear; they stop mattering.

That framing also tells you which knobs exist. Millidge et al. show that classical Hopfield networks, sparse distributed memory and modern continuous Hopfield networks differ only in similarity and separation, and report that replacing the dot product with Euclidean or Manhattan similarity performs substantially better on many tasks. Attention's dot product is a choice, not a consequence.

Where the equivalence stops

The derivation needs the bilinear form in the energy to be symmetric in \(X\), and attention learns separate projections, so \(W_Q \neq W_K\) and the Lyapunov argument does not apply. The honest statement is narrow: one attention operation is one retrieval step in a dense associative memory, and a stack of them is not descending an energy. The output side differs too, because attention projects through \(V\) rather than returning the matched key, so an attention output is a combination of payloads and not of addresses.

[IMAGE: Heatmap of attention entropy by layer and head for a 12-layer encoder, normalised so 1.0 is uniform, with early layers mostly pale (near-uniform, global averaging) and later layers mostly dark (peaked, retrieval), and a mid-stack band of intermediate values annotated "metastable". Caption: "Entropy per head read as a regime map rather than a quality score."]

Seeing It in Motion

The same three-stage pipeline describes classical, dense and modern memories; only the separation function changes, and that one change moves the capacity by orders of magnitude.

flowchart TB
    subgraph CLASSICAL["Classical Hopfield 1982"]
        C1["Binary states"] --> C2["Quadratic energy"]
        C2 --> C3["Capacity 0.138 N"]
    end
    subgraph DENSE["Dense associative memory 2016"]
        D1["Binary states"] --> D2["Degree n polynomial"]
        D2 --> D3["Capacity alpha N to the n minus 1"]
    end
    subgraph MODERN["Modern Hopfield 2020"]
        M1["Continuous states"] --> M2["Log-sum-exp energy"]
        M2 --> M3["Capacity exponential in d"]
        M3 --> M4["Update rule is softmax attention"]
    end

    classDef slate fill:#334155,stroke:#64748b,stroke-width:1px,color:#e2e8f0
    classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
    classDef teal fill:#0e7490,stroke:#22d3ee,stroke-width:1px,color:#fff
    classDef emerald fill:#047857,stroke:#34d399,stroke-width:1px,color:#fff

    class C1,D1,M1 slate
    class C2,D2,M2 purple
    class C3,D3,M3 teal
    class M4 emerald

A single attention head, traced as a memory read, makes the division of labour between the projections visible.

sequenceDiagram
    participant T as Token stream
    participant K as Key store
    participant S as Softmax at 1 over root d
    participant V as Value store
    participant R as Residual stream
    T->>K: project tokens into addresses
    T->>V: project tokens into payloads
    T->>S: query from current position
    K->>S: similarity against every address
    Note over S: sharpening decides pool, cluster or retrieve
    S->>V: retrieval weights
    V->>R: convex combination of payloads
    Note over R: one step only; no iteration to a fixed point

Which of the three regimes a head lands in is a property of the query and the key distribution, not a property of the architecture, and it can change token by token.

stateDiagram-v2
    [*] --> Query
    Query --> Similarity: dot with all patterns
    Similarity --> Sharpen: scale by beta
    Sharpen --> Retrieved: one pattern dominates
    Sharpen --> Metastable: a cluster dominates
    Sharpen --> Pooled: nothing dominates
    Metastable --> Pooled: iterate on correlated patterns
    Retrieved --> [*]
    Pooled --> [*]
    Metastable --> [*]

Watch It Run

Animated diagram of a query flowing left to right through similarity, sharpening and projection stages, with a feedback loop from the projected state back to the similarity stage and a side branch into a metastable blend.
Solid animated edges carry the single retrieval pass, query to similarity to sharpening to projection. The animated self-loop on the sharpening stage is the iteration the theory permits and transformers decline to take; the amber feedback edge shows where that iteration lands on correlated patterns, a metastable blend rather than the nearest pattern. The static Mermaid figures above show the same structure if the animation is absent.

By the Numbers

Model Energy / separation Patterns storable Retrieval steps Reported basin behaviour
Hopfield 1982, approximate recall quadratic \(0.138N\) (565 at \(N{=}4096\)) iterate to a fixed point wide basins; first-order collapse above \(\alpha_c\)
Hopfield 1982, exact fixed points quadratic \(\approx N/(4\log N)\) (123 at \(N{=}4096\)) iterate to a fixed point same landscape, stricter criterion
Dense associative memory, degree \(n\) polynomial \(x^n\) \(\alpha_n N^{n-1}\), or \(N^{n-1}/(c_n\log N)\) exact iterate to a fixed point narrows as \(n\) grows
Demircigil et al. 2017 exponential \(e^x\) exponential in \(N\) iterate to a fixed point almost as large as classical
Modern Hopfield / attention log-sum-exp, softmax at \(\beta\) exponential in \(d\) one, for well-separated patterns three regimes set by \(\beta\) and separation
Sparse Hopfield, Fenchel-Young sparse transformations exact retrieval tied to sparsity one exact zeros on non-matches

[IMAGE: Log-scale scatter of model size on the x-axis against associative-recall accuracy on the y-axis, with attention models forming a curve that saturates by 70M parameters and gated-convolution models forming a lower curve still climbing at 1.4B. Caption: "A twenty-fold parameter disadvantage on the one task that needs content addressing."]

And the price of giving up attention-style retrieval, measured rather than argued:

Measurement Figure
Perplexity gap, best gated convolution vs attention on the Pile up to 2.1 points
Share of that gap attributable to in-context recall about 82%
Attention parameters needed to beat a 1.4B gated convolution at associative recall 70M
Gap closed by convolution-attention hybrids with input-dependent sparse attention 97.4%
Memory-layer parameters trained by Meta over 1T tokens 128B
Compute multiple of the dense baseline that memory layers beat more than 2x

Sources: capacity rows from Amit et al., 1985, McEliece et al., 1987, Krotov and Hopfield, 2016 as restated by Demircigil et al., 2017, Ramsauer et al., 2020 and Santos et al., 2024. Recall figures from Arora et al., 2023, measured across 17 pretrained attention and gated-convolution models. Memory-layer figures are the authors' own reported results in Berges et al., 2024 and have not been independently reproduced. The 565 and 123 entries are arithmetic from the published formulas, not separately measured.

A Concrete Example

Three patterns in \(d = 12\), entries in \(\{-1,+1\}\), deliberately correlated so the interesting regime shows up.

\[\xi^1 = (+1)^{12}, \quad \xi^2 = \xi^1 \text{ with the last two entries flipped}, \quad \xi^3 = (+1^6, -1^6)\]

Overlaps: \(\xi^1\!\cdot\xi^2 = 8\), \(\xi^1\!\cdot\xi^3 = 0\), \(\xi^2\!\cdot\xi^3 = 4\). So \(\xi^1\) and \(\xi^2\) are near-duplicates, which is what a real key matrix looks like.

Step 1. Corrupt the query. Take \(\xi^1\) and flip three entries (positions 2, 5 and 8). The query is

\[q = (+1,-1,+1,+1,-1,+1,+1,-1,+1,+1,+1,+1)\]

Step 2. Similarities. \(X^{\!\top}q = (6,\,2,\,-2)\). The right pattern leads, but not by much, because \(\xi^2\) shares ten of twelve entries with \(\xi^1\).

Step 3. Sharpen at the transformer's own temperature. With \(\beta = 1/\sqrt{12} = 0.2887\), the exponentials are \(e^{1.732} = 5.6522\), \(e^{0.577} = 1.7813\), \(e^{-0.577} = 0.5614\), summing to 7.9949. Weights:

\[\operatorname{softmax}(\beta X^{\!\top}q) = (0.7070,\ 0.2228,\ 0.0702)\]

Step 4. Project. The weighted sum of patterns has overlap \(0.856\) with \(\xi^1\) per dimension and \(0.718\) with \(\xi^2\). Take the sign of each entry and you recover \(\xi^1\) exactly. One step, three corrupted bits, correct answer.

Step 5. Now iterate, which the theory allows and the transformer declines. Feed the output back in as the new query and repeat:

Iteration Weights \((\xi^1,\xi^2,\xi^3)\) Overlap with \(\xi^1\) Overlap with \(\xi^2\)
1 0.707, 0.223, 0.070 0.856 0.718
2 0.587, 0.364, 0.050 0.829 0.771
3 0.521, 0.426, 0.053 0.805 0.791
4 0.482, 0.460, 0.058 0.788 0.800

By the fourth step the state is closer to \(\xi^2\) than to \(\xi^1\). The dynamics are converging on the metastable blend of the two correlated patterns, which is a genuine minimum of the energy and the wrong answer to the query. Running the retrieval once was not an approximation; it was the only thing that worked.

Step 6. Turn the sharpness up. At \(\beta = 1\) the weights at the first step are \((0.9817, 0.0180, 0.0003)\), the overlap with \(\xi^1\) is 0.9937, and iterating is stable: step 2 gives 0.9931 and step 3 gives 0.9929. Higher \(\beta\) bought a clean, iterable retrieval of this query. It also shrank every basin, so a query with six flipped bits instead of three falls on the wrong side of the divide between \(\xi^1\) and \(\xi^2\).

Step 7. Compare with what the classical network would do. Three patterns in twelve units is a load of \(\alpha = 0.25\), which is above \(\alpha_c \approx 0.138\). The quadratic-energy network at this load is in its spin-glass phase and retrieves nothing reliably. The only difference between the two machines is the separation function.

(Arithmetic in steps 2 to 6 computed directly from the definitions above; it reproduces in a few lines of NumPy.)

[IMAGE: Grouped bar chart of the three softmax weights at beta = 0.1, 0.2887, 0.5 and 1.0, with the weight on the correct pattern rising from 0.47 to 0.98 and the weight on the near-duplicate falling from 0.32 to 0.02. Caption: "The same query, four sharpness settings. At beta = 0.1 the retrieval returns a blend whose overlap with the wrong pattern is higher than with the right one."]

Where It Breaks

The energy story does not survive contact with a real transformer

Three load-bearing assumptions fail. The similarity must be symmetric, and \(W_Q \neq W_K\). The retrieved object must be the stored pattern, and attention returns values instead. And the patterns must be fixed during retrieval, whereas a transformer's keys are functions of the same input the query came from, recomputed every forward pass. The memory view is a correct account of one operation and an incorrect account of the stack.

Exponential capacity has a separation condition, and real keys violate it

Every capacity theorem in this literature assumes patterns drawn at random, so that crosstalk averages out. Keys derived from natural language are nothing like random: the same token in similar contexts produces near-identical keys, and that is precisely the configuration in which the softmax spreads weight over a cluster. The practical ceiling is therefore not the exponential bound but the separation of your actual keys, which no theorem gives you. The metastable regime is not a pathology in this light; it is what the mechanism does with duplicated memories.

[IMAGE: Two-panel schematic of token representations across depth. Left panel, no sink: twelve token vectors drawn as arrows that progressively align until nearly parallel by layer 12. Right panel, with a sink position: the same arrows retain their spread while surplus weight collects on a greyed-out first position. Caption: "Over-mixing, and the null memory that prevents it."]

Over-mixing at the bottom of the stack

If \(\beta\) is too small relative to the spread of the similarities, every query returns roughly the mean of the values, and depth compounds it: repeated near-uniform mixing drives token representations together. The long-context literature arrived at the same phenomenon from the other direction and gave it a name, the attention sink. One recent argument is that the first-token sink is a mechanism for avoiding over-mixing, with context length, depth and data packing all influencing how strongly it appears (Barbero et al., 2025, arXiv:2504.02732). Under the memory reading, a sink is a null memory, a pattern the head can put its weight on when it has nothing to retrieve, and an architecture that forbids one leaves the head no way to abstain.

Softmax cannot say "not found"

The weights are strictly positive, so there is no query for which an attention head returns nothing. The nearest thing to an abstention is a near-uniform distribution, which is indistinguishable from pooling. This is a real limitation of the separation function, and it is the one that sparse alternatives target: Fenchel-Young-derived Hopfield energies produce updates with exact zeros, and tie sparsity directly to exact retrieval (Santos, Niculae, McNamee and Martins, 2024, ICML, PMLR 235).

Capacity is not compression

"Exponential storage capacity" means an exponential number of patterns stay individually addressable before interference, not that they are stored cheaply. The patterns are still held explicitly, \(M\) vectors of \(d\) numbers, which in a transformer is the KV cache, growing linearly with sequence length and dominating inference memory. Quoting the capacity result next to a memory-footprint discussion mixes two unrelated quantities.

[IMAGE: Diagram of a query vector landing midway between two clusters of stored patterns, with the retrieved output drawn as a vector at the midpoint and labelled "confident blend, never stored". A second copy of the output is shown alongside a genuine retrieval with no visible difference in form. Caption: "Nothing in a softmax output marks it as an average of four things."]

The failure modes of the classical model are still there, rescaled

Reversed patterns are stable in the binary model because \(E(-s) = E(s)\). Mixture states are stable because odd combinations of three patterns survive the sign function. Both have continuous analogues: a query that lands between clusters retrieves their average, and the retrieved average is a confident output with no marker distinguishing it from a real memory. Nothing in a softmax flags "this answer is a blend of four things".

Alternative Designs

Design How it works Key advantage Key limitation Best when
Classical Hopfield quadratic energy, binary states, iterate to a fixed point wide basins, provable convergence capacity \(0.138N\), catastrophic collapse above it teaching, small attractor dynamics, theory
Dense associative memory polynomial interaction of degree \(n\) capacity \(N^{n-1}\) from one function change narrower basins, numerical fragility fixed pattern set, known separation
Modern Hopfield / attention log-sum-exp energy, one softmax step exponential capacity, one-step retrieval, trainable end to end asymmetric in practice, cannot abstain, quadratic in sequence length general sequence modelling
Sparse / structured Hopfield Fenchel-Young separation, exact zeros true abstention, sparsity tied to exact retrieval extra implementation cost, less standard tooling interpretable retrieval, multiple-instance learning
Sparse distributed memory high-dimensional addresses, hypersphere intersections biological plausibility, graceful degradation fixed address space, weaker trainability neuroscience-facing models
Explicit memory layers trainable key-value lookup beside the FFN parameters without FLOPs, strong factual gains large parameter footprint, routing complexity factual knowledge capacity at fixed compute
Gated convolution / SSM state compress history into a fixed-size state sub-quadratic in sequence length loses in-context recall, measurably long sequences where recall is not the task

The last row is the honest cost accounting the field took a while to produce. Arora et al. pretrained 17 attention and gated-convolution models, found the best convolutional architectures trailing attention by up to 2.1 perplexity points on the Pile, and attributed about 82% of the gap to in-context recall, with a 70M-parameter attention model beating a 1.4B-parameter gated convolution on associative recall. Hybrids with input-dependent sparse attention patterns closed 97.4% of the gap while keeping sub-quadratic scaling, which is the same conclusion the memory view suggests: you need a mechanism that can address content, and it does not have to be dense.

How It Is Used in Practice

As a drop-in layer. The clearest production use of the identity is a Hopfield layer substituted where pooling would otherwise go. Widrich et al. built immune repertoire classification on it, treating transformer attention explicitly as the update rule of a modern Hopfield network with exponential storage capacity, on a multiple-instance problem with orders of magnitude more instances than standard benchmarks and an extremely low witness rate (Widrich et al., 2020, NeurIPS). That is the shape of problem where the memory framing earns its keep: many candidate instances, few informative ones, retrieval rather than aggregation.

As an architectural principle. The Energy Transformer goes the other way, designing an energy first and deriving layers that lower it, which makes its attention deliberately unlike conventional attention (Hoover et al., 2023, NeurIPS, arXiv:2302.07253). It is a research architecture, not a production one, and its value so far is as a demonstration that the energy constraint can be kept rather than approximated.

As an interpretability frame. Reading weight matrices as associative memories turns training dynamics into a question about which associations are learned when. Bietti et al. do exactly this on a synthetic task mixing global and in-context bigram statistics, and find global bigrams learned quickly while the induction head that handles in-context bigrams develops more slowly (Bietti, Cabannes, Bouchacourt, Jégou and Bottou, 2023, NeurIPS, arXiv:2306.00802). The same frame supports scaling analysis: Cabannes, Dohmatob and Bietti model inner transformer layers as high-dimensional matrices of outer products of embeddings and derive scaling laws in sample and parameter size under heavy-tailed token distributions (ICLR 2024, arXiv:2310.02984). And the empirical finding that feed-forward layers act as key-value memories, keys matching training-data patterns and values inducing output distributions, is the non-attention half of the same picture (Geva, Schuster, Berant and Levy, 2021, EMNLP, arXiv:2012.14913).

As a reason to build memory back in explicitly. If the attention block is a memory whose contents are recomputed every forward pass, why should the model not also have one that persists? Meta's memory layers are that: a trainable key-value lookup adding parameters without adding FLOPs, scaled to 128B memory parameters over 1T tokens against dense baselines up to 8B, reported to beat dense models with more than twice the compute and compute-matched mixture-of-experts models, with the gains concentrated on factual tasks. Those are the authors' own measurements, so read the direction rather than the margin.

As a bridge to biology. Bricken and Pehlevan showed attention relates closely, under certain data conditions, to Kanerva's sparse distributed memory, whose read operation over intersections of high-dimensional hyperspheres approximates attention's softmax, and confirmed those conditions hold in pretrained GPT-2 models (NeurIPS 2021, arXiv:2111.05498). It is the most load-bearing of the biological-plausibility arguments because it is conditional and checkable.

Insights Worth Remembering

  1. Capacity is a property of the separation function, not of the network size. Classical and modern Hopfield networks can have identical patterns, identical dimensions and capacities differing by many orders of magnitude. When someone attributes memory capacity to parameter count, ask which function is doing the sharpening.

  2. Attention entropy is a regime indicator, not a quality metric. A near-uniform head is in the global-averaging regime, doing pooling; a peaked head is retrieving a pattern. Ramsauer et al.'s observation that early layers favour averaging and later layers metastable states means "sharpen the early heads" is usually fighting the architecture rather than fixing it.

  3. The single attention step is load-bearing. Depth in a transformer is not iteration of one retrieval, because each layer re-derives its own keys, and the worked example shows why that matters: four iterations of the same retrieval hand the query to the wrong pattern.

  4. \(\beta\) trades capacity against noise tolerance, and \(1/\sqrt{d_k}\) was chosen for neither. The scaling exists to keep logit variance stable as dimension grows. That it also lands in a usable retrieval regime is luck, and it is one of the few genuinely free knobs in an attention block.

  5. Softmax cannot abstain, so models build abstention out of whatever is available. The first-token sink is the clearest example: a position to dump weight on when there is nothing to retrieve. Architectures that remove it, including some efficient-attention variants, remove the escape hatch and pay in over-mixing.

  6. The recall gap is the measurable cost of giving up content addressing. 2.1 perplexity points, 82% of it attributable to in-context recall, and a 20x parameter disadvantage on associative recall. Any architecture claiming attention-free parity should report an associative-recall number, and the ones that close the gap do it by putting input-dependent addressing back in.

  7. "Exponential capacity" is about interference, not storage cost. The theorem bounds how many patterns stay individually retrievable. How much memory they occupy is the KV cache, a different quantity entirely.

Open Questions

Does the metastable regime explain anything about hallucination? The mechanism is suggestive: a query between clusters returns their average, confidently, with no marker separating a blend from a retrieval. That is a hypothesis, not a result. Nobody has shown that measured metastable attention states predict factual errors in a production model, and the obvious experiment, correlating head-level retrieval regime with error type, has not been published as far as I can find.

What is the effective separation of real key matrices? Every capacity bound in this literature assumes random patterns. The quantity that would actually predict a head's retrieval quality is the separation statistic of its learned keys on real text. It is cheap to measure and rarely reported.

Can the symmetry be restored cheaply? Tying \(W_Q\) and \(W_K\) would make the energy well defined and the dynamics provably convergent, at some cost in expressiveness. Whether that cost is large is an empirical question with a small experiment attached, and the absence of a clear published answer is odd given how often the energy framing is invoked.

Do sparse Hopfield updates pay off at scale? The theory ties sparsity to exact retrieval cleanly, but the published results are on multiple-instance learning and text rationalisation, not language-model pretraining. Whether exact zeros help or hurt a 10B-parameter model is unknown.

Where does explicit memory belong? Memory layers show that parameters without FLOPs buy factual accuracy. What is not established is how that interacts with retrieval augmentation, which solves a similar problem outside the weights, or whether the two are substitutes or complements at a fixed budget.

Sources and Further Reading

  1. Hopfield, J. J. (1982). "Neural networks and physical systems with emergent collective computational abilities." PNAS, 79(8), 2554-2558. doi:10.1073/pnas.79.8.2554
  2. Amit, D. J., Gutfreund, H., & Sompolinsky, H. (1985). "Storing infinite numbers of patterns in a spin-glass model of neural networks." Physical Review Letters, 55(14), 1530-1533. PDF mirror
  3. McEliece, R. J., Posner, E. C., Rodemich, E. R., & Venkatesh, S. S. (1987). "The capacity of the Hopfield associative memory." IEEE Transactions on Information Theory, 33(4), 461-482. Caltech record
  4. Krotov, D., & Hopfield, J. J. (2016). "Dense Associative Memory for Pattern Recognition." NeurIPS 2016. arXiv:1606.01164
  5. Demircigil, M., Heusel, J., Löwe, M., Upgang, S., & Vermet, F. (2017). "On a model of associative memory with huge storage capacity." arXiv:1702.01929
  6. Ramsauer, H., et al. (2020). "Hopfield Networks is All You Need." ICLR 2021. arXiv:2008.02217
  7. Widrich, M., et al. (2020). "Modern Hopfield Networks and Attention for Immune Repertoire Classification." NeurIPS 2020. Proceedings
  8. Bricken, T., & Pehlevan, C. (2021). "Attention Approximates Sparse Distributed Memory." NeurIPS 2021. arXiv:2111.05498
  9. Geva, M., Schuster, R., Berant, J., & Levy, O. (2021). "Transformer Feed-Forward Layers Are Key-Value Memories." EMNLP 2021. arXiv:2012.14913
  10. Millidge, B., Salvatori, T., Song, Y., Lukasiewicz, T., & Bogacz, R. (2022). "Universal Hopfield Networks: A General Framework for Single-Shot Associative Memory Models." ICML 2022, PMLR 162, 15561-15583. arXiv:2202.04557
  11. Hoover, B., et al. (2023). "Energy Transformer." NeurIPS 2023. arXiv:2302.07253
  12. Bietti, A., Cabannes, V., Bouchacourt, D., Jégou, H., & Bottou, L. (2023). "Birth of a Transformer: A Memory Viewpoint." NeurIPS 2023. arXiv:2306.00802
  13. Arora, S., et al. (2023). "Zoology: Measuring and Improving Recall in Efficient Language Models." ICLR 2024. arXiv:2312.04927
  14. Cabannes, V., Dohmatob, E., & Bietti, A. (2024). "Scaling Laws for Associative Memories." ICLR 2024. arXiv:2310.02984
  15. Santos, S., Niculae, V., McNamee, D., & Martins, A. F. T. (2024). "Sparse and Structured Hopfield Networks." ICML 2024, PMLR 235, 43368-43388. Proceedings
  16. Berges, V.-P., Oğuz, B., Haziza, D., Yih, W.-t., Zettlemoyer, L., & Ghosh, G. (2024). "Memory Layers at Scale." ICML 2025. arXiv:2412.09764
  17. Barbero, F., et al. (2025). "Why do LLMs attend to the first token?" arXiv:2504.02732
  18. Royal Swedish Academy of Sciences (2024). "The Nobel Prize in Physics 2024." nobelprize.org

Free to read, no ads, no sign-up. If it was useful you can buy me a coffee.