Boltzmann Machines and Hidden-Unit Energies
How adding temperature and hidden units turns an attractor network into a generative distribution, why the restricted architecture makes sampling cheap, and the exact mapping that shows a Hopfield network is an RBM with Gaussian hidden units.
A Hopfield network answers one question: which stored pattern is this query closest to? It cannot tell you how probable a state is, cannot generate novel states, and cannot represent any structure that is not expressible as pairwise interactions between visible units. Two changes fix all three at once. Make the dynamics stochastic rather than deterministic, so the same energy defines a distribution instead of a descent path. Then add units that are never observed, so the model can express correlations among visible units that no pairwise term could capture. The result is the Boltzmann machine, and the pair of ideas it rests on is what the 2024 Nobel Prize in Physics recognised alongside Hopfield's own work (Royal Swedish Academy of Sciences, 2024).
From descent to distribution
Keep the energy and replace \(\operatorname{sign}\) with a coin flip whose bias depends on the local field:
At \(T \to 0\) this is the Hopfield update. At \(T > 0\) the chain can climb out of a local minimum, so it explores the landscape rather than freezing in the first basin it meets, and the stationary distribution is the Boltzmann distribution over states. The cost arrives immediately in the form of \(Z\), the sum of \(e^{-E(s)}\) over all \(2^N\) states, which is the same obstacle that defines the whole energy-based family (see Energy-Based Models and the Partition Function).
Why "restricted" made it usable
A general Boltzmann machine allows any connection, which makes even conditional sampling slow. The restricted version bipartites the graph into visible units \(v\) and hidden units \(h\) with no within-layer connections:
Now \(p(h \mid v)\) factorises, every hidden unit is conditionally independent given the visible layer, and \(p(h_j = 1 \mid v) = \sigma(b_j + \sum_i W_{ij}v_i)\). One matrix multiply samples the entire hidden layer, and block Gibbs sampling alternates between the two layers. The maximum-likelihood gradient has the shape every energy-based model has, a positive phase from the data and a negative phase from the model:
The second expectation needs samples from the model, which is what contrastive divergence approximates with a short chain started at the data; one Gibbs step is usually enough in practice (Hinton, 2002, Training products of experts by minimizing contrastive divergence, Neural Computation 14(8)). That route and its failure modes are covered in Contrastive Divergence and the Negative Phase.
Hidden units are stored patterns
The connection back to attractor memory is exact, not analogical. Any Hopfield network with \(N\) binary variables and \(p < N\) patterns, including correlated ones, can be transformed into a restricted Boltzmann machine with \(N\) binary visible variables and \(p\) Gaussian hidden variables (Smart and Zilman, 2021, On the mapping between Hopfield networks and Restricted Boltzmann Machines, ICLR, arXiv:2101.11744). Earlier work had established the correspondence only for orthogonal, uncorrelated patterns (Barra, Bernacchia, Santucci and Contucci, 2012, On the equivalence of Hopfield Networks and Boltzmann Machines, arXiv:1105.2790), and Smart and Zilman's extension to correlated patterns is what makes it practically interesting: their MNIST experiments use the mapping to initialise RBM weights from Hebbian patterns.
Read in that direction, each hidden unit is a stored pattern, and the hidden layer's activation is the vector of similarities between the input and the memory, which is the first stage of the similarity, separation, projection decomposition that also describes attention.
When it breaks
No likelihood without an estimator. \(Z\) is unknown, so a trained RBM cannot report a log-likelihood. Published figures come from annealed importance sampling and should be read as estimates with wide and asymmetric error bars (see Annealed Importance Sampling and Estimating log Z).
Contrastive divergence optimises something else. The short chain does not reach the model distribution, so the negative-phase samples are biased, and the bias is neither small nor well characterised. Models trained this way often generate acceptable samples while assigning poor likelihoods.
Depth makes inference approximate too. Stacking hidden layers breaks conditional independence between them, so even the positive phase needs mean-field or sampling approximations. Deep Boltzmann machines pay twice.
Pairwise energies are a real limitation. The hidden layer buys expressiveness, but the energy is still bilinear in \((v, h)\). Any structure requiring higher-order interactions must be represented through many hidden units rather than directly, which is the limitation that dense associative memories address by changing the interaction function instead.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Royal Swedish Academy of Sciences, 2024 nobelprize.org
- Hinton, 2002, Training products of experts by minimizing contrastive divergence, Neural Computation 14(8) doi.org
- Smart and Zilman, 2021, On the mapping between Hopfield networks and Restricted Boltzmann Machines, ICLR, arXiv:2101.11744 arxiv.org
- Barra, Bernacchia, Santucci and Contucci, 2012, On the equivalence of Hopfield Networks and Boltzmann Machines, arXiv:1105.2790 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.