Energy-Based & Score Models
Unnormalised densities, score matching, Langevin dynamics, and the SDE view that unifies the generative families.
5concepts
58flashcards
37minutes of reading
- 01 Contrastive Divergence and the Negative Phase How energy-based training approximates an intractable expectation with a short MCMC chain, what bias that introduces, and why the same positive-negative structure appears in contrastive learning and reward modelling.
- 02 Energy-Based Models and the Partition Function The most flexible way to specify a probability distribution, why the normalising constant makes it untrainable by direct likelihood, and the three escape routes that define the rest of the field.
- 03 Langevin Dynamics for Sampling How a noisy gradient ascent on log density becomes a valid sampler, why the noise term is what separates sampling from optimisation, and the mixing failure that makes it impractical alone.
- 04 Score Matching and Its Denoising Form How targeting the gradient of log density eliminates the partition function, why the naive form requires an intractable Hessian trace, and how adding noise makes the objective a simple regression.
- 05 The SDE View of Generative Models The continuous-time framework in which diffusion, score matching and denoising are the same object, why every SDE has a deterministic twin with identical marginals, and what the unification actually buys.