Energy-Based Learning §3 says Lenticulum has energies and no loss functionals. This note is the modern answer to what a loss functional should be — and the finding is that two of the three standard answers are already implemented in this repository, under other names.
Sources: Song & Kingma, How to Train Your Energy-Based Models (arXiv:2101.03288), and Du & Mordatch, Implicit Generation and Modeling with Energy-Based Models (arXiv:1903.08689)
Theory (CT-ML wiki): Statistical Game · Variational Free Energy
1. The problem
An EBM defines with unknown. Maximum likelihood needs
— push down on the data, push up on the model’s own samples. That second term is Energy-Based Learning §3’s contrastive term, made concrete: it is an expectation under the model, and sampling from the model is the whole difficulty.
Song & Kingma organise the escape routes into three.
2. The three families, and where each already lives here
| family | avoids by | already in this repo as |
|---|---|---|
| MLE with MCMC | sampling the model (Langevin, contrastive divergence) | — nothing |
| Score matching | matching , which kills | VariationalDiffusion.jl |
| Noise-contrastive estimation | classifying data against known noise | Adversarial.RatioFactor |
The middle two rows are the finding.
2.1 Score matching is the diffusion package
, and the gradient is taken in , so — a constant in — differentiates away. That is why score matching needs no normaliser and no sampling.
Denoising score matching is the diffusion objective. VariationalDiffusion’s
predictor.md §2 records the identity
which is exactly Song & Kingma’s score-matching estimator with the noise scale as the
smoothing parameter. So the package built for The Diffusion Family is, without ever
saying so, an energy-based model trained by score matching — and ImplicitREDDiff’s
regulariser is that trained energy being reused at inference time.
The connection is not decorative. It explains a fact reddiff.md §4 found empirically and
could not account for: the calibration constant , derived there so that
RED-Diff’s implied prior matches a Gaussian’s, is doing the job the missing normaliser would
have done. A score model knows and not ; the constant of
integration is exactly what is standing in for, and that is why one scalar cannot
fit a correlated prior (RED-Diff as a Statistical Game).
2.2 NCE is the ratio factor — and it recovers the normaliser
Noise-contrastive estimation trains a classifier to separate data from a known noise
distribution . Adversarial’s ratio.md §1 has the identity already:
In a GAN, is the generator and is unknown, so you get a ratio and stop. In NCE, is chosen and known, so
and you have the normalised log-density, partition function included. That is the whole
trick, and it changes what RatioFactor is worth:
RatioFactorwith a known noise distribution isbelief_logdensitymessages §1 has recorded since the beginning that the package’s main gap is
combinefor particle beliefs, which needsbelief_logdensity, which nothing implements.Implicit Generative Models §5 got as far as: a ratio is enough for importance reweighting. NCE goes one step further — with known you get the density, which is what
messages.mdactually asked for, and it unlocks the generic path rather than a special-cased reweighting.The code change is small:
RatioFactoralready computes the logit; it needs a field holding and a method adding the two. What it does not have is any way to train the classifier, which is the real work.
3. Du & Mordatch: composition is the factor graph
Du & Mordatch scale EBM training with Langevin dynamics plus a replay buffer, and report three capabilities. Two of them are things this project gets structurally rather than as a technique.
Compositionality. Independently trained EBMs compose by adding energies — — giving concept conjunction with no retraining. That is presented as a notable property of EBMs.
It is what a factor graph is. is Energy-Based Factor Graphs §1, it is
GradedEnergy’s , and it is Mycelium.combine in the log domain. Lenticulum does not
have to discover compositionality; the whole architecture is that operation.
The difference is what happens next: Du & Mordatch compose energies and then sample by Langevin; Lenticulum composes energies and then passes messages. Same composition, two inference algorithms — and for Dirac-valued factors Lenticulum’s is min-sum, which is the limit of the Gibbs distribution Langevin is sampling.
Inpainting and corrupt-image reconstruction. Clamp some coordinates, minimise over the
rest. That is ImplicitREDDiff’s selection matrices
and DiffusionFactor’s precision_vector — The Diffusion Factor. Again: a capability
there, a polarity here.
Langevin as the missing inference mode. The one thing Du & Mordatch have that this project
does not is sampling from a composed energy. RED-Diff’s prox is gradient descent on an
energy — Langevin without the noise, i.e. MAP rather than a sample. Adding the noise term is
a two-line change to reddiff_solve and would turn a point estimate into a sample, which is
the difference between and in Energy-Based Factor Graphs §3.
4. What is actually missing
Putting Energy-Based Learning §3 together with this note, the gap is specific rather than diffuse:
- No loss functional. The free energy is evaluated at the inferred configuration; nothing raises the energy elsewhere. Every training method above exists to supply that second term, and the framework has no slot to put one in.
- No sampling from a composed energy. MCMC is the one family of §2 with no counterpart here, and it is the one Du & Mordatch show scales.
- No training at all, in fact.
VariationalDiffusionconsumes a trained ;Adversarialconsumes a trained discriminator. Both packages deliberately have no AD dependency. The energy-based reading says what the training objective would be; nothing computes it. - The normaliser is nowhere.
logpartitionexists forGaussianBeliefand for nothing else, which is exactly the situation an EBM is designed to tolerate — but it meansscalar_free_energyreturning (The Linear Gaussian Chain §4) is a Gaussian-only guarantee, and no note has said so plainly.
Sources
- Song & Kingma, How to Train Your Energy-Based Models, arXiv:2101.03288 — MLE with MCMC, score matching, NCE.
- Du & Mordatch, Implicit Generation and Modeling with Energy-Based Models, arXiv:1903.08689 — Langevin training at scale, the replay buffer, compositionality, inpainting.
- Gutmann & Hyvärinen, Noise-Contrastive Estimation, AISTATS 2010 — §2.2’s identity, and the reason a known recovers the normaliser.
- Hyvärinen, Estimation of Non-Normalized Statistical Models by Score Matching, JMLR 2005 — the original of §2.1.
- Vincent, A Connection Between Score Matching and Denoising Autoencoders, 2011 — why the diffusion objective is score matching.
Related: Energy-Based Learning, Energy-Based Factor Graphs, Implicit Generative Models, The Diffusion Family, RED-Diff as a Statistical Game, The Diffusion Factor, messages, ratio, reddiff