annotation definition

§6 of Energy-Based Learning is titled Efficient Inference: Non-Probabilistic Factor Graphs. It is Mycelium.jl without the beliefs.

And it settles a complaint this vault has filed three times in three different packages.

Sources: original to this vault (design and analysis; no single paper).

Theory (CT-ML wiki): Rig (semirings) · Bayesian Inversion · Variational Free Energy

1. The structure, and it is the same structure

LeCun’s energy-based factor graph:

Energies add over factors; inference minimises the sum. On a tree the minimisation factorises and the algorithm is min-sum — Viterbi — with exactly the message structure of sum-product and a different semiring.

LeCun §6Mycelium
a factor a factor
a variable in an edge
GradedEnergy’s ; the Bethe sum
min-sum messagesfactor_message with DiracBeliefs
min-out a latentLatent() polarity
no partition functioncombine returns unnormalised beliefs
tree ⇒ exactistree(g) ⇒ tree_schedule is exact

The last row holds in both semirings and for the same reason: on a tree the exclusion principle makes each message summarise a disjoint subtree, and that argument never mentions which operations are being used.

2. The reframing: a Dirac message is a min-sum message

Every factor in every lib/ package returns a DiracBelief.

packagefactorreturns
VariationalDiffusionDiffusionFactorDiracBelief (RED-Diff’s is a point mass)
ImplicitLayersDEQFactorDiracBelief (a root-find returns a point)
ImplicitLayersNeuralODEFactorDiracBelief (an ODE solve returns a point)
AdversarialGeneratorFactorSampleBelief — particles, still no density

Three notes have recorded this as a defect, each blaming the same cause — GaussianBelief being stranded in the top-level package (The Equilibrium Family §5). That diagnosis is still true as far as it goes. But it is not the whole story, and the framing was wrong:

Those factors are not broken Bayesian factors. They are correct energy-based

factors. A DiracBelief message is a min-sum message. A root-find, an ODE solve and a RED-Diff prox are all minimisations, and a minimisation is what an energy-based factor graph does for inference. Three independent packages converged on point-valued messages not because each hit the same missing type, but because each wraps a model whose native inference is .

The project has an energy-based layer and a Bayesian layer, and most of the implementation lives in the first one.

3. Min-sum is the zero-temperature limit

The two semirings are not unrelated. Put a temperature on the Gibbs distribution, , and note

so sum-product at temperature becomes min-sum as . Marginalisation becomes minimisation; the Gibbs distribution concentrates on the argmin; a belief becomes a Dirac.

Two consequences worth writing down.

3.1 Latents may be minimised out, and at that is exact

Energy-Based Learning §5 gives both readings of a latent variable, and the min version is the limit of the marginal version.

constraint.md §4.1 records latent-channel marginalisation as “the single most valuable missing piece” in that file, and Copiers Cups and Caps calls marginalisation “the expensive one”. In a Dirac-valued graph the cheap alternative is not an approximation at all — it is the correct operation for the semiring the graph is actually running in.

3.2 The Bethe free energy degenerates correctly

Three packages have recorded a version of: these factors contribute energy but no entropy, so the counting correction has nothing to correct and the total is not .

Under the temperature reading that is not a bug report. The Bethe free energy is

and the entropy term carries a factor of . At it vanishes — regardless of being for a point mass, because — and what is left is

which is LeCun’s energy, exactly. So a graph of Dirac-valued factors is not failing to compute a free energy; it is computing the right object at the wrong temperature for the surrounding framework.

The real problem is the temperature mismatch, not the missing entropy

Lenticulum’s Bethe form has no explicit — it is implicitly . The Dirac factors are operating at inside it. Mixing them with a GaussianFactor sums a zero-temperature energy and a unit-temperature free energy and calls the result a free energy.

That is a sharper statement of the defect than “these factors have no entropy”, and it suggests a different fix: a temperature per factor, or a declared semiring per graph, rather than forcing every factor to produce a distribution it does not have.

There is a third semiring, and it is logic programming

— constraint satisfaction — sits beside and in the same framework. Dechter’s bucket elimination is one algorithm across all three, and Bistarelli–Montanari–Rossi’s semiring-based CSP makes the same point from the constraint side.

So Prolog and this project are the same algorithm at different semirings, which also explains why unification is idempotent and combine is not: is idempotent and is not. See Prolog and Logic Programming §4.

4. What this does and does not resolve

Resolved: the three complaints about Dirac-valued inversions were describing one thing — an energy-based sublayer — and describing it as a deficiency. It is a legitimate mode of operation with its own theory, its own exactness result on trees, and its own literature.

Not resolved: whether the project wants . The case for is everything AutoBayes is for — posteriors, uncertainty, the free energy as a model-comparison score, The Linear Gaussian Chain §4’s identity that holds to machine precision. None of that survives at zero temperature.

So GaussianBelief belonging in LenticulumCore is still the right call (The Equilibrium Family §5). What changes is the urgency and the reason: not “three packages are broken” but “three packages are energy-based and there is currently no way to be anything else from a lib/ package.”

Newly visible: the framework has no way to say which semiring a graph is running in, which The Type Discipline of a Factor Graph §3.3 identifies as a missing type index — if a belief carried its semiring, mixing would not typecheck, and the choice would be forced at construction rather than remembered. istree, isdag and isexact are all present; a MinSum / SumProduct distinction is not. A graph mixing the two is silently wrong, and nothing in validate(g) looks.

5. Loopy min-sum is a different subject

Off a tree the two semirings diverge, and the guarantees are not the same ones.

Loopy Message Passing and The Linear Gaussian Chain discuss loopy sum-product: for Gaussians, exact means and wrong variances (Weiss & Freeman), which ImplicitLayers’s resistive-divider test exhibits with a stable factor of five.

Loopy min-sum has its own analysis, and the flavour of the result is different: a fixed point of max-product/min-sum is locally optimal in a neighbourhood generated by trees and single loops, which is a genuine guarantee but a weaker and differently-shaped one than “the means are exact”. A graph that switches semirings switches which theory applies to it, and the vault currently documents only one of them.

6. What would follow

Not a plan — a note of what the reading suggests, in rough order of how much it buys:

  1. Declare the semiring. A trait on the graph or the factor saying whether messages are min-sum or sum-product, checked in validate. Cheap, and it turns a silent error into a loud one.
  2. A temperature. with explicit makes §3.2’s degeneration a limit rather than an inconsistency, and makes annealing expressible — which is what ImplicitREDDiff’s weights already are in disguise.
  3. min as an alternative to marginalisation for Latent(). Cheap, correct at , and it unblocks constraint.md §4.1 in the regime those factors actually run in.
  4. Loss functionals. The largest and the one Energy-Based Learning §3 argues is missing outright.

Related: Energy-Based Learning, Training Energy-Based Models, Factor Graphs, Prolog and Logic Programming, The Type Discipline of a Factor Graph, Bethe Free Energy, Messages are Inversions, Schedules, Loopy Message Passing, The Equilibrium Family, The Linear Gaussian Chain