application

The domain where the energy-based reading is not an analogy — a force field already is a sum of local energy terms over a graph, which is Energy-Based Factor Graphs §1 written in chemistry.

It is also the domain where the honest fit is narrower than it first looks, and §“What would be hard” says why.

Sources: original to this vault (design and analysis; no single paper).

A force field is an energy-based factor graph

A molecular potential decomposes into local terms:

Each term touches a few atoms; the total is their sum. That is over a factor graph whose variables are atomic positions — LeCun’s non-probabilistic factor graph, arrived at independently by computational chemistry decades earlier.

Some relations are exact constraints rather than soft energies: fixed bond lengths under SHAKE/RATTLE are holonomic constraints, which are acausal relations in the strict sense — a bond does not say which of its two atoms is the output.

The honest use case: structure from sparse data

MD as simulation — sampling a Boltzmann distribution by integrating dynamics — is not what message passing is for, and §“What would be hard” is blunt about that.

The fit is integrative structural biology: determining a structure from sparse, heterogeneous experimental restraints combined with a physical prior.

sourcecontributes
NMR NOEsnoisy distance restraints between specific atom pairs
cryo-EM densitya spatial likelihood, not per-atom
SAXS / SANSlow-resolution shape information, heavily averaged
crosslinking–MSsparse, ambiguous, sometimes wrong distance bounds
a force fieldthe physical prior over everything unconstrained

Each is a factor. The experimental restraints are sparse and disagree; the force field fills in everywhere they are silent. What comes out should be a posterior over structures, and in practice what usually comes out is one refined model plus an ensemble generated by an ad-hoc procedure.

What would be learned

Machine-learned interatomic potentials — ANI, NequIP, MACE and their successors — replace some or all of the classical terms with a fitted model trained on quantum-chemical reference data. That is the grey-box case: learned energies beside exact holonomic constraints, in one graph, with the experimental restraints beside both.

Also learnable: coarse-grained potentials, and instrument-specific noise models for each experimental modality (which are usually guessed).

What a residual means

Strain, and — more usefully — incompatibility between the experiment and the force field. A restraint that the physics cannot satisfy is either a bad measurement or a wrong potential, and per-factor attribution is what distinguishes “this one NOE is misassigned” from “this whole region is modelled wrongly”.

That is a real workflow question in structural biology and it is currently answered by inspection.

What would be hard

Bluntly, more here than in the other five.

  • Message passing is the wrong algorithm for sampling. MD’s task is to sample a high-dimensional, strongly-coupled Boltzmann distribution, and that is MCMC and integrator territory. Belief propagation on a dense nonbonded interaction graph has nothing to recommend it. The fit is refinement against sparse data, not simulation — and conflating the two would be overreach.
  • Nonbonded terms make the graph dense. Bonded terms are sparse and local; Lennard-Jones and electrostatics are all-pairs, cut off but still dense. Sparsity is what makes factor-graph inference cheap, and this domain does not have it where it matters.
  • Multimodality. Conformational states are distinct modes, and Gaussian beliefs cannot represent them. This is the messages §1 gap again, and in a domain where the multimodality is the science.
  • Scale and timescale. – atoms, and the interesting transitions are rare events many orders of magnitude beyond the integration step.

The defensible claim is therefore narrow and worth stating as such: not molecular dynamics, but Bayesian structure refinement on a molecular energy graph — where the sparsity is in the experimental restraints, the prior is a force field, and the product is a posterior with per-restraint attribution.

Related: Motivating Examples, Geometric Deep Learning and Physical Laws, Energy-Based Factor Graphs, Energy-Based Learning, messages, Loopy Message Passing