The domain where the energy-based reading is not an analogy — a force field already is a sum of local energy terms over a graph, which is Energy-Based Factor Graphs §1 written in chemistry.
It is also the domain where the honest fit is narrower than it first looks, and §“What would be hard” says why.
Sources: original to this vault (design and analysis; no single paper).
A force field is an energy-based factor graph
A molecular potential decomposes into local terms:
Each term touches a few atoms; the total is their sum. That is over a factor graph whose variables are atomic positions — LeCun’s non-probabilistic factor graph, arrived at independently by computational chemistry decades earlier.
Some relations are exact constraints rather than soft energies: fixed bond lengths under SHAKE/RATTLE are holonomic constraints, which are acausal relations in the strict sense — a bond does not say which of its two atoms is the output.
The honest use case: structure from sparse data
MD as simulation — sampling a Boltzmann distribution by integrating dynamics — is not what message passing is for, and §“What would be hard” is blunt about that.
The fit is integrative structural biology: determining a structure from sparse, heterogeneous experimental restraints combined with a physical prior.
| source | contributes |
|---|---|
| NMR NOEs | noisy distance restraints between specific atom pairs |
| cryo-EM density | a spatial likelihood, not per-atom |
| SAXS / SANS | low-resolution shape information, heavily averaged |
| crosslinking–MS | sparse, ambiguous, sometimes wrong distance bounds |
| a force field | the physical prior over everything unconstrained |
Each is a factor. The experimental restraints are sparse and disagree; the force field fills in everywhere they are silent. What comes out should be a posterior over structures, and in practice what usually comes out is one refined model plus an ensemble generated by an ad-hoc procedure.
What would be learned
Machine-learned interatomic potentials — ANI, NequIP, MACE and their successors — replace some or all of the classical terms with a fitted model trained on quantum-chemical reference data. That is the grey-box case: learned energies beside exact holonomic constraints, in one graph, with the experimental restraints beside both.
Also learnable: coarse-grained potentials, and instrument-specific noise models for each experimental modality (which are usually guessed).
What a residual means
Strain, and — more usefully — incompatibility between the experiment and the force field. A restraint that the physics cannot satisfy is either a bad measurement or a wrong potential, and per-factor attribution is what distinguishes “this one NOE is misassigned” from “this whole region is modelled wrongly”.
That is a real workflow question in structural biology and it is currently answered by inspection.
What would be hard
Bluntly, more here than in the other five.
- Message passing is the wrong algorithm for sampling. MD’s task is to sample a high-dimensional, strongly-coupled Boltzmann distribution, and that is MCMC and integrator territory. Belief propagation on a dense nonbonded interaction graph has nothing to recommend it. The fit is refinement against sparse data, not simulation — and conflating the two would be overreach.
- Nonbonded terms make the graph dense. Bonded terms are sparse and local; Lennard-Jones and electrostatics are all-pairs, cut off but still dense. Sparsity is what makes factor-graph inference cheap, and this domain does not have it where it matters.
- Multimodality. Conformational states are distinct modes, and Gaussian beliefs cannot represent them. This is the messages §1 gap again, and in a domain where the multimodality is the science.
- Scale and timescale. – atoms, and the interesting transitions are rare events many orders of magnitude beyond the integration step.
The defensible claim is therefore narrow and worth stating as such: not molecular dynamics, but Bayesian structure refinement on a molecular energy graph — where the sparsity is in the experimental restraints, the prior is a force field, and the product is a posterior with per-restraint attribution.
Related: Motivating Examples, Geometric Deep Learning and Physical Laws, Energy-Based Factor Graphs, Energy-Based Learning, messages, Loopy Message Passing