Entry point for the third of the three Implicit Learners families, worked out the way Algebraic Implicit Learners works out the first.
The claim being tested: a diffusion model is a statistical game whose inversion is a proximal operator. Not an analogy —
lib/VariationalDiffusion.jlimplements it, and the piecesLenticulumCoreneeded for it were already there.
Sources: Song, Sohl-Dickstein, Kingma, Kumar, Ermon, Poole, Score-Based Generative Modeling through Stochastic Differential Equations, ICLR 2021, arXiv:2011.13456; Mardani, Song, Kautz, Vahdat, A Variational Perspective on Solving Inverse Problems with Diffusion Models, arXiv:2305.04391; Efron, Tweedie’s Formula and Selection Bias, JASA 2011; Romano, Elad, Milanfar, The Little Engine that Could: Regularization by Denoising (RED), 2017; code:
lens.jlTheory (CT-ML wiki): Bayesian Inversion · Statistical Game · Bayesian Lens · Variational Free Energy · Lens
The one-paragraph version
Train to denoise samples corrupted by a known Gaussian process (The VP-SDE). Tweedie’s formula turns it into an MMSE denoiser, which makes it a prior rather than merely a sampler. Then, to condition on data, do not sample the reverse SDE — optimise: minimise data misfit plus a score-matching regulariser whose gradient is one forward pass of (RED-Diff as a Statistical Game). Wrap the result as a factor whose channels are blocks of one state space and whose polarity supplies the selection matrices (The Diffusion Factor).
The notes
- The VP-SDE — the forward process, its perturbation kernel, and why is the only property that matters downstream. Song et al. 2021.
- RED-Diff as a Statistical Game — the variational reformulation; the stop-gradient; the energy/entropy split; and λ is derivable, not merely tunable, with the calculation.
- The Diffusion Factor — becomes a
Polarity; what breaks when a Dirac-valued message meets a factor graph. - ProxDM and Proximal Alternatives — the other way to build the prox, and how the family relates to DPS, ΠGDM and plain plug-and-play.
The implicit learner, carried through (2026-10-02):
- Implicit Diffusion Learners — the relation as the zero set of a residual; the stop-gradient field is the exact gradient of a smoothed log-density; inference is a proximal point; the smoothing scale decides which relation you get.
- Inference Signatures — point, anytime, message, variational, sampling: what each returns and which algorithm computes it.
- Backpropagation through Implicit Inference — the Lagrangian worked out: state, adjoint, gradient; no stored noise; what to do when inference has not converged; learning a parabola from a circle.
- Deterministic Relaxation — fixed nodes, and the single noise-free level, which is a DEQ whose layer is the Tweedie denoiser.
- The Implicit Diffusion Factor as a Statistical Game — the element-by-element reading; the two training semantics; what is missing.
Implementation notes sit next to the code: VariationalDiffusion, schedule, predictor, reddiff, factor, implicit, analytic.
Where it sits among the three families
Implicit Learners’s table, with the diffusion column filled in from experience rather than expectation:
| algebraic | equilibrium | diffusion | |
|---|---|---|---|
| approximator | varieties | fixed points | score / denoiser network |
| inference | Gröbner, homotopy | fixed-point iteration | proximal descent |
| backward pass | implicit function theorem | IFT at the fixed point | stop-gradient (RED-Diff); IFT with fixed nodes (Backpropagation through Implicit Inference) |
| inversion type | SolverInversion | SolverInversion | ProximalInversion |
| posterior | a point (or a branch) | a point | a point — is a Dirac |
| exactness | exact where the Jacobian is invertible | exact at convergence | biased, by a computable amount |
The last row is the interesting one and is the subject of RED-Diff as a Statistical Game §4. The other two families are approximate because they stop early; this one is approximate at its fixed point, because the method deliberately discards the denoiser Jacobian. A solver that stopped early is an inexact inversion and Inversions and Bayesian Lenses says that is legal — the loss just gets worse. A method whose fixed point is the wrong point is a different situation, and the vault had not previously distinguished them.
What this family cost the framework
Nothing structural, which is the headline. ProximalInversion was already in lens.jl,
named after this package, before this package existed. Per-channel precisions in Polarity
were already exactly . GradedEnergySpace already expressed the split.
Two things it did expose:
energy’s signature presumes a causal factor. The core’senergy(factor, x, a, y, ps, st)wants AutoBayes’ split, which a diffusion factor does not have (it has one state space and a polarity). ModelingToolkit as an Acausal Relation’sLinearConstraintFactorhit the same wall from the acausal side. Two families, same complaint.- The Bethe machinery assumes every factor’s is an entropy. For this factor it is a
score-matching term standing in for one, because the actual posterior is a Dirac whose
differential entropy is . Mixing a
DiffusionFactorand aGaussianFactorin one graph gives a total free energy that is not for anything. See The Diffusion Factor §5.
Sources
- Song, Sohl-Dickstein, Kingma, Kumar, Ermon, Poole, Score-Based Generative Modeling through Stochastic Differential Equations, ICLR 2021, arXiv:2011.13456.
- Mardani, Song, Kautz, Vahdat, A Variational Perspective on Solving Inverse Problems with Diffusion Models, arXiv:2305.04391.
- Efron, Tweedie’s Formula and Selection Bias, JASA 2011 — the denoiser identity.
- Romano, Elad, Milanfar, The Little Engine that Could: Regularization by Denoising (RED), 2017 — the “RED” that RED-Diff is named after.
Related: Implicit Learners, ImplicitREDDiff, Factors are Parameterized Statistical Games, Algebraic Implicit Learners, Channels and Polarity