The founding note of the project. Carried through — inference, backpropagation, the deterministic relaxation and the statistical-game reading — in Implicit Diffusion Learners. Worked out in full in The Diffusion Family and implemented in
lib/VariationalDiffusion.jl; the energy below is the one the code computes (reddiff.jl,factor.jl).
Sources: original to this vault (design and analysis; no single paper). Builds on Mardani, Song, Kautz & Vahdat, A Variational Perspective on Solving Inverse Problems with Diffusion Models (2023) — arXiv:2305.04391; Fang, Díaz, Buchanan & Sulam, Beyond Scores: Proximal Diffusion Models (2025) — arXiv:2507.08956; the Implicit Layers tutorial; code:
reddiff.jl,factor.jl
Implicit learning: relations instead of functions
Implicit means replacing learned functions by learned relations (Implicit Layers tutorial).
| Explicit machine learning | Implicit learning |
|---|---|
| approximator: functions | approximator: relations |
How do we learn a relation?
- Introduce an error (energy) space , assumed multivariate.
- Learn a function and read the relation off its approximate zero set:
A simple example: polynomials versus varieties
| Aspect | Explicit | Implicit |
|---|---|---|
| Approximator | multivariate polynomials | algebraic varieties |
| Inference | forward evaluation | root finding |
| Backpropagation | reverse-mode automatic differentiation | implicit function theorem / differential algebra |
| Universal approximation | continuous functions on compacta (Weierstraß) | compact smooth manifolds (Nash–Tognoli) |
| Well-posedness | always single-valued | may be multi-valued or empty; in general only the closest point to the variety |
| Loss | ||
| Symmetry | fixed direction from input to output | symmetric: no distinguished input or output |
| Cost of inference | cheap | expensive (Newton’s method, …) |
| Layer connections | directed acyclic graph | arbitrary connected graph |
Implicit layers, deep equilibrium networks (DEQs) and neural ODEs realise this to a limited
degree; see lib/ImplicitLayers.jl.
Implicit RED-Diff
Choosing input and output channels
Let be diagonal selection matrices with
so that and : the choice of input and output is a choice of masks, made at inference time, not at training time. Weight the three blocks with precisions:
In the code this is precision_vector in factor.jl; the limit
is implemented as a hard clamp (projection) rather than an infinite penalty.
The energy
The first term is the diffusion prior (the entropy side of the free energy, see RED-Diff as a Statistical Game); the second ties to the observed on the chosen channels. Its gradient in comes from the RED-Diff machinery: the score network supplies the prior gradient without differentiating through .
Alternatively, ProxDM learns a proximal operator instead of a score; see ProxDM and Proximal Alternatives.