model design

The founding note of the project. Carried through — inference, backpropagation, the deterministic relaxation and the statistical-game reading — in Implicit Diffusion Learners. Worked out in full in The Diffusion Family and implemented in lib/VariationalDiffusion.jl; the energy below is the one the code computes (reddiff.jl, factor.jl).

Sources: original to this vault (design and analysis; no single paper). Builds on Mardani, Song, Kautz & Vahdat, A Variational Perspective on Solving Inverse Problems with Diffusion Models (2023) — arXiv:2305.04391; Fang, Díaz, Buchanan & Sulam, Beyond Scores: Proximal Diffusion Models (2025) — arXiv:2507.08956; the Implicit Layers tutorial; code: reddiff.jl, factor.jl

Implicit learning: relations instead of functions

Implicit means replacing learned functions by learned relations (Implicit Layers tutorial).

Explicit machine learningImplicit learning
approximator: functions approximator: relations

How do we learn a relation?

  • Introduce an error (energy) space , assumed multivariate.
  • Learn a function and read the relation off its approximate zero set:

A simple example: polynomials versus varieties

AspectExplicitImplicit
Approximatormultivariate polynomialsalgebraic varieties
Inferenceforward evaluationroot finding
Backpropagationreverse-mode automatic differentiationimplicit function theorem / differential algebra
Universal approximationcontinuous functions on compacta (Weierstraß)compact smooth manifolds (Nash–Tognoli)
Well-posednessalways single-valuedmay be multi-valued or empty; in general only the closest point to the variety
Loss
Symmetryfixed direction from input to outputsymmetric: no distinguished input or output
Cost of inferencecheapexpensive (Newton’s method, …)
Layer connectionsdirected acyclic grapharbitrary connected graph

Implicit layers, deep equilibrium networks (DEQs) and neural ODEs realise this to a limited degree; see lib/ImplicitLayers.jl.

Implicit RED-Diff

Choosing input and output channels

Let be diagonal selection matrices with

so that and : the choice of input and output is a choice of masks, made at inference time, not at training time. Weight the three blocks with precisions:

In the code this is precision_vector in factor.jl; the limit is implemented as a hard clamp (projection) rather than an infinite penalty.

The energy

The first term is the diffusion prior (the entropy side of the free energy, see RED-Diff as a Statistical Game); the second ties to the observed on the chosen channels. Its gradient in comes from the RED-Diff machinery: the score network supplies the prior gradient without differentiating through .

Alternatively, ProxDM learns a proximal operator instead of a score; see ProxDM and Proximal Alternatives.