derivation design

The implicit diffusion learner with its randomness removed. The relaxations are, in order: fixing the noise (deterministic, still a smoothed MAP), using one noise level without noise (a fixed point of the Tweedie denoiser, a deep equilibrium model whose layer is the denoiser), and the zero-noise limit (modes of the data density, but an ill-conditioned field). The DEQ correspondence is exact and is checked to .

Sources: original to this vault (design and analysis); Romano, Elad & Milanfar, The Little Engine that Could: Regularization by Denoising (RED), SIAM J. Imaging Sci. 10(4) (2017); Bai, Kolter & Koltun, Deep Equilibrium Models, NeurIPS 2019; Efron, Tweedie’s Formula and Selection Bias, JASA 2011; Song et al. arXiv:2011.13456 (the probability-flow ODE); code: implicit.jl (noisefree_nodes, field_nodes)

Theory (CT-ML wiki): Least Fixed Point · Initial Algebra · Contextual Equivalence

1. Where the randomness is, and what each relaxation removes

levelwhat is randomafter relaxationoutput
RED-Diff as published redrawn every step—a stochastic approximation of a smoothed MAP
fixed nodes (sample-average approximation)nothingthe expectation replaced by a fixed quadraturea deterministic smoothed MAP
one level, no noisenothingone , a fixed point of the denoiser at that level
zero-noise limit nothingthe data density itselfmodes of — sharp, but stiff
probability-flow ODEthe initial noise onlya deterministic map from noise to samplea sample, determined by its seed

The probability-flow ODE is deterministic given its initial noise, so it is a deterministic sampler rather than a deterministic relation. The rest are relations.

2. Fixed nodes: the deterministic smoothed MAP

Replacing by a fixed set of nodes gives a deterministic field whose roots are stationary points of , with (Implicit Diffusion Learners §3). Two refinements are worth having:

  • Antithetic pairs cancel the control-variate tilt exactly, so has no spurious linear term.
  • Few nodes break symmetries of the true expectation: on the circle the symmetry is lost and a merged branch sits slightly off the axis. More nodes shrink this.

This is the relaxation the package uses by default: it is cheap, it keeps the smoothing over several scales, and it makes inference a deterministic function, so the adjoint is exact.

3. One level, no noise: a DEQ whose layer is the denoiser

Take a single node, fixed and : . Tweedie’s denoiser at level is , so

With hard inputs and free outputs () the relation is the fixed-point set of the denoiser on the free coordinates:

That is a deep equilibrium model in the sense of DEQ as a Relation: the layer is , the input enters by clamping, and the backward pass is the same IFT, with (equivalently up to scale). Checked: for every root found, (test suite, “deterministic relaxation”).

This is Romano–Elad–Milanfar’s RED fixed point, , with the diffusion model’s Tweedie denoiser as . RED-Diff is its multi-level, stochastic generalisation.

So the three implicit families of Implicit Learners meet here. The diffusion family’s deterministic relaxation is an equilibrium-family model, with a layer that was trained as a denoiser rather than end to end. That has a consequence for training: a DEQ trained end to end need not be a denoiser of anything, whereas this one is, so it carries a density the DEQ lacks (it can be sampled, and its relation has a probabilistic reading).

With a soft anchor () the root solves , i.e.

an anchored fixed point. For an exact score it is the proximal step of the level- smoothed log-density in the metric : the step a plug-and-play ADMM iteration takes.

4. The level is a bias–conditioning trade-off

At a single level, the inferred radius of a unit circle (component width ) follows :

inferred radius at fixed-point error
0.0020.0150.9981
0.0100.0450.9977
0.0300.1090.9927
0.0600.2020.9766

Small means little smoothing bias, but varies on the scale , so the field is stiff, basins of attraction are small, and for a learned network the region where is accurate shrinks too: networks are worst at small . Multi-level fixed nodes (§2) are the compromise. Coarse levels give wide basins, fine levels give little bias, and annealing from coarse to fine (solve at a coarse node set, warm-start a finer one) gets both.

5. What the deterministic relaxation loses

  • Multimodality. A point returns one branch, and which one depends on the start. A sampler would return all branches in proportion.
  • Uncertainty. There are no error bars; the posterior is a Dirac (The Diffusion Factor §4).
  • Calibration. The smoothing makes the relation the ridge of , not the support of ; the bias is computable (§4) but present.

And what it gains: a deterministic, differentiable, warm-startable map. That is exactly what an implicit learner needs, and what a factor graph’s message schedule can call repeatedly.

Related: Implicit Diffusion Learners, Inference Signatures, Backpropagation through Implicit Inference, DEQ as a Relation, The Equilibrium Family, ProxDM and Proximal Alternatives