The implicit diffusion learner with its randomness removed. The relaxations are, in order: fixing the noise (deterministic, still a smoothed MAP), using one noise level without noise (a fixed point of the Tweedie denoiser, a deep equilibrium model whose layer is the denoiser), and the zero-noise limit (modes of the data density, but an ill-conditioned field). The DEQ correspondence is exact and is checked to .
Sources: original to this vault (design and analysis); Romano, Elad & Milanfar, The Little Engine that Could: Regularization by Denoising (RED), SIAM J. Imaging Sci. 10(4) (2017); Bai, Kolter & Koltun, Deep Equilibrium Models, NeurIPS 2019; Efron, Tweedie’s Formula and Selection Bias, JASA 2011; Song et al. arXiv:2011.13456 (the probability-flow ODE); code:
implicit.jl(noisefree_nodes,field_nodes)Theory (CT-ML wiki): Least Fixed Point · Initial Algebra · Contextual Equivalence
1. Where the randomness is, and what each relaxation removes
| level | what is random | after relaxation | output |
|---|---|---|---|
| RED-Diff as published | redrawn every step | — | a stochastic approximation of a smoothed MAP |
| fixed nodes (sample-average approximation) | nothing | the expectation replaced by a fixed quadrature | a deterministic smoothed MAP |
| one level, no noise | nothing | one , | a fixed point of the denoiser at that level |
| zero-noise limit | nothing | the data density itself | modes of — sharp, but stiff |
| probability-flow ODE | the initial noise only | a deterministic map from noise to sample | a sample, determined by its seed |
The probability-flow ODE is deterministic given its initial noise, so it is a deterministic sampler rather than a deterministic relation. The rest are relations.
2. Fixed nodes: the deterministic smoothed MAP
Replacing by a fixed set of nodes gives a deterministic field whose roots are stationary points of , with (Implicit Diffusion Learners §3). Two refinements are worth having:
- Antithetic pairs cancel the control-variate tilt exactly, so has no spurious linear term.
- Few nodes break symmetries of the true expectation: on the circle the symmetry is lost and a merged branch sits slightly off the axis. More nodes shrink this.
This is the relaxation the package uses by default: it is cheap, it keeps the smoothing over several scales, and it makes inference a deterministic function, so the adjoint is exact.
3. One level, no noise: a DEQ whose layer is the denoiser
Take a single node, fixed and : . Tweedie’s denoiser at level is , so
With hard inputs and free outputs () the relation is the fixed-point set of the denoiser on the free coordinates:
That is a deep equilibrium model in the sense of DEQ as a Relation: the layer is , the input enters by clamping, and the backward pass is the same IFT, with (equivalently up to scale). Checked: for every root found, (test suite, “deterministic relaxation”).
This is Romano–Elad–Milanfar’s RED fixed point, , with the diffusion model’s Tweedie denoiser as . RED-Diff is its multi-level, stochastic generalisation.
So the three implicit families of Implicit Learners meet here. The diffusion family’s deterministic relaxation is an equilibrium-family model, with a layer that was trained as a denoiser rather than end to end. That has a consequence for training: a DEQ trained end to end need not be a denoiser of anything, whereas this one is, so it carries a density the DEQ lacks (it can be sampled, and its relation has a probabilistic reading).
With a soft anchor () the root solves , i.e.
an anchored fixed point. For an exact score it is the proximal step of the level- smoothed log-density in the metric : the step a plug-and-play ADMM iteration takes.
4. The level is a bias–conditioning trade-off
At a single level, the inferred radius of a unit circle (component width ) follows :
| inferred radius at | fixed-point error | ||
|---|---|---|---|
| 0.002 | 0.015 | 0.9981 | |
| 0.010 | 0.045 | 0.9977 | |
| 0.030 | 0.109 | 0.9927 | |
| 0.060 | 0.202 | 0.9766 |
Small means little smoothing bias, but varies on the scale , so the field is stiff, basins of attraction are small, and for a learned network the region where is accurate shrinks too: networks are worst at small . Multi-level fixed nodes (§2) are the compromise. Coarse levels give wide basins, fine levels give little bias, and annealing from coarse to fine (solve at a coarse node set, warm-start a finer one) gets both.
5. What the deterministic relaxation loses
- Multimodality. A point returns one branch, and which one depends on the start. A sampler would return all branches in proportion.
- Uncertainty. There are no error bars; the posterior is a Dirac (The Diffusion Factor §4).
- Calibration. The smoothing makes the relation the ridge of , not the support of ; the bias is computable (§4) but present.
And what it gains: a deterministic, differentiable, warm-startable map. That is exactly what an implicit learner needs, and what a factor graph’s message schedule can call repeatedly.
Related: Implicit Diffusion Learners, Inference Signatures, Backpropagation through Implicit Inference, DEQ as a Relation, The Equilibrium Family, ProxDM and Proximal Alternatives