model

Implemented in lib/VariationalDiffusion.jl/src/reddiff.jl; see reddiff.

Sources: Mardani et al. 2023, arXiv:2305.04391, read as AutoBayes Definition 20 — and the sharpest thing this vault has to say about the method: **its regularisation weight λ is derivable, not merely tunable, and one scalar λ cannot fit a correlated prior

Theory (CT-ML wiki): Bayesian Inversion · Statistical Game · Bayesian Lens · Variational Free Energy · Lens

1. Don’t sample the reverse SDE — optimise

The standard way to condition a diffusion model on data is to run the reverse SDE with a guidance term (DPS, ΠGDM). RED-Diff refuses: it posits a variational posterior , minimises , and turns inference into an optimisation over .

Notation

RED-Diff writes for the clean signal and for the measurement; ImplicitREDDiff writes for the state and for the clamp. This vault writes for the joint state and for the evidence (clamp values and anchors), as fixed in Channels and Polarity §“Notation”.

Specialised to ImplicitREDDiff’s linear clamp, the objective is

which is exactly the energy in the prompt and in ImplicitREDDiff. So the note’s sketch was right, and this is the formal reading of it.

2. The energy/entropy split, and it is not arbitrary

Factors are Parameterized Statistical Games (Definition 20) wants a factor to supply an energy and an entropy . The two terms above sort themselves:

termiswhy
the energy pointwise; a function of the data point
the entropy a function of the learned distribution, not of the datum

This was the vault’s first reading, and the code still splits the terms this way. The Implicit Diffusion Factor as a Statistical Game §2 revises it: the score-matching term is the energy of the prior game (a cross-entropy, by Remark 24’s “priors are games too”), and the entropy slot belongs to the inversion’s output. The split itself is unchanged; only its labels differ. LenticulumCore.GradedEnergySpace expresses it as (clamp = ..., score = ...) with no new type.

But the "entropy" is not an entropy

is a point mass, so . The score term stands in for the entropy because it plays the entropy’s structural role (it depends on the learned distribution and it regularises the inversion), not because it equals one. So a DiffusionFactor and a GaussianFactor in the same graph produce a Bethe total that is not for anything. See The Diffusion Factor §5.

Superseded reading (2026-10-02)

The Implicit Diffusion Factor as a Statistical Game §2 files the score-matching term as the energy of a prior game (Remark 24), not as the entropy: it is a cross-entropy , and the entropy slot holds . The table above is kept for the record.

3. Proposition 2: the stop-gradient is the method

Two separate things happen, and conflating them is easy:

  • the reparametrisation is kept — is differentiable in , and its is absorbed into ;
  • the denoiser Jacobian is dropped — the true gradient carries in front of the residual, and RED-Diff replaces it with the identity.

So RED-Diff is not an unbiased estimator of the regulariser’s gradient; it is a preconditioned one. In the vocabulary of abstract_types’s AbstractGradientCoupling, the sampling path is PathwiseCoupling and the network Jacobian is stopped — which is not any of the four constructors, and is worth recording as a fifth case rather than forced into DiagonalCoupling.

This is why the implementation needs no AD

One forward pass of per Monte-Carlo draw; the clamp’s gradient is in closed form. lib/VariationalDiffusion.jl therefore depends on LuxCore, Random and LinearAlgebra — and not on Lux, Zygote or Enzyme. The stop-gradient is usually sold as a memory saving; here it is the difference between a package with an AD dependency and one without. Training through inference does need network derivatives; they come from a backend the user chooses, through a package extension (backends).

4. λ is derivable, and this is the finding

Both halves below are asserted in lib/VariationalDiffusion.jl/test/runtests.jl.

4.1 For isotropic Gaussian data there is exactly one correct λ

Gaussian data is the one case where has a closed form (The VP-SDE §3), so the expectation can be done by hand. With and :

RED-Diff’s regulariser is a Gaussian prior of precision . The true prior gradient is , so the regulariser is correct iff , which pins λ uniquely. With that λ, the fixed point gives

which is exactly the Gaussian posterior mean. For the default schedule and , — while the paper’s tuned value is .

Tuning λ is not adjusting “how much regularisation feels right”. It is choosing how strong the learned prior is. A wrong λ biases every posterior by a computable factor.

4.2 For correlated data no single λ works

Take . Both and are functions of , hence simultaneously diagonalisable, so per eigenvalue the requirement is where . That needs constant in . It is not:

0.250.5124
0.2070.3890.7241.3302.415

A factor of ~12 across a 16× spread. Calibrating at therefore

  • over-regularises high-variance directions (: about too strong),
  • under-regularises low-variance ones (: about ).

Suppressing exactly the directions that carry the most signal variance is an over-smoothing bias — and over-smoothed, mode-seeking reconstructions are precisely what RED-Diff is criticised for empirically. The theory predicts the known artefact, from a five-line calculation, on the one data family where it can be checked.

5. Why the weighting is mode-seeking

is unbounded as : the most heavily corrupted times weigh most. Compare the ELBO’s , which is RED-Diff’s weight squared — so relative to a genuine likelihood bound, RED-Diff systematically reweights toward the coarse, low-frequency end of the process. That is a reasonable thing to want from a reconstruction method and a bad thing to want from a posterior, and the paper is explicit that it is chasing the former.

6. What this is not

  • Not a posterior. is a point mass; there are no error bars. The general case is derived in the paper’s §3 and dropped in its experiments, and dropped here too — it is the most valuable missing piece (The Diffusion Factor §5.2).
  • Not exact at its own fixed point — for the regulariser it starts from. For an exact score its field is the exact gradient of a smoothed log-density (Implicit Diffusion Learners §3), so it is exact for a different, computable prior.
  • (original wording) Not exact at its own fixed point. Unlike the algebraic and equilibrium families, which are approximate because they stop early, this one converges to the wrong point by construction. Inversions and Bayesian Lenses permits inexact inversions and the free energy is supposed to measure the cost — but here the cost is a Jacobian nobody computes, so it is not measured either.
  • Not trainable here — by RED-Diff itself. Training through inference is now possible by the adjoint: Backpropagation through Implicit Inference.

Related: The Diffusion Family, The VP-SDE, The Diffusion Factor, ProxDM and Proximal Alternatives, Factors are Parameterized Statistical Games, Implicit Learners, ImplicitREDDiff, Bethe Free Energy