Implemented in
lib/VariationalDiffusion.jl/src/schedule.jl; see schedule.
Sources: Song et al. 2021, arXiv:2011.13456 §3.4 — the forward corruption process, and the three quantities a trained hands you for free
Theory (CT-ML wiki): Statistical Game
1. One SDE, and why this one
Song’s paper unifies the two families that existed separately before it:
| SDE | drift | ends at | discrete ancestor |
|---|---|---|---|
| VP (variance preserving) | DDPM | ||
| VE (variance exploding) | SMLD / NCSN | ||
| sub-VP | , smaller noise | — |
VP is the one implemented, because it is DDPM’s continuous limit and therefore what almost every pretrained checkpoint speaks.
2. The perturbation kernel is the whole interface
Solving the SDE gives a closed-form Gaussian marginal:
and because is affine, in closed form — the only reason this schedule is convenient in an inner loop.
The defining identity:
Variance preserving: unit-variance data has unit variance at every . More generally , which equals iff — so the name is a statement about normalised data, and it is why every diffusion pipeline normalises to unit variance first. Feed it and the process is not variance preserving at all; it merely converges to anyway.
The implementation exposes the kernel, not the SDE
AbstractNoiseSchedulerequires onlyalphaandsigma.driftanddiffusionexist for completeness and nothing calls them, because RED-Diff replaces sampling with optimisation. A schedule with a non-Gaussian marginal — cold diffusion, discrete diffusion — does not fit this interface at all.
3. Three quantities from one network
A network trained on gives you three things, of which it only learned one:
The score identity is exact, not approximate: for the perturbation kernel,
so predicting the noise and estimating the score are the same task in different units. This is why -parametrisation won: it is the score, scaled to unit variance at every , which is far better conditioned to regress.
Tweedie’s formula is what makes a diffusion model a prior rather than a sampler: , the MMSE denoiser. For Gaussian data this is checkable in closed form, and the test suite checks it:
and denoise reproduces it to floating point. That agreement is the single most useful
sanity check available on this whole family, because it is the only place a diffusion model
has a closed form to be checked against.
4. The weighting is where the objective’s identity lives
The integrand is always the same. What multiplies it by decides what you are computing:
| objective | |
|---|---|
| the ELBO — an actual likelihood bound | |
| DDPM’s “simple loss” — what people train | |
| RED-Diff’s regulariser |
That the trained objective () is not the likelihood objective is well known and is why diffusion models are good samplers and mediocre density estimators. That the inference-time regulariser is a third weighting again is the subject of RED-Diff as a Statistical Game §4 — and it turns out to determine the strength of the implied prior.
5. The floor nobody documents
At : , so the score divides by zero and RED-Diff’s weight vanishes. Every implementation therefore samples from with , and almost none say so.
It is a modelling choice disguised as a numerical guard
appears as the lower limit of every integral in the package — including the λ calibration of RED-Diff as a Statistical Game §4, whose answer depends on it. Two implementations with different floors compute different posteriors, and neither is wrong.
A second numerical point, small but real: must be computed as , not . The latter evaluates near and loses half its significant digits to cancellation.
6. What is not here
- No sampler. No ancestral sampling, no Euler–Maruyama, no probability-flow ODE. Which is consistent — RED-Diff never samples — but means the model cannot be checked by generating from it.
- No VE or sub-VP, though both are just another pair of formulas.
- No training loop.
lib/VariationalDiffusion.jlconsumes a trained ; training it needs AD and this package deliberately has none (reddiff §2). - No per-element time.
alpha(s,t)takes a scalar; a real training batch draws a different per element. Fine for RED-Diff’s inner loop, wrong for training.
Related: The Diffusion Family, RED-Diff as a Statistical Game, schedule, predictor, Implicit Learners