The VP-SDE of Song et al. 2021, presented through its perturbation kernel rather than its SDE.
Sources: code:
schedule.jl
1. What is implemented
with closed-form marginals
Because is affine, in closed form. That is the only reason the VP-SDE is convenient here: a schedule needing quadrature for would put a numerical integral inside every message.
The defining identity is — variance preserving: unit-variance data has unit variance at every . It is asserted in the test suite at eight values of , because it is the one property the whole file exists to have.
2. The interface is the kernel, not the SDE
AbstractNoiseSchedule requires only alpha(s,t) and sigma(s,t). drift and
diffusion are provided and nothing in this package calls them — RED-Diff replaces
sampling by optimisation, so the SDE is never integrated.
This is a deliberate narrowing and it has a cost: a schedule whose marginal is not Gaussian does not fit. Cold diffusion, discrete/categorical diffusion and flow matching with a non-Gaussian path all fall outside. What does fit trivially is sub-VP and VE — both are another pair of formulas — and neither is implemented.
3. Three weightings, one integrand
denoising_loss returns the unweighted . That is
not laziness: the weighting is where the objective’s identity lives.
| gives | |
|---|---|
| the ELBO / exact likelihood bound | |
| DDPM’s “simple” loss — the one everybody trains | |
| RED-Diff’s regulariser |
So the same network, the same integrand, three different meanings. reddiff.md §3 is about
what the third choice does.
4. Implementation difficulties
4.1 tmin is a floor everyone has and nobody documents
At , : the score divides by zero and
RED-Diff’s weight vanishes. VPSDE therefore carries
tmin = 1e-3 and sample_time draws from .
This is a modelling choice smuggled in as a numerical guard. It changes the value of every
integral in the package — including calibrate_lambda, whose answer depends on
tmin through the lower limit. Two implementations with different floors compute different
posteriors and neither is wrong. Recorded because the constant is invisible in the maths and
load-bearing in the code.
4.2 sigma must not be sqrt(1 - alpha^2)
Near , and sqrt(1 - alpha^2) evaluates
, losing roughly half the significant digits to cancellation. The
code uses sqrt(-expm1(-B(t))) instead, which is accurate to full precision. Asserted in the
test suite at against the small- asymptote .
4.3 Times are scalars, so a batch is not batched
alpha(s, t) takes a scalar t, and perturb broadcasts it over x. A real training loop
draws a different per batch element. Nothing here forbids it — t could be an array
and the broadcasts would work — but nothing tests it either, and sample_time returns one
scalar. This package is written for the inference-time inner loop, where one per step is
what RED-Diff’s Algorithm 1 asks for.
4.4 No discretisation, hence no sampler
There is no ancestral sampler, no Euler–Maruyama, no probability-flow ODE, and therefore no way to draw from the prior. That is consistent — RED-Diff never samples — but it means the package cannot be sanity-checked by generating from the model, only by the analytic oracle in the test suite.
Related: The VP-SDE, predictor, reddiff