definition theorem design

Is an implicit diffusion learner a parameterized statistical game? Yes, once a polarity is chosen: the composite of a prior game, whose energy is the smoothed negative log-density, and a likelihood game, whose energy is the clamp. Its inversion is the proximal inference of Implicit Diffusion Learners §4. Two things are missing: a real entropy for the inversion (the posterior is a Dirac) and, for a learned non-conservative network, a scalar loss. The two ways of training it — the game’s block-diagonal gradient and the Lagrangian of the original sketch — are different semantics, and both are now implemented.

Sources: St Clere Smithe & Perin, AutoBayes, arXiv:2503.18608, Definitions 1, 20, 22, 27–29, Theorem 23, Remarks 24, 30; Mardani et al. arXiv:2305.04391 §3 (the KL decomposition behind RED-Diff); Ho, Jain & Abbeel, Denoising Diffusion Probabilistic Models, NeurIPS 2020 (the denoising loss as a likelihood bound); code: implicit.jl, factor.jl, statistical_game.jl

Theory (CT-ML wiki): Statistical Game · Bayesian Lens · Open Model · Variational Free Energy · Para Construction · Lax Functor

1. The polarity chooses the game

A diffusion model is one joint density over ; it has no direction. A statistical game does. The polarity supplies it, with the dictionary of Channels and Polarity:

AutoBayeshere
(unobserved, solved for)output coordinates,
(observed, clamped)input coordinates, or large
(latent)latent coordinates

Different polarities give different games from the same network: a -element family, one for each way to read the relation (The Diffusion Factor §2).

2. The elements, one by one

Following RED-Diff’s derivation, minimising splits into three parts:

Every element of Definition 20 then has a home:

Definition 20 / 27the implicit diffusion factorstatus
generative model, prior on and kernel the diffusion joint , read as prior on times conditional of present, implicitly: only its (smoothed) score is available
inversion proximal inference, implicit_inferpresent; a Dirac at
energy (pointwise)the clamp, of a Gaussian observation of precision present
the prior as its own game (Remark 24)energy , the smoothed of Implicit Diffusion Learners §3, entropy 0present for an exact score; for a learned network only exists (the residual)
entropy of the inversionmissing: for a Dirac; needs
loss , up to the entropy constantpresent for an exact score and for an energy-parametrised network (implicit_energy, energy); a denoising-loss estimate otherwise
parameters (Definition 27, )network weights; also the precisions and λpresent; is differentiable by the adjoint

A correction to RED-Diff as a Statistical Game §2

That note files the score-matching term as the game’s entropy and then warns that it “is not an entropy”. The decomposition above puts it where it belongs: it is the energy of the prior game (Remark 24, “priors are games too”), a cross-entropy . The entropy slot holds , the inversion’s own entropy. The composite of prior game and likelihood game then has loss , the variational free energy with the smoothed prior, by Theorem 23. The remaining defect is not a mislabelled term but a degenerate .

3. Composition

Two implicit diffusion factors sharing a coordinate compose like any games: energies add, entropies chain (Definition 22), and with vector energies the sum becomes a direct sum (Scalar and Multivariate Energy). The multivariate version is the natural one here, because the residual is defined for every network while the scalar loss is not (§2, prior-game row). The composite residual of a factor graph is the stacked residuals of its factors, and inference on the graph is a root of the stack.

What composition does not fix is the message problem of The Diffusion Factor §4: a factor whose prior cannot be divided out sends posteriors, not likelihoods.

4. Two ways to train it, and what each means

(A) The game’s gradient: block-diagonal, generative

Definition 29 differentiates the loss with the inversion held fixed, i.e. the DiagonalCoupling of Factors are Parameterized Statistical Games:

Since is, up to -independent constants, a weighted denoising loss (an upper bound on under the ELBO weighting), this is denoising-score-matching training on the completed configuration . Alternating with inference is EM with a diffusion prior: E-step, infer the missing coordinates; M-step, train the denoiser on the completed data. It learns the whole relation, because it fits a density.

(B) The Lagrangian of the original sketch: exact through the inversion, discriminative

Differentiating a downstream loss through inference keeps the term Definition 29 drops: the dependence of the inversion on (Remark 30’s laxness). The adjoint of Backpropagation through Implicit Inference computes it exactly. This is supervised implicit learning, the way a DEQ is trained.

(A) game gradient(B) bilevel / adjoint
differentiates through inferenceno (stop-gradient)yes (implicit function theorem)
objectivefree energy, i.e. density fita task loss on the output
needsparameter gradient of the denoising lossone adjoint solve, plus one parameter-VJP per node
learnsevery branch of the relationthe branches inference visits (§7 of the backpropagation note: the old upper arc of the circle survived)
coupling in Factors are Parameterized Statistical GamesDiagonalCouplingExactCoupling on the inversion

They are complementary. (A) shapes the density; (B) shapes the answer the inference procedure actually returns, including its smoothing bias, which (A) ignores. A practical scheme is (A) to pretrain the relation and (B) to fine-tune for a task.

5. The deterministic relaxation as a game

With the noise-free single-level relaxation (Deterministic Relaxation §3) the prior energy is at one level and the inversion is the denoiser’s fixed point. That is the same game as a DEQ factor with an energy, except that the energy has a density reading. The equilibrium family’s factors “contribute no entropy” (DEQ as a Relation §5); this one contributes a prior energy, and would contribute an entropy if its inversion were Gaussian.

6. What is left to do

  1. A Gaussian inversion. Keep in RED-Diff’s variational family, iterate on , and becomes finite. The Laplace approximation at is already computed by the adjoint, so a first version costs nothing extra.
  2. A scalar loss for learned networks — done for energy-parametrised networks: with the field is the gradient of implicit_energy, exactly (energy). A network that outputs directly still has only the denoising-loss estimate.
  3. AD for epsilon_vjp_params on a Lux network — done, through any AD backend (backends).
  4. The factor interface — done: DiffusionFactor(…; prox = ImplicitProx(nodes)) inverts with the deterministic solver; implicit_solution reports, implicit_factor_pullback differentiates (implicit_factor).

Related: Implicit Diffusion Learners, Inference Signatures, Backpropagation through Implicit Inference, Deterministic Relaxation, RED-Diff as a Statistical Game, Factors are Parameterized Statistical Games, Scalar and Multivariate Energy