open-problem

The honest ledger of the implicit diffusion learner, in two parts. Theoretical problems are open questions or intrinsic limits: nobody knows the answer, or there is provably no free lunch. Missing implementation is work whose design is known and that simply has not been done. Each entry links to where it was observed or measured.

Sources: original to this vault (design and analysis), collecting results from the notes and tutorials linked below; Du et al., Reduce, Reuse, Recycle, ICML 2023 (composition at ); Chung et al., Diffusion Posterior Sampling, ICLR 2023 (conditional sampling)

Theory (CT-ML wiki): Statistical Game · Bayesian Inversion

The companion ledger for the algebraic family is Open Problems in Algebraic Implicit Learning; several entries below are the same problem in a different family.

Part I. Theoretical problems

T1. Finding all answers — open; multi-start search implemented

A query on a multivalued relation has several stable roots (the circle’s , the robot arm’s elbow up and down). Point inference returns the root whose basin contains the start. There is no guarantee that a set of starts finds every stable root, and no way to know when all have been found. Sampling the conditional distribution would find them with the right frequencies, but then the answer is a distribution, and how to turn a learned relation into well-calibrated branch probabilities is itself open. implicit_roots now searches from several starts and returns every distinct stable answer it reaches (implicit §5), which finds both circle branches; completeness remains unguaranteed. Observed: Implicit Diffusion Learners §6; tutorials A relation without training and Robot arm.

T2. Bias against conditioning — intrinsic

The relation recovered is the ridge of a smoothed density, displaced inwards by about for a ring of radius . Smaller noise levels shrink the bias and make the field stiff and ill-conditioned ( near the data). There is no setting with neither; the open part is a principled choice of the levels for a given relation and data density. Measured: Implicit Diffusion Learners §5 (radius 0.98 at the default levels, 0 at RED-Diff’s range).

T3. Branch points — intrinsic

Where branches meet (the circle at ), smoothing merges them and creates saddles; the solver lands on them and reports stable = false. This is the smoothed form of the discriminant (Branches and the Discriminant; Open Problems in Algebraic Implicit Learning §7). Near a branch point the answer is ill-conditioned for every method, not only this one. Observed: tutorial Train a small diffusion model (unstable answers near ).

T4. What a query off the relation should return — open (semantics)

Clamping , just off the circle, returns the ridge point on that line with zero residual: the clamped problem has a root. Whether the right answer is that point, a refusal, or a point with a flag is a modelling decision without an agreed answer. Energy-parametrised models at least provide the number to decide with ( at the answer, energy §3); ε-networks provide nothing comparable.

T5. Coverage and out-of-distribution queries — open

The relation is only learned where data were seen. Outside, the field is extrapolation, and answers degrade silently (robot arm: misses up to 0.18 at targets needing angles outside the training range). Where data are sparse along the relation, the ridge can thin out or break. There is no measure of how well a learned relation covers a region, and no guarantee that it is connected where the true one is. Measured: tutorial Robot arm (in- vs out-of-distribution statistics).

T6. A non-conservative field defines no energy — intrinsic to ε-networks; resolved by energy networks at a cost

A network that outputs directly is not the gradient of anything (2.6% Jacobian asymmetry on the circle). Then the relation has no energy, “stable” uses the symmetric part of the Jacobian as a heuristic, and two answers cannot be compared. The energy parametrisation removes the problem by construction (energy); what remains open is whether the extra cost (about 3× in training, second derivatives everywhere) is always worth it, and how much the asymmetry of a well-trained ε-network matters in practice.

T7. Training through inference shapes only the visited branches — intrinsic to the bilevel objective

Backpropagating a task loss through inference moves the branches inference actually lands on. When a parabola was learned from a circle this way, the old upper arc survived as another branch. Density training (the game’s own gradient) shapes everything but ignores the task. How to combine the two with guarantees is open. Observed: Backpropagation through Implicit Inference §7; The Implicit Diffusion Factor as a Statistical Game §4.

T8. Hyperparameters without theory — partly open

The field nodes (levels, samples), the weighting λ and the scale of the precisions decide what relation is learned and how strongly evidence counts. λ is derivable for Gaussian data (calibrate_lambda, RED-Diff as a Statistical Game §4); in general it is not. Precision interacts with the field’s scale: on the trained circle, a precision of 1 already pulled an answer most of the way from the relation to the anchor.

T9. Composing diffusion factors at — known obstruction, approximate remedies

Noising does not commute with products, , so adding the scores of two diffusion factors is exact only at . Remedies (MCMC corrections at each level) exist but are approximate. See Language Models §7.

T10. Messages are posteriors, not likelihoods — open design problem

A diffusion factor carries its own prior, which cannot be divided out of its inversion, so it sends posteriors where a factor graph expects likelihoods, and neighbours double-count. See The Diffusion Factor §4.

T11. No guarantees — open

Nothing guarantees that the solver converges on a learned field, that a stable root exists for a given query, or that the relation is identifiable from data (which relations produce the same smoothed ridges?). In flat regions far from data, descent stalls and the solver reports it (Implicit Diffusion Learners §6).

Part II. Missing implementation

I1. All branches: mixture beliefs and conditional sampling

Design known: a mixture belief type (Belief Algebra §4, item 2), multi-start inference returning the distinct stable roots, and a conditional sampler (DPS-like, or ProxDM’s sampler with the clamp) for branch frequencies. None exists; point inference only.

I2. Uncertainty: a Gaussian answer — done (as a covariance)

implicit_laplace returns the Laplace covariance at an answer, and density_lambda chooses the weighting under which it is in data units (implicit §5). Still missing: returning it as a GaussianBelief (I3) and using its entropy in the statistical game (The Implicit Diffusion Factor as a Statistical Game §6).

I3. Gaussian messages from learned factors

GaussianBelief lives above the lib/ packages, so diffusion factors cannot emit Gaussians even with I2. Fix: move all belief types into one low layer (Belief Algebra §6).

I4. Automatic recovery from unstable answers — done

implicit_roots restarts from perturbed points and reports only converged, stable answers (implicit §5).

I5. Cost: batching and matrix-free solves

Each field evaluation calls the network once per node (128 times by default), sequentially; the Jacobian costs one derivative pass per coordinate, and the linear algebra is dense, . Known fixes: evaluate all nodes in one batched call (also the right shape for Reactant), and replace dense Newton by Newton–Krylov with Jacobian-vector products. Required before anything beyond a handful of dimensions (backends §5).

I6. Energy networks on every backend

Second derivatives work with Zygote and ForwardDiff-over-Zygote; Enzyme’s forward-over-reverse fails on Lux layers, and Reactant is not wired for energy models (energy §4, §7).

I7. ProxDM, completed

No adjoint for prox_infer, no ProxDM inversion for DiffusionFactor, no ad field on ProxNetwork, and the sampler is first order (spread 0.29 against 0.30) (proxdm §5).

I8. Corrected composition

MCMC correction steps for products of diffusion factors (T9’s remedy). Not built.

I9. Coverage diagnostics

A practical flag for T5: the energy at the answer (energy networks), the distance to the nearest training data, or a density estimate on . Any of them would mark out-of-distribution answers. Not built.

I10. Real data and higher dimensions

Everything validated is two- to four-dimensional and synthetic or closed-form. A real-data benchmark, and a problem with tens of coordinates, are the obvious next tests; I5 comes first.

I11. Nonlinear factors around the learner

Latent states with nonlinear dynamics (the particle tutorial’s noisy positions, SLAM) need nonlinear Gaussian factors and Gauss–Newton in the graph, which Lenticulum does not have yet.

Summary

problemkind
T1finding all answersopen; multi-start search implemented
T2smoothing bias vs conditioningintrinsic
T3branch pointsintrinsic
T4queries off the relationopen (semantics)
T5coverage, out of distributionopen
T6non-conservative ε-networksintrinsic; resolved by energy networks at a cost
T7training shapes only visited branchesintrinsic to the bilevel objective
T8hyperparameterspartly open
T9composition at known obstruction, approximate remedies
T10posterior messagesopen design problem
T11convergence, existence, identifiabilityopen
I1mixture beliefs, conditional samplingmissing
I2Laplace (Gaussian) answersdone as a covariance; not yet a belief
I3Gaussian messages from lib/missing (package layout)
I4restarts after unstable solvesdone
I5batched nodes, Newton–Krylovmissing, needed for scale
I6energy networks on Enzyme and Reactantmissing
I7ProxDM adjoint and factormissing
I8corrected compositionmissing
I9coverage diagnosticsmissing
I10real data, higher dimensionsmissing
I11nonlinear factorsmissing (Lenticulum-wide)

Related: Implicit Diffusion Learners, Inference Signatures, Backpropagation through Implicit Inference, The Diffusion Factor, The Implicit Diffusion Factor as a Statistical Game, energy, proxdm, backends, Belief Algebra, Open Problems in Algebraic Implicit Learning