implementation

Sources: code: lens.jl

Theory (CT-ML wiki): Dagger Category · Lax Functor · Bayesian Inversion · Bayesian Lens · Variational Free Energy · Reverse Derivative Category · Lens · Functor · Statistical Game

Implements: Inversions and Bayesian Lenses (Defs. 9, 10), Composition of Bayesian Lenses (Defs. 12, 15).

The idea in one comparison

The prior is the linearisation point. A put needs the cached forward value; an invert needs the prior it was conditioned on. BayesianLens(model, inversion) is the pair.

The inversion zoo

typerealisationfamily
ExactInversionBayes’ lawconjugate
AmortisedInversion(net)a Lux layerVAE encoder
SolverInversion(solver)root-findingalgebraic, DEQ, NeuralODE
ProximalInversion(prox)denoising stepsRED-Diff, ProxDM
TrivialInversionnothing to inferpriors, clamped channels

All five inhabit the same slot. That is the point of Definition 9: is any function of the right type, and its quality is measured by the free energy rather than enforced by the type. The three solver-ish ones are the three families of Implicit Learners.

Implementation difficulties

1. A factor has two independently parametrised halves

AmortisedInversion holds a Lux layer with its own ps. So does the forward model. This is the structural reason a factor cannot be an AbstractLuxLayer: a Lux layer has one direction and one parameter tree; a Bayesian lens has two directions and (often) two trees.

Practical consequence: the parameter tree of a factor is naturally (; model = ..., inversion = ...), and the two halves may want different optimisers — in a VAE the encoder and decoder are usually trained together, but in wake-sleep or in amortised inference with a fixed generative model they are not. The interface must not assume a single ps.

2. invert must return , not just

Easy to get wrong, and it fails silently: dropping the latent part gives a belief of the right type that breaks the chain rule, because the next factor upstream consumes exactly that latent. Documented on the invert docstring; will need a test once a concrete inversion exists.

3. TensorLens is lossy and there is no way around it

Remark 16: parallel composition of inversions can only feed each branch the marginal of a joint prior, so and the composite is mean-field. The gap is the mutual information between branches (Remark 26).

This is not a defect to fix. It is the formal content of “mean-field VI is wrong, and by this much”. The honest options are (a) accept it and report the gap, or (b) refuse to factor across correlated branches and keep a joint inversion. Both should be available at the graph level. Currently TensorLens just holds the parts and the docstring warns; the gap is not computed.

4. Exact inversions are only almost surely well defined

Footnote 3 of the paper: inversions may not be fully supported and are defined only up to a.s. equality, so is only a.s. a pseudofunctor. Numerically this is “conditioning on a null set” — a zero denominator in Bayes’ law. ExactInversion will need a support check and a policy (throw? return TrivialBelief? widen with a floor?) before it does anything. Unresolved.

5. compose argument order

compose(c, d) builds — the arguments follow the wiring (first c, then d) while the mathematical notation is written right-to-left. This will confuse anyone reading the paper with the code open, so the docstring says it explicitly. The alternative, matching the notation, would confuse everyone else.

Related: Inversions and Bayesian Lenses, Composition of Bayesian Lenses, Implicit Learners, open_model