Lenticulum
The user-facing package: Gaussian beliefs and the linear-Gaussian factors built on them. If you want to try the framework on something, start here — everything in this package has a closed form, so you can check the answers.
Gaussian beliefs
GaussianBelief stores a Gaussian in canonical form — information vector η = Λμ and precision Λ = Σ⁻¹ — rather than as a mean and covariance. Two reasons:
- Pooling becomes addition. Combining two beliefs adds their canonical parameters, which is exact, associative, commutative and total.
- Improper beliefs are representable. A factor → variable message is a likelihood, and a likelihood constraining only some directions has a singular
Λ. The moment form cannot write that down at all.
g = Gaussian([1.0], fill(2.0, 1, 1)) # from mean and covariance
belief_mean(g), belief_cov(g) # back again
uninformative(1) # η = 0, Λ = 0
isproper(g) # is Λ positive definite?belief_mean on an improper belief throws rather than returning a plausible-looking pseudo-mean.
The factors
| factor | is | channels |
|---|---|---|
GaussianFactor | $p(y \mid x) = \mathcal{N}(y; Ax+b, Q)$ | two, named |
GaussianPrior | $\mathcal{N}(\mu, \Sigma)$ on one channel | one |
LinearConstraintFactor | $0 = \sum_i A_i x_i - c + \varepsilon$ | any number |
GaussianFactor is bidirectional: both (x observed, y unobserved) and the reverse, each exact. Parameters are (A, b); noise is a fixed hyperparameter, not learned.
LinearConstraintFactor is the acausal generalisation — one equation over n channels, none of them distinguished, with one supported polarity per channel. It subsumes GaussianFactor exactly (same messages, same energy) and is what you want for conservation laws, balance equations and anything written as a relation rather than an assignment:
# Kirchhoff's current law at a three-way node: i₁ + i₂ + i₃ = 0
f = LinearConstraintFactor((i1 = 1, i2 = 1, i3 = 1), 1; noise = 1e-8)
ps = (A = (i1 = ones(1,1), i2 = ones(1,1), i3 = ones(1,1)), c = [0.0])Known gaps
Qis a fixed hyperparameter everywhere; learning it needs a positive-definite parametrisation, and a contrastive term to stop it collapsing to zero.- A
LinearConstraintFactorreturns an uninformative message if any non-target channel is unconstrained, which is blunter than necessary. - A loopy graph of constraint factors cannot bootstrap without a weak prior on every variable.
API
Lenticulum.LenticulumLenticulum.GaussianBeliefLenticulum.GaussianFactorLenticulum.GaussianPriorLenticulum.LinearConstraintFactorLenticulum.LinearConstraintModelLenticulum.LinearGaussianModelLenticulum.GaussianLenticulum.belief_covLenticulum.belief_meanLenticulum.isproperLenticulum.kl_divergenceLenticulum.logpartitionLenticulum.residualLenticulum.residualLenticulum.residual_statisticsLenticulum.residual_statisticsLenticulum.uninformative
Lenticulum.Lenticulum — Module
LenticulumImplicit machine learning on factor graphs: learning relations $R_\theta \subseteq X_1\times\cdots\times X_n$ rather than functions.
This package is the user-facing layer, mirroring Lux.jl. It sits on
LenticulumCore— what a factor is (a parameterized statistical game);Mycelium— how factors are wired and in what order they talk.
What is here so far
The linear-Gaussian factor and its beliefs: the first factor in the project that implements the entire interface with nothing stubbed — a real open model, an exact Bayesian inversion, a genuine vector energy and entropy, bidirectional polarity, and learnable parameters.
Its purpose is to be a test oracle. Every quantity it produces has a closed form, so the framework's claims can be checked against arithmetic instead of against itself. Two of them turned out to be wrong; see The Gaussian Factor.md §"What this caught".
beliefs.jl—GaussianBeliefin canonical form, which makesMycelium.combineaddition and thereby unblocks message passing.gaussian.jl—GaussianFactorandGaussianPrior.constraint.jl—LinearConstraintFactor, the acausal n-ary sibling ofGaussianFactor: one equation $0 = \sum_i A_i x_i - c + \varepsilon$ over any number of channels, none of them distinguished. This is ModelingToolkit's0 ~ ...equation as a statistical game; seeModelingToolkit as an Acausal Relation.md.
Lenticulum.GaussianBelief — Type
GaussianBelief(η, Λ)A Gaussian in canonical form: information vector η = Λμ and precision Λ = Σ⁻¹.
Λ may be singular or zero, and that is not an error: a factor → variable message is a likelihood, not a distribution, and a likelihood that constrains only some directions has a rank-deficient precision. The moment form cannot represent this at all, which is the second reason for the canonical parametrisation.
Use Gaussian to build one from a mean and covariance.
Lenticulum.GaussianFactor — Type
GaussianFactor(nin => nout; noise = 0.1, channels = (:x, :y), init_weight, init_bias)The linear-Gaussian conditional $p(y \mid x) = \mathcal{N}(y;\ Ax + b,\ Q)$.
noise is Q: a matrix, a vector of variances, or a scalar variance. It is a fixed hyperparameter, not a parameter — learning Q needs a positive-definite parametrisation, which is a separate concern (see gaussian.md §3).
Parameters are (A, b), produced by initialparameters exactly as a Lux layer's are: the factor is a description, ps is an inhabitant.
Bidirectional: both (x observed, y unobserved) and (x unobserved, y observed) are supported, and each is exact. That is what makes it a relation rather than a function, and it is the smallest non-trivial witness that assemble — the operation with no Lux counterpart — does real work.
Lenticulum.GaussianPrior — Type
GaussianPrior(channel, μ, Σ)The prior factor $\pi : 1 \multimap X$ of AutoBayes Remark 24: emits $\mathcal{N}(\mu,\Sigma)$, charges energy $-\log p_\pi(x)$, and has zero entropy of its own.
Without it the composite free energy is an open free energy, missing the $-\log p_\pi(x)$ term. With it, $F^{c\pi}(\ast, y) = \mathrm{VFE}(c,c')(\pi,y)$, and when the inversion is exact that equals $-\log p_{c_*\pi}(y)$ on the nose. That identity is the project's main correctness test.
Non-learnable in v0: a learnable prior wants its natural parameters in ps, and Σ needs a positive-definite parametrisation — the same deferred concern as Q in GaussianFactor.
Lenticulum.LinearConstraintFactor — Type
LinearConstraintFactor(dims::NamedTuple, m::Int; noise = 0.1, init_coeff, init_constant)The soft linear relation $0 = \sum_i A_i x_i - c + \varepsilon$, $\varepsilon \sim \mathcal{N}(0, Q)$ — equivalently the density
\[f(x_1,\ldots,x_n) \;\propto\; \exp\Bigl(-\tfrac12 \bigl\|\textstyle\sum_i A_i x_i - c\bigr\|^2_{Q^{-1}}\Bigr)\]
dims names the channels and their dimensions, m is the dimension of the residual (the number of scalar equations). Q is a fixed hyperparameter, exactly as in GaussianFactor; parameters are (A, c), with A a NamedTuple keyed by channel.
Acausal: no channel is privileged. supported_polarities returns one polarity per channel — "solve this equation for x_t, given the others" — and every one of them is exact. As $Q \to 0$ the factor becomes the hard relation $\{x : \sum_i A_i x_i = c\}$; a finite Q is the soft relation, which is what makes it a statistical game rather than a set.
# Kirchhoff's current law at a three-way node: i₁ + i₂ + i₃ = 0
f = LinearConstraintFactor((i1 = 1, i2 = 1, i3 = 1), 1; noise = 1e-8)
ps = (A = (i1 = ones(1,1), i2 = ones(1,1), i3 = ones(1,1)), c = [0.0])See constraint.md for the message algebra and the honest list of what is missing.
Lenticulum.LinearConstraintModel — Type
LinearConstraintModel(factor, polarity)The open model the polarity selects: "solve this equation for the unobserved channel". Pure — a linear relation with additive Gaussian noise hides nothing.
Lenticulum.LinearGaussianModel — Type
LinearGaussianModel(factor, polarity)The open model $c : X \nrightarrow Y$ of AutoBayes Definition 1, in the direction the polarity selects. Pure ($\llbracket c \rrbracket \cong 1$): a linear-Gaussian conditional hides nothing.
Lenticulum.Gaussian — Method
Gaussian(μ, Σ)
Gaussian(μ::Real, σ²::Real)A Gaussian from its moments. Converts to canonical form immediately; Σ must be positive definite.
Lenticulum.belief_cov — Method
belief_cov(b) -> Matrix$\Sigma = \Lambda^{-1}$.
Lenticulum.belief_mean — Method
belief_mean(b) -> Vector$\mu = \Lambda^{-1}\eta$. Errors on an improper belief, rather than returning a least-squares pseudo-mean that would be silently wrong.
Named belief_mean rather than extending Statistics.mean to avoid a fifth name collision in this project — see Mycelium.md §2.
Lenticulum.isproper — Method
isproper(b::GaussianBelief) -> BoolWhether Λ is positive definite, i.e. whether the belief is a normalisable distribution rather than a likelihood. belief_mean, belief_cov, variable_entropy and belief_logdensity all require this.
Lenticulum.kl_divergence — Method
kl_divergence(q::GaussianBelief, p::GaussianBelief) -> Real$D_{KL}(q \,\|\, p)$. Both must be proper. Used to measure the laxness quantities the framework keeps insisting should be reported rather than hidden — see Composition of Bayesian Lenses.md Remark 16.
Lenticulum.logpartition — Method
logpartition(b) -> Real$\log \int \exp(-\tfrac12 x^\top\Lambda x + \eta^\top x)\,dx = \tfrac12\eta^\top\Lambda^{-1}\eta - \tfrac12\log\det\Lambda + \tfrac n2\log 2\pi$.
Lenticulum.residual — Method
residual(f, x, y, ps) -> Vector$r = y - Ax - b$: the pointwise vector energy of Scalar and Multivariate Energy.md, in the residual space where the observation noise lives (Algebraic versus Geometric Distance.md §"what the energy space is").
Q⁻¹ is the noise precision, hence the natural scalarisation's metric — derived, not chosen.
Lenticulum.residual — Method
residual(f::LinearConstraintFactor, vals::NamedTuple, ps) -> Vector$r = \sum_i A_i x_i - c$, the pointwise vector energy, in the space where the equation's own noise lives.
Note the signature: a NamedTuple of all channel values, with no input/output split, because the factor has none. See constraint.md §3 for why LenticulumCore.energy's (x, a, y) signature cannot express this directly.
Lenticulum.residual_statistics — Method
residual_statistics(f, msgs, ps) -> (r̄, Cov_r)The mean and covariance of the residual $r = y - Ax - b$ under the factor's own belief $b_c \propto f_c \prod_i \mu_{i\to c}$.
These are what make the two energies comparable on this factor:
\[\underbrace{\mathbb{E}\bigl[\tfrac12 r^\top Q^{-1} r\bigr]}_{\text{scalar (exact)}} \;-\; \underbrace{\tfrac12 \bar r^\top Q^{-1}\bar r}_{\text{multivariate (lax)}} \;=\; \tfrac12\operatorname{tr}\bigl(Q^{-1}\operatorname{Cov}(r)\bigr)\]
which is precisely the Jensen gap $\tfrac12\operatorname{tr}\operatorname{Cov}$ of Scalar and Multivariate Energy.md §5, here in closed form on a real model. Asserted in the test suite.
Lenticulum.residual_statistics — Method
residual_statistics(f::LinearConstraintFactor, msgs, ps) -> (r̄, Cov_r, H)Mean and covariance of $r = \sum_i A_i x_i - c$ under the factor's own belief $b_c \propto f_c \prod_i \mu_{i\to c}$, plus that belief's entropy.
Dirac channels are substituted out rather than represented: a clamped channel moves its contribution into the constant, $c \mapsto c - A_j v_j$, and drops out of the joint. That is what keeps the joint precision finite, and it is the n-ary version of the three hand-written cases in gaussian.jl. With every channel clamped the residual is deterministic and the entropy is zero, which falls out of the same code rather than needing its own branch.
Lenticulum.uninformative — Method
uninformative(n)The improper uniform belief on $\mathbb{R}^n$: η = 0, Λ = 0. The identity for Mycelium.combine, and the canonical-form spelling of LenticulumCore.TrivialBelief.
LenticulumCore.energyspace — Method
energyspace(f::GaussianFactor)$E_c = \mathbb{R}^3$, graded as (fit, complexity, negentropy):
| summand | is | role |
|---|---|---|
fit | $\mathbb{E}_{b_c}[\tfrac12 r^\top Q^{-1} r]$ | how badly the residual misses |
complexity | $\tfrac12\log\det(2\pi Q)$ | the log-normaliser; what stops $Q \to 0$ |
negentropy | $-H(b_c)$ | the entropy term of the free energy |
The grading is not decoration: it is the fit/complexity/entropy decomposition every Gaussian model has, made visible for free by $E_G = \bigoplus_f E_f$. Summing it recovers the exact free energy, so the scalarisation is IdentityScalarisation on each summand and is therefore linear — hence composition is strict, per Scalar and Multivariate Energy.md §5.
The alternative grading — keep the residual $r$ as a vector with SquaredNorm(Q⁻¹) — is lax, and by exactly $\tfrac12\operatorname{tr}(Q^{-1}\operatorname{Cov}(r))$. residual_statistics exposes both sides of that identity; see gaussian.md §2.
LenticulumCore.invert — Method
LenticulumCore.invert(lens, π, inputs, ps, st) -> (belief, st)The exact Bayesian inversion $c^\dagger_\pi(y)$ of AutoBayes Definition 10 — the posterior, including the prior.
[!warning] This is not the belief-propagation message $c^\dagger_\pi(y)$ is the posterior; a BP message is the likelihood, with the prior divided out. If a factor sent the posterior, a variable of degree $d$ would count the prior $d$ times. In canonical form the two differ by one subtraction, and
Mycelium.factor_messagereturns the likelihood while this returns the posterior.The relation between them is exactly
combine(π, message) == posterior, which is asserted in the test suite. On a chain (every variable of degree 2) the distinction is invisible, which is why it does not appear in the paper. Seegaussian.md§4.
LenticulumCore.supported_polarities — Method
LenticulumCore.supported_polarities(f::LinearConstraintFactor)One polarity per channel: that channel Unobserved(), every other Observed().
Contrast GaussianFactor, which has exactly two regardless of anything. Here the count is n, and none of them is the factor's "real" direction — the equation was never written in a direction. This is the enumerated form of "a relation, not a function".
LenticulumCore.supports_polarity — Method
LenticulumCore.supports_polarity(f::LinearConstraintFactor, p)Any polarity over exactly this channel set with exactly one Unobserved() channel.
Latent() channels are accepted by this predicate but produce an uninformative message — see Mycelium.factor_message below. Accepting-then-returning-nothing rather than refusing is deliberate: a latent channel is a legal modelling choice (marginalise it out), and the answer "this equation tells you nothing about x_t once x_j is free" is a correct answer, not an error.
Mycelium.belief_distance — Method
Mycelium.belief_distance(a::GaussianBelief, b::GaussianBelief)$\max(\|\Delta\eta\|_\infty, \|\Delta\Lambda\|_\infty)$ on the canonical parameters.
Chosen over a KL divergence because it is defined for improper beliefs too, and messages are routinely improper. A convergence criterion that throws on half the messages is not a convergence criterion. kl_divergence is available separately when both beliefs are proper.
Mycelium.belief_logdensity — Method
Mycelium.belief_logdensity(b::GaussianBelief, x)``\log p_b(x) = -\tfrac n2\log 2\pi + \tfrac12\log\det\Lambda
- \tfrac12 (x-\mu)^\top\Lambda(x-\mu)``. Requires a proper belief.
Mycelium.combine — Method
Mycelium.combine(a::GaussianBelief, b::GaussianBelief)Addition of canonical parameters. Exact, associative, commutative, total — no approximation, no density evaluation, no failure case.
This is the operation messages.md §1 records as the package's blocking gap. For Gaussians it is free; for anything else it is importance reweighting. That asymmetry is why Gaussian belief propagation is the workhorse it is.
Mycelium.factor_message — Method
Mycelium.factor_message(f::GaussianFactor, target, polarity, inputs, prior, ps, st)The factor → variable message: the likelihood contribution, in canonical form.
Forward ($x$ observed, target $y$), for an incoming $\mathcal{N}(m, S)$ on $x$:
\[\mu_{c\to y} = \mathcal{N}\bigl(Am + b,\ ASA^\top + Q\bigr)\]
Backward ($y$ observed, target $x$), for an incoming $\mathcal{N}(m_y, S_y)$ on $y$, writing $R = Q + S_y$:
\[\mu_{c\to x} = \mathcal{N}^{-1}\bigl(A^\top R^{-1}(m_y - b),\ A^\top R^{-1}A\bigr)\]
Note the backward message is generally improper — $A^\top R^{-1} A$ is rank-deficient whenever A is not full column rank, because a likelihood constrains only the directions the model can see. The moment form cannot express this; the canonical form does, which is the second reason beliefs.jl uses it.
The prior argument is deliberately unused: including it would double-count. See LenticulumCore.invert above.
Mycelium.factor_message — Method
Mycelium.factor_message(f::LinearConstraintFactor, target, polarity, inputs, prior, ps, st)The factor → variable message, in canonical form. Writing $d = c - \sum_{i \ne t} A_i m_i$ and $R = Q + \sum_{i \ne t} A_i S_i A_i^\top$ for incoming $\mathcal{N}(m_i, S_i)$:
\[\mu_{c\to t} \;=\; \mathcal{N}^{-1}\bigl(A_t^\top R^{-1} d,\ A_t^\top R^{-1} A_t\bigr)\]
One formula for every direction — the target channel appears only as "which $A_i$ is $A_t$". Specialising to two channels with $A_x = -A$, $A_y = I$, $c = b$ reproduces both of GaussianFactor's messages exactly, forward and backward.
The message is improper whenever $A_t$ is not full column rank — an equation constrains only the directions it can see, and with fewer equations than unknowns ($m < \dim x_t$) that is the normal case, not an edge case. Only the canonical form can say this; see beliefs.md.
Returns TrivialBelief() if any non-target channel carries no information: a free variable elsewhere in the equation makes the whole thing vacuous. See constraint.md §4 for the sharper answer this gives up.
Mycelium.local_free_energy — Method
Mycelium.local_free_energy(f::GaussianFactor, msgs, ps, st)The factor's Bethe contribution $F_c = U_c - H(b_c)$, as a graded energy (fit, complexity, negentropy), where
\[b_c(x,y) \;\propto\; \mathcal{N}(y;\,Ax+b,\,Q)\ \mu_{x\to c}(x)\ \mu_{y\to c}(y)\]
Note msgs are the incoming messages $\mu_{i\to c}$, not the variable marginals. That is what the Bethe formula asks for, and getting it wrong is the bug recorded in free_energy.md §7.
Mycelium.local_free_energy — Method
Mycelium.local_free_energy(f::LinearConstraintFactor, msgs, ps, st)$F_c = U_c - H(b_c)$, graded as (fit, complexity, negentropy) — the same decomposition as GaussianFactor, computed from residual_statistics.
Mycelium.variable_entropy — Method
Mycelium.variable_entropy(b::GaussianBelief)$H = \tfrac n2(1 + \log 2\pi) - \tfrac12\log\det\Lambda$.
Note this can be negative — differential entropy of a concentrated Gaussian is negative, and the Bethe counting correction of free_energy.md depends on the sign being carried correctly. This is the first belief type for which the correction is not identically zero, and implementing it is what exposed the sign bug recorded in free_energy.md §6.