model derivation

The observation that makes this whole family more than a curiosity for Lenticulum: conditional independence is a polynomial constraint, so a Bayesian network is an algebraic variety, and marginalisation is elimination. The algebraic learner and the AutoBayes factor graph are the same object seen twice.

Sources: original to this vault (design and analysis; no single paper).

Theory (CT-ML wiki): Lax Functor · Bayesian Inversion · Conditional Independence · Statistical Game · Lens · Bayesian Lens

Conditional independence is determinantal

Let be discrete with , , and let be the joint probability table, . Then

Independence is the vanishing of a set of quadrics. Conditional independence is the same statement slice-by-slice in : the minors of each conditional slice.

So the model “all distributions satisfying this CI statement” is a determinantal variety intersected with the probability simplex. This is the founding observation of algebraic statistics (Drton–Sturmfels–Sullivant, Lectures on Algebraic Statistics; Pistone–Riccomagno–Wynn).

Therefore a graphical model is a variety

A Bayesian network’s model — the set of joint distributions that factorise according to the DAG — is cut out by the CI statements implied by the graph (its global Markov property), each of which is a set of minors. Discrete graphical models are toric or determinantal varieties in the simplex.

This is exactly the object an algebraic implicit factor learns. Fitting a variety to data and fitting a graphical model to data are, in the discrete case, the same problem in different notation. In particular:

  • the Veronese parametrisation with spans all quadrics, hence all pairwise CI constraints;
  • the nullspace fit is then structure learning: which quadrics vanish on the empirical distribution is which CI statements hold;
  • the Grassmannian non-identifiability is the statement that many generating sets encode the same CI structure.

Marginalisation is elimination

Here the correspondence becomes sharp and useful. Consider a graphical model with a hidden variable . The marginal model — the set of distributions over the observed variables obtained by summing out — is the image of a polynomial map

Its Zariski closure is computed by elimination: eliminate the parameters and the hidden coordinates from the ideal. The resulting equations are the constraints the marginal model satisfies — classically, the tetrad constraints of factor analysis ( minors vanishing) are exactly the elimination ideal of a one-hidden-factor model.

Now line this up with Composition is Elimination:

algebraic geometryprobabilityAutoBayes
projection marginalisation pushforward
elimination ideal the marginal model’s equationsthe composite model
image is only constructible (Chevalley)the marginal model is only semialgebraiccomposition is lax
Zariski closure adds spurious pointsinequality constraints are lostthe 2-cell witnessing laxness
doubly exponential (Mayr–Meyer)marginalisation is P-hard”computing is expensive”

Every row is the same fact stated in a different language. The paper’s closing warning — that “priors are propagated by pushforward (i.e. marginalization), and computing these is similarly expensive to computing exact inversions”, so belief propagation is warranted — is the probabilistic form of “do not run Gröbner, keep the factors separate”.

Why this matters for Lenticulum

Composition is Elimination argues from complexity that you must not compose algebraic factors. This note shows the argument is not special to the algebraic family: it is the same argument as the one AutoBayes already makes for Bayesian factors. The two justifications for Mycelium.jl — algebraic and probabilistic — are one justification.

The semialgebraic gap, made concrete

Chevalley says the image of a variety is constructible; over , Tarski–Seidenberg says semialgebraic. In statistics this is not a technicality: hidden-variable models are famously not varieties. The set of distributions realisable by a latent-class model with classes is cut out by equations and inequalities, and the inequalities are essential — the “expected” variety contains distributions that are not achievable by any non-negative parameter values.

So the honest type of a marginalised graphical model is semialgebraic set, and:

  • the correct tool is real quantifier elimination (Collins’ CAD, doubly exponential) or Positivstellensatz certificates, not Gröbner bases;
  • an implementation that models the marginal as a variety is systematically too permissive, admitting parameter values that no distribution realises.

This is the precise algebraic content of the laxness theme, and it says the laxness is not merely quantitative (a KL gap) but type-level: the composite is a different kind of object than the parts.

What is worth borrowing

Algebraic statistics has fifteen years of results on exactly the objects Lenticulum needs, and they are directly transferable:

  • Model invariants — the elimination ideal of a graphical model is a checkable signature of its structure. Two graphs with the same invariants are indistinguishable from data; this is model identifiability, and it is decidable.
  • Maximum likelihood degree — the number of complex critical points of the likelihood on the model, the exact analogue of the Euclidean Distance Degree in Algebraic versus Geometric Distance. It counts how many local maxima EM can get stuck in. For many standard models it is known exactly.
  • Toric structure — decomposable/hierarchical models are toric varieties, for which everything (Gröbner bases, degree, ML estimation) is far better behaved. If a factor’s relation can be made toric, do so; it is the single most valuable structural property in the family.

Related: Composition is Elimination, Composition of Statistical Games, The Algebraic Factor as a Statistical Game, Factors are Parameterized Statistical Games