model derivation

Assembling everything into AutoBayes Definition 20 and Definition 27. Every abstract slot gets a concrete, computable inhabitant — which is the point of doing the algebraic case first.

Sources: original to this vault (design and analysis; no single paper).

Theory (CT-ML wiki): Lax Functor · Bayesian Inversion · Statistical Game · Variational Free Energy

The quadruple

A parameterized statistical game is with . For an algebraic factor of degree on channels with generators:

slotAutoBayesalgebraic realisation
parameter space any space, — The Parameter is a Grassmannian
(update space)tangent bundle (Rem. 30), via
forward kernel root-solve in the chosen polarity
latent space scratchthe branch index — Branches and the Discriminant
inversion root-solve in the opposite polarity, weighted by
vector energy
energy space ours — the residual noise space, Algebraic versus Geometric Distance
scalarisation ours; (algebraic) or (Sampson)
vector entropy over branches, times a direction
gradient couplingDef. 29’s laxnessExactCoupling — the IFT gives the exact Jacobian

Every entry is computable. That is not true of the other two families in Implicit Learners, and it is why this one is worth doing even though it does not scale: it is the case where the framework can be debugged.

The vector energy is forced, not chosen

is -valued by construction. There is no scalar residual — a variety of codimension needs equations. So the algebraic family is the existence proof for Scalar and Multivariate Energy: it is a family in which the paper’s -valued cannot be used without first throwing away the model, because

  • the well-posedness condition is , a statement about the vector residual’s Jacobian, unstatable for a scalar;
  • the adjoint solves , requiring the Jacobian;
  • the correct scalarisation is derived from the Jacobian.

The energy/entropy split is genuinely two things here

Proposition 18 insists energy and entropy behave differently. In this family the difference is visible:

Note the first term vanishes when inference succeeds exactly. So for a well-posed algebraic factor the free energy reduces to : the loss is entirely the branch entropy. That is a striking and correct consequence — an exactly-solvable factor is scored only by how ambiguous its answer was.

The energy becomes non-zero exactly when the system is overdetermined () or the solver fails. So the two terms cleanly separate model misfit (energy) from inferential ambiguity (entropy), which is precisely the reading the paper’s decomposition is meant to support.

Composition: local only

By Composition is Elimination the factors must not be fused. The composite game is formed by the Definition 22 laws applied to the local data:

The energy direct sum here is concrete: the composite residual is the stacked residual vector, graded by factor. That is exactly the of Scalar and Multivariate Energy, and in this family it is literally how you would write the joint system down — which is a good sign the grading is the right abstraction rather than an imposition.

The expectation in the entropy law is over branches, so it is a finite sum, not a Monte Carlo estimate. In this family the chain rule is exact.

Learning: EM, with both halves in closed form

From Fitting is a Nullspace Problem §“missing channels”:

  • E-step the inversion : root-find the unobserved channels; obtain the branch belief .
  • M-step descent on : with branches fixed, the energy is quadratic in , so the update is the bottom- eigenspace of — a globally optimal M-step.

This instantiates Example 2 exactly, and adds something the paper does not claim: the M-step is not merely a descent step but the exact maximiser. That is generalised EM’s ideal case, and it comes free from the linearity in .

GradientCoupling is ExactCoupling

Definition 29 is lax in general because moves the pushforward prior and the sampling distribution. Here:

  • the “sampling distribution” is a finite branch set, and its dependence on is differentiable away from the discriminant, with derivative given by the IFT;
  • so the terms Definition 29 drops are computable, and the coupling can be ExactCoupling rather than DiagonalCoupling.

The exception is at the discriminant, where the branch count changes — a discrete jump that no derivative captures. So the honest statement is: exact away from , discontinuous on it. The GradientCoupling annotation should be ExactCoupling with a conditioning guard on .

What this factor cannot do

To be explicit, since the table above is otherwise flattering:

  • , (parameter count and sample complexity);
  • unobserved channels per message (Bézout);
  • coefficients must be fitted, not composed (Composition is Elimination);
  • no guarantee the learned real locus is nonempty or of the intended dimension (Varieties Ideals and Real Nullstellensatz);
  • gradients meaningless near the discriminant.

Related: Factors are Parameterized Statistical Games, Scalar and Multivariate Energy, Branches and the Discriminant, Open Problems in Algebraic Implicit Learning