model derivation

Derivation. The parameter of an algebraic factor is not a matrix; it is a -plane. This is the cleanest concrete instance in the whole vault of Definition 2.5’s insistence that .

Sources: original to this vault (design and analysis; no single paper).

Theory (CT-ML wiki): Statistical Game · Parametric Lens · Lens · Gradient Based Learning with Parametric Lenses

The non-identifiability

If then , so

The variety, the ideal, and the relation all depend on only through its row space. The parameter is therefore genuinely a point of the Grassmannian

not of (dimension ). The missing dimensions are pure gauge.

This is not pedantry. Three things go wrong if you ignore it:

  1. Gradient descent on drifts along the gauge directions, changing the parameter without changing the model. Step sizes become meaningless; two runs that learned the same relation look maximally different.
  2. is a global minimiser of any residual objective, and it is reachable by descent. The normalisation of Fitting is a Nullspace Problem is not a regulariser bolted on to prevent collapse — it is a chart on , and collapse is the symptom of using the wrong space.
  3. The “spurious polynomial” pathology is the same phenomenon: a polynomial with tiny coefficients that “almost vanishes” everywhere is a point of near the origin, which is not a point of the Grassmannian at all.

The Riemannian gradient, derived

Represent a point by with (a Stiefel representative; the Grassmannian is the quotient by the action). The horizontal (tangent-to-) space at is

i.e. row-wise orthogonal to the current row space. Given a Euclidean gradient , the projection onto is

Check: . ✓

The update is then a retraction, not an addition:

where orthonormalises the rows (thin QR of the transpose, or a polar factor). Edelman–Arias–Smith is the standard reference for the geometry; the point here is only that .

This is exactly

Definition 2.5 carries a parameter pair with , and the vault’s note on it says the distinction matters “for a parameter constrained to a manifold, where the update lives in the tangent space, not in the manifold”. Here that is not a hypothetical:

and Definition 3.11’s gradient-update lens is type-incorrect. The correct lens is

which is still a perfectly good lens — the framework accommodates it without modification. The categorical setup was right and the naive implementation was wrong, which is a satisfying vindication of taking the distinction seriously.

And it is the natural gradient Definition 27 asks for

Definition 27 says the default semantics is gradient descent with respect to the Fisher information metric — the Bayesian learning rule. On the canonical (Riemannian) metric is restricted to the horizontal space, and above is the metric-correct gradient. So in this family the natural gradient is not an expensive approximation of an intractable Fisher matrix — it is a projection costing one matrix product.

Interaction with the rank-one structure

Backpropagation by the Implicit Function Theorem derives that each data point contributes a rank-one cotangent . So a minibatch of size gives of rank , and

is also rank . The Riemannian update is a rank- perturbation of a -plane. That is a well-studied object (subspace tracking; incremental SVD), and it means the update costs rather than — which is the difference between feasible and not at .

Caveat: the Grassmannian is the right space only for a fixed and

Model selection — choosing how many generators and what degree — moves you between Grassmannians of different dimensions, and there is no smooth path. The numerical rank decision of Fitting is a Nullspace Problem is therefore a discrete jump between manifolds, not a continuous shrinkage. No continuous relaxation of it is known that behaves well; see Open Problems in Algebraic Implicit Learning §4.

Related: Fitting is a Nullspace Problem, Parametric Lens, Factors are Parameterized Statistical Games, Backpropagation by the Implicit Function Theorem