Derivation. The convexity of Fitting is a Nullspace Problem is bought by using the wrong distance. Correcting it produces, exactly, a scalarisation — and shows what the energy space is: the space the noise lives in.
Sources: original to this vault (design and analysis; no single paper).
Theory (CT-ML wiki): Lax Functor · Statistical Game
The two distances
Algebraic distance. . Cheap, quadratic in , hence the closed-form fit. But it is not a distance: it depends on how the variety is written. Scaling scales it; a point far from the variety in a region where is large scores better than a near point where is small.
Geometric distance. . The thing you actually mean. Independent of the equations. Also very expensive: the number of complex critical points of on is the Euclidean Distance Degree (Draisma–Horobeț–Ottaviani–Sturmfels–Thomas). For a generic hypersurface of degree in ,
Sanity check: a plane conic, , : — the classical fact that the foot of the perpendicular to a conic solves a quartic. For a general hypersurface the EDD is exponential in , so evaluating the exact geometric distance is itself an algebraic inference problem of the same difficulty as the one we are trying to solve.
This is the fundamental tension of the family, and it is not resolvable: exact fitting requires exact projection, which is as hard as inference.
The first-order correction: Sampson distance
Linearise. Near , the variety is approximately the affine subspace with . The minimum-norm displacement to that subspace solves subject to , whose solution is the pseudoinverse , giving
the Sampson distance (Sampson 1982; Taubin 1991). For a single equation () it reduces to the familiar .
The derivation that matters: is where the noise lives
Now read the same formula statistically rather than geometrically. Posit the honest errors-in-variables model: the true point lies on , and we observe with in the data space.
Push the noise forward through the residual. To first order,
So the negative log-likelihood of the observation, expressed in residual coordinates, is
which for is exactly the Sampson distance. Two conclusions, and they are the payoff of this note:
What the energy space is
The residual space of Scalar and Multivariate Energy is the space the observation noise lives in after pushforward, and the scalarisation is the noise precision. Choosing is choosing a noise model, not choosing a convenience.
Why the energy must stay multivariate
cannot be computed from the scalar energy . It needs and its Jacobian — i.e. the vector energy. A factor that scalarises early has irrecoverably thrown away its own noise model. This is the concrete instance of Scalar and Multivariate Energy §6’s claim that “scalarising early destroys the metric”.
The cost: three properties lost
depends on , and therefore on . That breaks three things that made the algebraic distance attractive:
- Convexity in is gone. is no longer quadratic, so Fitting is a Nullspace Problem’s global optimum does not apply. The standard workaround is iteratively reweighted least squares: freeze , solve the eigenproblem, recompute , repeat. Each step is globally optimal; the alternation is not. (This is Taubin’s algorithm, and in the conic-fitting literature it is known to work well and to have no convergence proof.)
- is not linear, so by Scalar and Multivariate Energy §5 the composition of games is lax: the scalar chain rule and the multivariate one differ by the Jensen gap . The algebraic family is therefore a lax member of the framework, by construction, and the laxness is quantified.
- blows up on the discriminant, where loses rank
(Branches and the Discriminant). Correctly — the noise really is unconstrained along
the collapsing direction — but uselessly. The
SquaredNorm(M)caveat recorded inenergy.md§4 (that must be PSD and is not validated) is exactly this failure, met in the wild.
The honest summary
| distance | convex in ? | statistically correct? | cost |
|---|---|---|---|
| algebraic | yes, closed form | no — biased by | once |
| Sampson | no (IRLS) | first order | per IRLS step |
| geometric | no | yes (MLE, isotropic noise) | EDD-many roots per data point |
The practical recommendation is Sampson via IRLS, initialised at the algebraic solution. The theoretically interesting statement is that the three rows are three scalarisations of the same multivariate energy, and moving between them is a change of , not a change of model. That is precisely the reweighting flexibility Scalar and Multivariate Energy §6.4 claims the vector energy buys.
Related: Fitting is a Nullspace Problem, Scalar and Multivariate Energy, Branches and the Discriminant, The Algebraic Factor as a Statistical Game