open-problem

Learn a relation as a formula: not a polynomial with fixed monomials, not a network, but a short symbolic expression found by search. It would make learned relations readable and comparable, and it is slow and hard. Noted for later; nothing is implemented.

Sources: Schmidt & Lipson, Distilling free-form natural laws from experimental data, Science 2009; Mangan, Brunton, Proctor & Kutz, Inferring biological networks by sparse identification of nonlinear dynamics, IEEE TMBMC 2016 (implicit SINDy); Kaheman, Kutz & Brunton, SINDy-PI, Proc. R. Soc. A 2020; Cranmer, Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl, arXiv:2305.01582, 2023; Cranmer et al., Discovering Symbolic Models from Deep Learning with Inductive Biases, NeurIPS 2020

Theory (CT-ML wiki): Statistical Game

1. The idea

Symbolic regression searches over expression trees for . The implicit version searches for with on the data, so that the formula is a relation with no privileged output, the way this family’s polynomials are (Algebraic Implicit Learners). The algebraic family is the special case of a fixed library of monomials, where fitting is a nullspace problem (Fitting is a Nullspace Problem); implicit SINDy is the same with any fixed library of candidate terms. Symbolic implicit learning drops the fixed library.

2. Why it would be worth it

  • Readable relations. The output is a law, , not a weight vector.
  • Model selection is native: a Pareto front of complexity against residual is a family of candidate relations, and the free energy is a principled way to choose among them (Open Problems in Algebraic Implicit Learning §4 is the same jump-between-models problem).
  • Distillation. A learned factor (a pair potential, an energy network) can be explained afterwards by a formula, as in Cranmer et al. 2020. The force-law and planetary examples discussed for the tutorials need exactly this last step.

3. Why it is hard

  • Trivial solutions. fits every dataset, and so does for any that vanishes on it. Explicit regression avoids this by fixing ; implicit regression needs a normalisation. The nullspace formulation normalises the coefficient vector; Schmidt and Lipson match ratios of partial derivatives instead; SINDy-PI tries each library term as the left-hand side in turn.
  • Search cost. Genetic search over expression trees is slow, and the implicit version has no target column to guide it.
  • Real locus. A formula can vanish on the data and on extra spurious components, the same problem as Open Problems in Algebraic Implicit Learning §1.

4. What exists to build on

  • SymbolicRegression.jl (and PySR on top of it): evolutionary search, Pareto fronts, custom losses, dimensional constraints, templates that fix part of a formula. Explicit regression; an implicit loss is possible through its custom-loss interface.
  • DataDrivenDiffEq.jl (SciML): SINDy and an implicit optimiser for fixed libraries.
  • This family’s own nullspace fitting, which is implicit and exact for polynomial libraries.

A first experiment: recover the circle, then the robot arm’s kinematic relation, from samples, with an implicit loss in SymbolicRegression.jl, and compare with the polynomial nullspace fit.

Related: Algebraic Implicit Learners, Fitting is a Nullspace Problem, Open Problems in Algebraic Implicit Learning, Geometric Deep Learning and Physical Laws