model derivation

Where differential algebra genuinely belongs — and it is not backpropagation. README’s table pairs “implicit function theorem / differential algebra” as the backward pass. The IFT does the backward pass (Backpropagation by the Implicit Function Theorem); differential algebra does a different and arguably more valuable job.

Sources: original to this vault (design and analysis; no single paper).

The differential setting

A differential ring is with a derivation. The differential polynomial ring is

— polynomials in the variables and all their derivatives. A differential ideal is closed under .

The basic finiteness result is weaker than Hilbert’s. Ritt–Raudenbush: every radical differential ideal is the radical differential ideal generated by finitely many elements. Non-radical differential ideals need not be finitely generated at all — the first sign that this setting is harder.

The computational engine is the Rosenfeld–Gröbner algorithm (Boulier–Lazard–Ollivier–Petitot), which decomposes a radical differential ideal into regular differential chains; it is what Maple’s DifferentialAlgebra and BLAD implement.

When a factor is a DAE

A Lenticulum factor whose relation involves derivatives,

is a differential-algebraic equation. This is not exotic: it is the natural home of

  • mechanical systems in descriptor form (Lagrangian dynamics with constraints — the constraint forces make it a DAE, not an ODE);
  • circuit models (modified nodal analysis is index-1 or index-2 DAE);
  • NeuralODEs with algebraic constraints, and any equilibrium model with a conservation law;
  • chemical kinetics with quasi-steady-state assumptions.

It sits between the algebraic and equilibrium families of Implicit Learners and inherits difficulties from both.

The differentiation index

The central invariant. The index is the number of times the system must be differentiated to extract an explicit ODE for all components.

  • : an ODE.
  • : is invertible after one differentiation. Standard solvers (IDA/Sundials, DAE mode in DifferentialEquations.jl) handle these.
  • : there are hidden constraints — consequences of the equations that are not visible without differentiating. Direct numerical integration is unstable: the constraint drift grows, and consistent initial conditions cannot be found by inspection.

Differential elimination is exactly the computation of the hidden constraints. Index reduction — the standard preprocessing that makes a high-index DAE solvable — is a differential-elimination step, whether performed by Pantelides’ graph-theoretic algorithm or by Rosenfeld–Gröbner symbolically. So:

Differential algebra’s first job is not differentiation; it is making the inference problem well posed in the first place.

The adjoint of a DAE — where it does touch backprop

Backpropagating through a DAE solve is the adjoint DAE (Cao–Li–Petzold–Serban). The key fact for us:

The adjoint of an index- DAE is itself a DAE of index (and for index , the adjoint has its own hidden constraints).

So for you must run index reduction twice — once on the forward system and once on the adjoint — and the two reductions must be consistent or the computed gradient is wrong in a way that does not show up as a solver failure. That is the one place where differential algebra is genuinely part of the backward pass, and it vindicates the README’s pairing, but narrowly and for a specific reason.

For index-1 DAEs the adjoint is index-1 and everything reduces to the ordinary implicit function theorem applied at each step — nothing new.

The second job: structural identifiability = polarity legality

This is the payoff, and it is a genuine reframing of something the vault already has.

Given a model , , the question

can be recovered from perfect, noise-free observation of ?

is structural identifiability, and the classical method (Ljung–Glad) answers it by differential elimination: eliminate the unobserved states from the system to obtain the input–output equations, polynomial relations among , its derivatives, and alone. The parameter is identifiable iff the coefficients of those input–output equations determine .

Now compare Channels and Polarity:

supports_polarity(factor, p) — whether the factor can answer this input/output split.

These are the same question. Structural identifiability is supports_polarity for the differential case, and — unlike the general case, where the vault says the predicate must be declared by the factor author — it is decidable, by differential elimination.

That is a satisfying result: in the algebraic and differential-algebraic families, polarity legality is not a matter of the author’s judgement but a computation. In Julia the relevant tool is StructuralIdentifiability.jl, which does exactly this (differential elimination / power-series methods) for ODE models.

The costs, which are severe

Everything in Composition is Elimination gets worse:

  • Effective differential Nullstellensatz bounds are enormous — multiply exponential in the number of variables and the order of the system. The known bounds (Golubitsky–Kondratieva–Ovchinnikov–Szanto; Ovchinnikov–Pogudin–Vo) are far from anything practical, and it is not known whether they are tight.
  • Rosenfeld–Gröbner is practical only for small systems — a handful of states and parameters. Modern identifiability tools sidestep the worst of it with randomised / power-series methods rather than full symbolic elimination, which is a strong hint that the symbolic route is not the one to build on.
  • Numerical instability carries over: differential Gröbner-type computations inherit the leading-term discontinuity of the commutative case, so they cannot be applied to learned coefficients. There is no differential analogue of border bases in common use.

The practical verdict

Differential algebra is a compile-time, exact-coefficient, small-system tool. Use it to answer structural questions about a factor’s design (index, hidden constraints, identifiability, legal polarities), once, in exact arithmetic, before any data is seen. Do not put it in a training loop, and do not apply it to fitted floating-point coefficients.

This note is the prerequisite for the temporal extension

Time as a Base designs a Lenticulum in which variables carry trajectories rather than values. Everything above — index, hidden constraints, consistent initialisation — has to happen before any belief is propagated over such a graph, and all of it is symbolic. That is why Time as a Base §8 concludes with a division of labour rather than a merged library.

Related: Composition is Elimination, Channels and Polarity, Implicit Learners, Backpropagation by the Implicit Function Theorem, Time as a Base, The Structural Gap to ModelingToolkit