definition example theorem

For a (strict) symmetric Monoidal Category — or more generally an Actegory — the parametrised category has

  • objects: the objects of ;
  • morphisms : pairs of a parameter object and a map ;
  • identity on : ;
  • composition of and : the pair

Parameter spaces multiply under composition: composing a layer with 20 weights and a layer with 30 weights gives one arrow with 50 weights, and the composite’s parameter wire is the pair of the two. A reparametrisation of along is — “feed the parameter wire through first”. Reparametrisations are the 2-cells that make a Bicategory; quotienting them out (Cruttwell et al.) gives a category.

Sources: Cruttwell, Gavranović, Ghani, Wilson & Zanasi, Categorical Foundations of Gradient-Based Learning arXiv:2103.01931 (notes) Definitions 2.1, 2.3, Example 2.2, Remark 2.1; Capucci, Gavranović, Hedges & Fjeldgren Rischel arXiv:2105.06332 (notes) Definition 2, Remarks 3–4, Proposition 5; Fong, Spivak & Tuyéras, Backprop as Functor arXiv:1711.10455 (notes) (the category of parametrised functions); Gavranović, Fundamental Components of Deep Learning arXiv:2403.13001 (notes) ch. 3.

Q¬AP¬AB®¬1A(Q;(®¬1);f)fQ¬AP¬AB®¬1A(Q;(®¬1);f)f

Example: neural networks (Cruttwell et al., Example 2.2)

Let have natural numbers as objects and smooth maps as morphisms, with on objects. Then

is a morphism of (15 weights plus 5 biases), and a two-layer chain has parameter object — the second layer’s parameters first, exactly as in the composition rule. Deep-learning libraries store this as a nested record, e.g. Lux.jl’s (layer_1 = …, layer_2 = …), which is an order-insensitive representative of the associativity class.

Reparametrisations are everywhere

reparametrisation what it does
weight tying (two layers share one parameter)
a hypernetwork parameters generated by another network
LoRA: a low-rank reparametrisation of a frozen weight
a quantisation maplow-precision weights
an optimiser’s getoptimisers — see Gradient-Based Learning with Parametric Lenses

Para is a monad, and CoPara is its dual

Capucci et al. (Proposition 5) show is a monad on -actegories: a map parametrised twice, by and then by , is a map parametrised once by . The dual (Remark 4) has morphisms — the extra object sits in the codomain. Probabilistic models with latent variables are coparametrised: an open model in the AutoBayes framework is a coparametrised Markov kernel whose “coparameter” is the latent space, composed by copy-composition (St Clere Smithe’s , thesis arXiv:2212.12538 (notes) §5.2.1).

Where it goes next

  • is the category of parametric lenses — the shape of a trainable layer with a backward pass.
  • , for a reverse derivative , is automatic differentiation.
  • Parameterized statistical games (AutoBayes Definition 27) are applied to a category of Bayesian lenses with losses.
  • Open games are equipped with a selection (best-response) functor.

Docs: plain Julia — Catlab has no dedicated API for this; related: Catlab v0.16 docs · GATlab standard library

# Para(Smooth) on dense layers: a morphism is (parameter shape, f(p, x)).
struct ParaMap{F}; nparams::Int; f::F; end
(φ::ParaMap)(p, x) = φ.f(p, x)
function dense(n, m, σ = identity)
    ParaMap(m * n + m, (p, x) -> σ.(reshape(p[1:m*n], m, n) * x .+ p[m*n+1:end]))
end
# composition: parameters concatenate (second layer first, as in Q ⊗ P)
compose(g::ParaMap, f::ParaMap) =
    ParaMap(g.nparams + f.nparams, (p, x) -> g(p[1:g.nparams], f(p[g.nparams+1:end], x)))
net = compose(dense(5, 2), dense(3, 5, tanh))
net.nparams                                            # 12 + 20 = 32
p = randn(net.nparams); x = randn(3)
size(net(p, x))                                        # (2,)
# reparametrisation = weight tying: one 2-vector drives both halves of a 4-parameter map
twoscales = ParaMap(2, (p, x) -> p[1] * x + p[2])
tied = ParaMap(1, (q, x) -> twoscales([q[1], q[1]], x))
tied([3.0], 2.0) == 9.0                                # true
import Mathlib
open CategoryTheory MonoidalCategory
-- a morphism of Para(C): a parameter object and a map P ⊗ A ⟶ B
structure ParaHom {C : Type*} [Category C] [MonoidalCategory C] (A B : C) where
  P : C
  f : P ⊗ A ⟶ B
 
def ParaHom.comp {C : Type*} [Category C] [MonoidalCategory C] {A B D : C}
    (φ : ParaHom A B) (ψ : ParaHom B D) : ParaHom A D where
  P := ψ.P ⊗ φ.P
  f := (α_ ψ.P φ.P A).hom ≫ (ψ.P ◁ φ.f) ≫ ψ.f
-- Para over (Hask, (,)): parameters multiply under composition
newtype Para p a b = Para { run :: p -> a -> b }
 
(>>>) :: Para p a b -> Para q b c -> Para (q, p) a c
Para f >>> Para g = Para (\(q, p) a -> g q (f p a))
 
reparam :: (q -> p) -> Para p a b -> Para q a b          -- a 2-cell of Para
reparam alpha (Para f) = Para (f . alpha)
 
scale :: Para Double Double Double
scale = Para (*)
tied :: Para Double Double Double                           -- weight tying via the diagonal
tied = reparam (\w -> (w, w)) (scale >>> scale)
-- run tied 3 2 == 18