For a (strict) symmetric Monoidal Category — or more generally an Actegory — the parametrised category has
- objects: the objects of ;
- morphisms : pairs of a parameter object and a map ;
- identity on : ;
- composition of and : the pair
Parameter spaces multiply under composition: composing a layer with 20 weights and a layer with 30 weights gives one arrow with 50 weights, and the composite’s parameter wire is the pair of the two. A reparametrisation of along is — “feed the parameter wire through first”. Reparametrisations are the 2-cells that make a Bicategory; quotienting them out (Cruttwell et al.) gives a category.
Sources: Cruttwell, Gavranović, Ghani, Wilson & Zanasi, Categorical Foundations of Gradient-Based Learning arXiv:2103.01931 (notes) Definitions 2.1, 2.3, Example 2.2, Remark 2.1; Capucci, Gavranović, Hedges & Fjeldgren Rischel arXiv:2105.06332 (notes) Definition 2, Remarks 3–4, Proposition 5; Fong, Spivak & Tuyéras, Backprop as Functor arXiv:1711.10455 (notes) (the category of parametrised functions); Gavranović, Fundamental Components of Deep Learning arXiv:2403.13001 (notes) ch. 3.
Example: neural networks (Cruttwell et al., Example 2.2)
Let have natural numbers as objects and smooth maps as morphisms, with on objects. Then
is a morphism of (15 weights plus 5 biases), and a two-layer chain has parameter object — the second layer’s parameters first, exactly as in the composition rule. Deep-learning libraries store this as a nested record, e.g. Lux.jl’s (layer_1 = …, layer_2 = …), which is an order-insensitive representative of the associativity class.
Reparametrisations are everywhere
| reparametrisation | what it does |
|---|---|
| weight tying (two layers share one parameter) | |
| a hypernetwork | parameters generated by another network |
| LoRA: a low-rank reparametrisation of a frozen weight | |
| a quantisation map | low-precision weights |
an optimiser’s get | optimisers — see Gradient-Based Learning with Parametric Lenses |
Para is a monad, and CoPara is its dual
Capucci et al. (Proposition 5) show is a monad on -actegories: a map parametrised twice, by and then by , is a map parametrised once by . The dual (Remark 4) has morphisms — the extra object sits in the codomain. Probabilistic models with latent variables are coparametrised: an open model in the AutoBayes framework is a coparametrised Markov kernel whose “coparameter” is the latent space, composed by copy-composition (St Clere Smithe’s , thesis arXiv:2212.12538 (notes) §5.2.1).
Where it goes next
- is the category of parametric lenses — the shape of a trainable layer with a backward pass.
- , for a reverse derivative , is automatic differentiation.
- Parameterized statistical games (AutoBayes Definition 27) are applied to a category of Bayesian lenses with losses.
- Open games are equipped with a selection (best-response) functor.
Docs: plain Julia — Catlab has no dedicated API for this; related: Catlab v0.16 docs · GATlab standard library
# Para(Smooth) on dense layers: a morphism is (parameter shape, f(p, x)).
struct ParaMap{F}; nparams::Int; f::F; end
(φ::ParaMap)(p, x) = φ.f(p, x)
function dense(n, m, σ = identity)
ParaMap(m * n + m, (p, x) -> σ.(reshape(p[1:m*n], m, n) * x .+ p[m*n+1:end]))
end
# composition: parameters concatenate (second layer first, as in Q ⊗ P)
compose(g::ParaMap, f::ParaMap) =
ParaMap(g.nparams + f.nparams, (p, x) -> g(p[1:g.nparams], f(p[g.nparams+1:end], x)))
net = compose(dense(5, 2), dense(3, 5, tanh))
net.nparams # 12 + 20 = 32
p = randn(net.nparams); x = randn(3)
size(net(p, x)) # (2,)
# reparametrisation = weight tying: one 2-vector drives both halves of a 4-parameter map
twoscales = ParaMap(2, (p, x) -> p[1] * x + p[2])
tied = ParaMap(1, (q, x) -> twoscales([q[1], q[1]], x))
tied([3.0], 2.0) == 9.0 # trueimport Mathlib
open CategoryTheory MonoidalCategory
-- a morphism of Para(C): a parameter object and a map P ⊗ A ⟶ B
structure ParaHom {C : Type*} [Category C] [MonoidalCategory C] (A B : C) where
P : C
f : P ⊗ A ⟶ B
def ParaHom.comp {C : Type*} [Category C] [MonoidalCategory C] {A B D : C}
(φ : ParaHom A B) (ψ : ParaHom B D) : ParaHom A D where
P := ψ.P ⊗ φ.P
f := (α_ ψ.P φ.P A).hom ≫ (ψ.P ◁ φ.f) ≫ ψ.f-- Para over (Hask, (,)): parameters multiply under composition
newtype Para p a b = Para { run :: p -> a -> b }
(>>>) :: Para p a b -> Para q b c -> Para (q, p) a c
Para f >>> Para g = Para (\(q, p) a -> g q (f p a))
reparam :: (q -> p) -> Para p a b -> Para q a b -- a 2-cell of Para
reparam alpha (Para f) = Para (f . alpha)
scale :: Para Double Double Double
scale = Para (*)
tied :: Para Double Double Double -- weight tying via the diagonal
tied = reparam (\w -> (w, w)) (scale >>> scale)
-- run tied 3 2 == 18