References
The works cited in the tutorials, the implementation notes and the theory vault. The same bibliography, docs/src/refs.bib, is rendered as the vault's Bibliography note.
- — (2001). Sequential Monte Carlo Methods in Practice (Springer).
- — (2026). Catlab.jl: A framework for applied category theory in Julia. Software. Accessed October 2026.
- — (2026). DifferentiationInterface.jl: An interface to various automatic differentiation backends in Julia. Software. Accessed October 2026.
- — (2026). ForneyLab.jl: Generating fast and flexible message passing algorithms. Software. Accessed October 2026.
- — (2026). IncrementalInference.jl: Non-parametric factor graph inference, part of Caesar.jl. Software. Accessed October 2026.
- — (2026). Lux.jl: Explicitly parameterized neural networks in Julia. Software. Accessed October 2026.
- — (2026). Reactant.jl: Optimize Julia functions with MLIR and XLA. Software. Accessed October 2026.
- Anderson, S.; Barfoot, T. D.; Tong, C. H. and Särkkä, S. (2015). Batch Nonlinear Continuous-Time Trajectory Estimation as Exactly Sparse Gaussian Process Regression. Autonomous Robots 39, 221–238, arXiv:1412.0630.
- Ansel, J.; Yang, E.; He, H.; Gimelshein, N.; Jain, A.; Voznesensky, M.; Bao, B.; Bell, P.; Berard, D.; Burovski, E.; Chauhan, G.; Chourdia, A.; Constable, W.; Desmaison, A.; DeVito, Z.; Ellison, E.; Feng, W.; Gong, J.; Gschwind, M.; Hirsh, B.; Huang, S.; Kalambarkar, K.; Kirsch, L.; Lazos, M.; Lezcano, M.; Liang, Y.; Liang, J.; Lu, Y.; Luk, C. K.; Maher, B.; Pan, Y.; Puhrsch, C.; Reso, M.; Saroufim, M.; Siraichi, M. Y.; Suk, H.; Zhang, S.; Suo, M.; Tillet, P.; Zhao, X.; Wang, E.; Zhou, K.; Zou, R.; Wang, X.; Mathews, A.; Wen, W.; Chanan, G.; Wu, P. and Chintala, S. (2024). PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. In: Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2; pp. 929–947.
- Atkey, R. (2018). Syntax and Semantics of Quantitative Type Theory. In: Proceedings of the 33rd Annual ACM/IEEE Symposium on Logic in Computer Science; pp. 56–65.
- Baez, J. C. and Courser, K. (2020). Structured Cospans. Theory and Applications of Categories 35, arXiv:1911.04630.
- Bagaev, D.; Podusenko, A. and de Vries, B. (2023). RxInfer: A Julia package for reactive real-time Bayesian inference. Journal of Open Source Software 8, 5161.
- Bai, S.; Kolter, J. Z. and Koltun, V. (2019). Deep Equilibrium Models. In: Advances in Neural Information Processing Systems (NeurIPS), arXiv:1909.01377. ↩1
- Batatia, I.; Kovacs, D. P.; Simm, G.; Ortner, C. and Csanyi, G. (2022). MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields. In: Advances in Neural Information Processing Systems 35; pp. 11423–11436.
- Batzner, S.; Musaelian, A.; Sun, L.; Geiger, M.; Mailoa, J. P.; Kornbluth, M.; Molinari, N.; Smidt, T. E. and Kozinsky, B. (2022). E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials. Nature Communications 13.
- Bezanson, J.; Chen, J.; Chung, B.; Karpinski, S.; Shah, V. B.; Vitek, J. and Zoubritzky, L. (2018). Julia: dynamism and performance reconciled by design. Proceedings of the ACM on Programming Languages 2, 1–23.
- Bezanson, J.; Edelman, A.; Karpinski, S. and Shah, V. B. (2017). Julia: A Fresh Approach to Numerical Computing. SIAM Review 59, 65–98, arXiv:1411.1607.
- Bigi, F.; Langer, M. and Ceriotti, M. (2025). The dark side of the forces: assessing non-conservative force models for atomistic machine learning. In: International Conference on Machine Learning (ICML), arXiv:2412.11569. ↩1
- Bishop, C. M. (1994). Mixture Density Networks. Technical Report NCRG/94/004 (Aston University). ↩1
- Bistarelli, S.; Montanari, U.; Rossi, F.; Schiex, T.; Verfaillie, G. and Fargier, H. (1999). Semiring-Based CSPs and Valued CSPs: Frameworks, Properties, and Comparison. Constraints 4, 199–240.
- Bistarelli, S.; Montanari, U. and Rossi, F. (1997). Semiring-based constraint satisfaction and optimization. Journal of the ACM 44, 201–236.
- Bonchi, F.; Sobociński, P. and Zanasi, F. (2017). Interacting Hopf algebras. Journal of Pure and Applied Algebra 221, 144–184.
- Bradbury, J.; Frostig, R.; Hawkins, P.; Johnson, M. J.; Leary, C.; Maclaurin, D.; Necula, G.; Paszke, A.; VanderPlas, J.; Wanderman-Milne, S. and Zhang, Q. (2018). JAX: composable transformations of Python+NumPy programs. Software.
- Bronstein, M. M.; Bruna, J.; Cohen, T. and Veličković, P. (2021). Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges, arXiv:2104.13478. ↩1
- Broyden, C. G. (1965). A class of methods for solving nonlinear simultaneous equations. Mathematics of Computation 19, 577–593.
- Carr, J. C.; Beatson, R. K.; Cherrie, J. B.; Mitchell, T. J.; Fright, W. R.; McCallum, B. C. and Evans, T. R. (2001). Reconstruction and representation of 3D objects with radial basis functions. In: Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH).
- Chen, R. T.; Rubanova, Y.; Bettencourt, J. and Duvenaud, D. (2018). Neural Ordinary Differential Equations. In: Advances in Neural Information Processing Systems (NeurIPS), arXiv:1806.07366.
- Chung, H.; Kim, J.; McCann, M. T.; Klasky, M. L. and Ye, J. C. (2023). Diffusion Posterior Sampling for General Noisy Inverse Problems. In: International Conference on Learning Representations (ICLR), arXiv:2209.14687.
- Cooper, R.; Dobnik, S.; Lappin, S. and Larsson, S. (2015). Probabilistic Type Theory and Natural Language Semantics. Linguistic Issues in Language Technology 10.
- Cranmer, M. (2023). Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl, arXiv:2305.01582. ↩1
- Cranmer, M.; Greydanus, S.; Hoyer, S.; Battaglia, P.; Spergel, D. and Ho, S. (2020). Lagrangian Neural Networks, arXiv:2003.04630.
- Cranmer, M.; Sanchez-Gonzalez, A.; Battaglia, P.; Xu, R.; Cranmer, K.; Spergel, D. and Ho, S. (2020). Discovering Symbolic Models from Deep Learning with Inductive Biases. In: Advances in Neural Information Processing Systems (NeurIPS), arXiv:2006.11287. ↩1
- Cruttwell, G. S.; Gavranović, B.; Ghani, N.; Wilson, P. and Zanasi, F. (2022). Categorical Foundations of Gradient-Based Learning. In: European Symposium on Programming (ESOP), arXiv:2103.01931.
- Cusumano-Towner, M. F.; Saad, F. A.; Lew, A. K. and Mansinghka, V. K. (2019). Gen: A General-Purpose Probabilistic Programming System with Programmable Inference. In: ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI).
- De Raedt, L.; Kimmig, A. and Toivonen, H. (2007). ProbLog: A Probabilistic Prolog and Its Application in Link Discovery. In: International Joint Conference on Artificial Intelligence (IJCAI).
- Dechter, R. (1999). Bucket elimination: A unifying framework for reasoning. Artificial Intelligence 113, 41–85.
- Dellaert, F. and Kaess, M. (2017). Factor Graphs for Robot Perception (Now Publishers). ↩1
- Dong, J.; Mukadam, M.; Dellaert, F. and Boots, B. (2016). Motion Planning as Probabilistic Inference using Gaussian Processes and Factor Graphs. In: Robotics: Science and Systems XII.
- Du, Y.; Durkan, C.; Strudel, R.; Tenenbaum, J. B.; Dieleman, S.; Fergus, R.; Sohl-Dickstein, J.; Doucet, A. and Grathwohl, W. (2023). Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMC. In: International Conference on Machine Learning (ICML), arXiv:2302.11552. ↩1
- Du, Y. and Mordatch, I. (2019). Implicit Generation and Generalization in Energy-Based Models. In: Advances in Neural Information Processing Systems (NeurIPS), arXiv:1903.08689.
- Efron, B. (2011). Tweedie's Formula and Selection Bias. Journal of the American Statistical Association 106, 1602–1614.
- Fang, Z.; Díaz, M.; Buchanan, S. and Sulam, J. (2025). Beyond Scores: Proximal Diffusion Models, arXiv:2507.08956. ↩1
- Feldbaum, A. A. (1960). Dual control theory I–IV. Automation and Remote Control.
- Fong, B. (2013). Causal Theories: A Categorical Perspective on Bayesian Networks. Master's thesis, University of Oxford, arXiv:1301.6201.
- Fong, B. (2016). The Algebra of Open and Interconnected Systems. Ph.D. Thesis, University of Oxford, arXiv:1609.05382.
- Fong, B. and Spivak, D. I. (2019). Hypergraph Categories. Journal of Pure and Applied Algebra, arXiv:1806.08304.
- Friston, K. (2010). The free-energy principle: a unified brain theory? Nature Reviews Neuroscience 11, 127–138.
- Friston, K.; FitzGerald, T.; Rigoli, F.; Schwartenbeck, P. and Pezzulo, G. (2017). Active Inference: A Process Theory. Neural Computation 29, 1–49.
- Fritz, T. (2020). A synthetic approach to Markov kernels, conditional independence and theorems on sufficient statistics. Advances in Mathematics 370, arXiv:1908.07021.
- Frostig, R.; Johnson, M. J. and Leary, C. (2018). Compiling machine learning programs via high-level tracing. In: Systems for Machine Learning (SysML).
- Ge, H.; Xu, K. and Ghahramani, Z. (2018). Turing: A Language for Flexible Probabilistic Inference. In: International Conference on Artificial Intelligence and Statistics (AISTATS).
- Genovese, C. R.; Perone-Pacifico, M.; Verdinelli, I. and Wasserman, L. (2014). Nonparametric ridge estimation. The Annals of Statistics 42.
- Ghani, N.; Hedges, J.; Winschel, V. and Zahn, P. (2018). Compositional game theory. In: Logic in Computer Science (LICS), arXiv:1603.04641.
- Giry, M. (1982). A categorical approach to probability theory. In: Lecture Notes in Mathematics; pp. 68–85.
- Grathwohl, W.; Chen, R. T.; Bettencourt, J.; Sutskever, I. and Duvenaud, D. (2019). FFJORD: Free-form Continuous Dynamics for Scalable Reversible Generative Models. In: International Conference on Learning Representations (ICLR), arXiv:1810.01367.
- Greydanus, S.; Dzamba, M. and Yosinski, J. (2019). Hamiltonian Neural Networks. In: Advances in Neural Information Processing Systems (NeurIPS), arXiv:1906.01563.
- Hinton, G. E. (2002). Training Products of Experts by Minimizing Contrastive Divergence. Neural Computation 14, 1771–1800.
- Ho, J.; Jain, A. and Abbeel, P. (2020). Denoising Diffusion Probabilistic Models. In: Advances in Neural Information Processing Systems (NeurIPS), arXiv:2006.11239. ↩1 ↩2
- Hoffmann, H. (2007). Kernel PCA for novelty detection. Pattern Recognition 40, 863–874.
- Innes, M. (2018). Don't Unroll Adjoint: Differentiating SSA-Form Programs, arXiv:1810.07951.
- Innes, M.; Saba, E.; Fischer, K.; Gandhi, D.; Rudilosso, M. C.; Joy, N. M.; Karmali, T.; Pal, A. and Shah, V. (2018). Fashionable Modelling with Flux, arXiv:1811.01457.
- Kaess, M.; Johannsson, H.; Roberts, R.; Ila, V.; Leonard, J. J. and Dellaert, F. (2012), iSAM2: Incremental smoothing and mapping using the Bayes tree. The International Journal of Robotics Research 31, 216–235.
- Kaheman, K.; Kutz, J. N. and Brunton, S. L. (2020). SINDy-PI: A Robust Algorithm for Parallel Implicit Sparse Identification of Nonlinear Dynamics. Proceedings of the Royal Society A, arXiv:2004.02322.
- Kalman, R. E. (1960). A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering 82, 35–45. ↩1
- Katsumata, S.-y. (2014). Parametric effect monads and semantics of effect systems. ACM SIGPLAN Notices 49, 633–645.
- Koller, D. and Friedman, N. (2009). Probabilistic Graphical Models: Principles and Techniques (MIT Press).
- Kolter, Z.; Duvenaud, D. and Johnson, M. (2020). Deep Implicit Layers: Neural ODEs, Deep Equilibrium Models, and Beyond. NeurIPS 2020 tutorial.
- Lavore, E. D.; Román, M. and Sobociński, P. (2025). Partial Markov Categories, arXiv:2502.03477.
- LeCun, Y.; Chopra, S.; Hadsell, R.; Ranzato, M. and Huang, F. J. (2006). A Tutorial on Energy-Based Learning. In: Predicting Structured Data (MIT Press).
- Li, X. L.; Holtzman, A.; Fried, D.; Liang, P.; Eisner, J.; Hashimoto, T.; Zettlemoyer, L. and Lewis, M. (2023). Contrastive Decoding: Open-ended Text Generation as Optimization. In: Annual Meeting of the Association for Computational Linguistics (ACL), arXiv:2210.15097.
- Liu, N.; Li, S.; Du, Y.; Torralba, A. and Tenenbaum, J. B. (2022). Compositional Visual Generation with Composable Diffusion Models. In: European Conference on Computer Vision (ECCV), arXiv:2206.01714.
- Livni, R.; Lehavi, D.; Schein, S.; Nachlieli, H.; Shalev-Shwartz, S. and Globerson, A. (2013). Vanishing Component Analysis. In: International Conference on Machine Learning (ICML).
- Loeliger, H.-A. (2004). An Introduction to factor graphs. IEEE Signal Processing Magazine 21, 28–41.
- Loeliger, H.-A.; Dauwels, J.; Hu, J.; Korl, S.; Ping, L. and Kschischang, F. R. (2007). The Factor Graph Approach to Model-Based Signal Processing. Proceedings of the IEEE 95, 1295–1322.
- Lou, A.; Meng, C. and Ermon, S. (2024). Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution. In: International Conference on Machine Learning (ICML), arXiv:2310.16834.
- Ma, Y.; Gowda, S.; Anantharaman, R.; Laughman, C.; Shah, V. and Rackauckas, C. (2021). ModelingToolkit: A Composable Graph Transformation System for Equation-Based Modeling, arXiv:2103.05244.
- Macêdo, I.; Gois, J. P. and Velho, L. (2011). Hermite Radial Basis Functions Implicits. Computer Graphics Forum 30, 27–42.
- Mangan, N. M.; Brunton, S. L.; Proctor, J. L. and Kutz, J. N. (2016). Inferring biological networks by sparse identification of nonlinear dynamics. IEEE Transactions on Molecular, Biological and Multi-Scale Communications, arXiv:1605.08368.
- Mardani, M.; Song, J.; Kautz, J. and Vahdat, A. (2024). A Variational Perspective on Solving Inverse Problems with Diffusion Models. In: International Conference on Learning Representations (ICLR), arXiv:2305.04391. ↩1
- Minka, T. P. (2001). Expectation Propagation for approximate Bayesian inference. In: Conference on Uncertainty in Artificial Intelligence (UAI), arXiv:1301.2294.
- Mohamed, S. and Lakshminarayanan, B. (2016). Learning in Implicit Generative Models, arXiv:1610.03483.
- Moses, W. S. and Churavy, V. (2020). Instead of Rewriting Foreign Code for Machine Learning, Automatically Synthesize Fast Gradients. In: Advances in Neural Information Processing Systems (NeurIPS), arXiv:2010.01709.
- Nie, S.; Zhu, F.; You, Z.; Zhang, X.; Ou, J.; Hu, J.; Zhou, J.; Lin, Y.; Wen, J.-R. and Li, C. (2025). Large Language Diffusion Models, arXiv:2502.09992.
- Noether, E. (1918). Invariante Variationsprobleme. Nachrichten von der Gesellschaft der Wissenschaften zu Göttingen, Mathematisch-Physikalische Klasse.
- Ozertem, U. and Erdogmus, D. (2011). Locally Defined Principal Curves and Surfaces. Journal of Machine Learning Research 12.
- Pantelides, C. C. (1988). The Consistent Initialization of Differential-Algebraic Systems. SIAM Journal on Scientific and Statistical Computing 9, 213–231.
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Köpf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J. and Chintala, S. (2019). PyTorch: An Imperative Style, High-Performance Deep Learning Library. In: Advances in Neural Information Processing Systems (NeurIPS), arXiv:1912.01703.
- Poole, D. (2003). First-order probabilistic inference. In: International Joint Conference on Artificial Intelligence (IJCAI).
- Rackauckas, C.; Innes, M.; Ma, Y.; Bettencourt, J.; White, L. and Dixit, V. (2019). DiffEqFlux.jl – A Julia Library for Neural Differential Equations, arXiv:1902.02376.
- Rauch, H. E.; Tung, F. and Striebel, C. T. (1965). Maximum likelihood estimates of linear dynamic systems. AIAA Journal 3, 1445–1450. ↩1
- Revels, J.; Lubin, M. and Papamarkou, T. (2016). Forward-Mode Automatic Differentiation in Julia, arXiv:1607.07892.
- Richardson, M. and Domingos, P. (2006). Markov logic networks. Machine Learning 62, 107–136.
- Romano, Y.; Elad, M. and Milanfar, P. (2017). The Little Engine that Could: Regularization by Denoising (RED). SIAM Journal on Imaging Sciences 10, arXiv:1611.02862.
- Sahoo, S. S.; Arriola, M.; Schiff, Y.; Gokaslan, A.; Marroquin, E.; Chiu, J. T.; Rush, A. and Kuleshov, V. (2024). Simple and Effective Masked Diffusion Language Models. In: Advances in Neural Information Processing Systems (NeurIPS), arXiv:2406.07524.
- Salimans, T. and Ho, J. (2021). Should EBMs model the energy or the score? In: Energy Based Models Workshop, International Conference on Learning Representations (ICLR). ↩1
- Sato, T. and Kameya, Y. (1997). PRISM: A Language for Symbolic-Statistical Modeling. In: International Joint Conference on Artificial Intelligence (IJCAI).
- Schmidt, M. and Lipson, H. (2009). Distilling Free-Form Natural Laws from Experimental Data. Science 324, 81–85.
- Schölkopf, B.; Platt, J. C.; Shawe-Taylor, J.; Smola, A. J. and Williamson, R. C. (2001). Estimating the Support of a High-Dimensional Distribution. Neural Computation 13, 1443–1471.
- Somogyi, Z.; Henderson, F. and Conway, T. (1996). The execution algorithm of mercury, an efficient purely declarative logic programming language. The Journal of Logic Programming 29, 17–64.
- Song, Y. and Kingma, D. P. (2021). How to Train Your Energy-Based Models, arXiv:2101.03288.
- Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S. and Poole, B. (2021). Score-Based Generative Modeling through Stochastic Differential Equations. In: International Conference on Learning Representations (ICLR), arXiv:2011.13456. ↩1 ↩2
- Sriperumbudur, B.; Fukumizu, K.; Gretton, A.; Hyvärinen, A. and Kumar, R. (2017). Density Estimation in Infinite Dimensional Exponential Families. Journal of Machine Learning Research 18, arXiv:1312.3516.
- St Clere Smithe, T. (2020). Bayesian Updates Compose Optically, arXiv:2006.01631.
- St Clere Smithe, T. and Perin, M. (2025). AutoBayes: A Compositional Framework for Generalized Variational Inference, arXiv:2503.18608.
- Stein, D. and Samuelson, R. (2023). A Category for unifying Gaussian Probability and Nondeterminism. In: Conference on Algebra and Coalgebra in Computer Science (CALCO), arXiv:2204.14024.
- Stein, D.; Zanasi, F.; Piedeleu, R. and Samuelson, R. (2024). Graphical Quadratic Algebra, arXiv:2403.02284.
- Särkkä, S. and García-Fernández, Á. F. (2021). Temporal Parallelization of Bayesian Smoothers. IEEE Transactions on Automatic Control 66, 299–306, arXiv:1905.13002.
- Särkkä, S.; Solin, A. and Hartikainen, J. (2013). Spatiotemporal Learning via Infinite-Dimensional Bayesian Filtering and Smoothing: A Look at Gaussian Process Regression Through Kalman Filtering. IEEE Signal Processing Magazine 30, 51–61.
- Tax, D. M. and Duin, R. P. (2004). Support Vector Data Description. Machine Learning 54.
- Tsitsiklis, J. and Van Roy, B. (1997). An analysis of temporal-difference learning with function approximation. IEEE Transactions on Automatic Control 42, 674–690.
- Turk, G. and O'Brien, J. F. (2002). Modelling with implicit surfaces that interpolate. ACM Transactions on Graphics 21, 855–873.
- Verlet, L. (1967). Computer "Experiments" on Classical Fluids. I. Thermodynamical Properties of Lennard-Jones Molecules. Physical Review 159, 98–103. ↩1 ↩2
- Vincent, P. (2011). A Connection Between Score Matching and Denoising Autoencoders. Neural Computation 23, 1661–1674. ↩1
- Willems, J. (2007). The Behavioral Approach to Open and Interconnected Systems. IEEE Control Systems Magazine 27, x1–x1.
- Williams, O. and Fitzgibbon, A. (2006). Gaussian Process Implicit Surfaces. In: Gaussian Processes in Practice workshop.
- Wu Fung, S.; Heaton, H.; Li, Q.; McKenzie, D.; Osher, S. and Yin, W. (2022). JFB: Jacobian-Free Backpropagation for Implicit Networks. In: AAAI Conference on Artificial Intelligence, arXiv:2103.12803.
- Zhu, Y.; Zhang, K.; Liang, J.; Cao, J.; Wen, B.; Timofte, R. and Van Gool, L. (2023). Denoising Diffusion Models for Plug-and-Play Image Restoration. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), arXiv:2305.08995.