references

References

Bibliography

1

Samson Abramsky and Adam Brandenburger. The sheaf-theoretic structure of non-locality and contextuality. New Journal of Physics, 13(11):113036, 2011.

2

Peter Aczel and Nax Paul Mendler. A final coalgebra theorem. In Category Theory and Computer Science, volume 389 of Lecture Notes in Computer Science, pages 357–365. Springer, 1989.

3

Jiří Adámek and Jiří Rosický. Locally Presentable and Accessible Categories, volume 189 of London Mathematical Society Lecture Note Series. Cambridge University Press, 1994.

4

Shun-ichi Amari. Natural gradient works efficiently in learning. Neural Computation, 10(2):251–276, 1998. DOI https://doi.org/10.1162/089976698300017746.

5

Shun-ichi Amari. Information Geometry and Its Applications. Springer, 2016. DOI https://doi.org/10.1007/978-4-431-55978-8.

6

Anil Ananthaswamy. Is our universe a hologram? Physicists debate famous idea on its 25th anniversary. Scientific American, 328(3):58, 2023. DOI https://doi.org/10.1038/scientificamerican0323-58. URL https://doi.org/10.1038/scientificamerican0323-58.

7

Vladimir Araujo, Marie-Francine Moens, and Tinne Tuytelaars. Learning to route for dynamic adapter composition in continual learning with language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 687–696, 2024. DOI https://doi.org/10.18653/v1/2024.findings-emnlp.38.

8

Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, and Sepp Hochreiter. RUDDER: Return decomposition for delayed rewards. In Advances in Neural Information Processing Systems, volume 32, 2019.

9

V. I. Arnold and B. A. Khesin. Topological Methods in Hydrodynamics, volume 125 of Applied Mathematical Sciences. Springer, New York, 1999. ISBN 978-0-387-94947-5.

10

Federico Barbero, Cristian Bodnar, Haitz Sáez de Ocáriz Borde, Michael M. Bronstein, Petar Veličković, and Pietro Liò. Sheaf neural networks with connection laplacians. In Topological, Algebraic and Geometric Learning Workshops, volume 196 of Proceedings of Machine Learning Research, pages 28–36, 2022.

11

Michael Barr. Terminal coalgebras in well-founded set theory. Theoretical Computer Science, 114(2):299–315, 1993. DOI https://doi.org/10.1016/0304-3975(93)90076-6.

12

Michael Barr and Charles Wells. On the limitations of sketches. Canadian Mathematical Bulletin, 35(3):287–294, 1992. DOI https://doi.org/10.4153/CMB-1992-040-7. URL https://doi.org/10.4153/CMB-1992-040-7.

13

Michael Barr and Charles Wells. Category Theory for Computing Science. Centre de Recherches Mathématiques, 3 edition, 1999.

14

Claudio Battiloro, Lucia Testa, Lorenzo Giusti, Stefania Sardellitti, Paolo Di Lorenzo, and Sergio Barbarossa. Generalized simplicial attention neural networks. IEEE Transactions on Signal and Information Processing over Networks, 10:1–16, 2024.

15

Amir Beck and Marc Teboulle. Mirror descent and nonlinear projected subgradient methods for convex optimization. Operations Research Letters, 31(3):167–175, 2003.

16

Jean Bénabou. Introduction to bicategories. In Reports of the Midwest Category Seminar I, volume 47 of Lecture Notes in Mathematics, pages 1–77. Springer, Berlin, 1967. DOI https://doi.org/10.1007/BFb0074299.

17

Jean-David Benamou and Yann Brenier. A computational fluid mechanics solution to the Monge–Kantorovich mass transfer problem. Numerische Mathematik, 84(3):375–393, 2000. DOI https://doi.org/10.1007/s002110050002.

18

Peter J. Bickel, Chris A. J. Klaassen, Ya’acov Ritov, and Jon A. Wellner. Efficient and Adaptive Estimation for Semiparametric Models. Johns Hopkins University Press, 1993.

19

Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Baber Abbasi, Jessica Zosa Forde, Leo Gao, Jonathan Tow, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, et al. Lessons from the trenches on reproducible evaluation of language models, 2026. URL https://arxiv.org/abs/2405.14782.

20

Cristian Bodnar, Fabrizio Frasca, Nina Otter, Yu Guang Wang, Pietro Liò, Guido Montúfar, and Michael Bronstein. Weisfeiler and lehman go cellular: CW networks. In Advances in Neural Information Processing Systems, volume 34, pages 2625–2640, 2021.

21

Cristian Bodnar, Francesco Di Giovanni, Benjamin P. Chamberlain, Pietro Liò, and Michael M. Bronstein. Neural sheaf diffusion: A topological perspective on heterophily and oversmoothing in GNNs. In Advances in Neural Information Processing Systems, volume 35, pages 18527–18541, 2022a.

22

Cristian Bodnar, Francesco Di Giovanni, Benjamin Paul Chamberlain, Pietro Liò, and Michael M. Bronstein. Neural sheaf diffusion: A topological perspective on heterophily and oversmoothing in GNNs. In Advances in Neural Information Processing Systems, volume 35, pages 18527–18541, 2022b.

23

Jérôme Bolte, Ryan Boustany, Edouard Pauwels, and Béatrice Pesquet-Popescu. On the complexity of nonsmooth automatic differentiation, 2022. URL https://arxiv.org/abs/2206.01730.

24

Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models, 2021. URL https://arxiv.org/abs/2108.07258.

25

Nicolas Bonneel, Julien Rabin, Gabriel Peyré, and Hanspeter Pfister. Sliced and Radon wasserstein barycenters of measures. Journal of Mathematical Imaging and Vision, 51(1):22–45, 2015. DOI https://doi.org/10.1007/s10851-014-0506-3.

26

Byron Boots, Sajid M. Siddiqi, and Geoffrey J. Gordon. Closing the learning-planning loop with predictive state representations. The International Journal of Robotics Research, 30(7):954–966, 2011.

27

Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.

28

Philippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien, and Alexandre Drouin. Differentiable causal discovery from interventional data, 2020. URL https://arxiv.org/abs/2007.01754.

29

Matthew Burke and Rory B. B. MacAdam. Involution algebroids: a generalisation of Lie algebroids for tangent categories, 2019. URL https://arxiv.org/abs/1904.06594.

30

Jérémie Cabessa, Hugo Hernault, and Umer Mushtaq. Argument mining with fine-tuned large language models. In Proceedings of the 31st International Conference on Computational Linguistics, pages 6624–6635. Association for Computational Linguistics, 2025. URL https://aclanthology.org/2025.coling-main.442/.

31

Ruichu Cai, Zhiyi Huang, Wei Chen, Zhifeng Hao, and Kun Zhang. Causal discovery with latent confounders based on higher-order cumulants. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 3380–3407. PMLR, 2023. URL https://proceedings.mlr.press/v202/cai23a.html.

32

Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural ordinary differential equations. In Advances in Neural Information Processing Systems, 2018.

33

N. N. Chentsov. Statistical Decision Rules and Optimal Inference, volume 53 of Translations of Mathematical Monographs. American Mathematical Society, Providence, R.I., 1982. ISBN 0-8218-4502-0. Translation from the Russian edited by Lev J. Leifman.

34

Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, 21(1):C1–C68, 2018. DOI https://doi.org/10.1111/ectj.12097.

35

David Maxwell Chickering. Optimal structure identification with greedy equivalence search. Journal of Machine Learning Research, 3:507–554, 2002.

36

Kenta Cho and Bart Jacobs. Disintegration and Bayesian inversion via string diagrams. arXiv preprint arXiv:1709.00322, 2017. URL https://arxiv.org/abs/1709.00322.

37

Noam Chomsky. Aspects of the Theory of Syntax. MIT Press, Cambridge, MA, 1965. URL https://mitpress.mit.edu/9780262530071/aspects-of-the-theory-of-syntax/.

38

Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, volume 30, 2017.

39

Kacper Chwialkowski, Heiko Strathmann, and Arthur Gretton. A kernel test of goodness of fit. In Proceedings of the 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 2606–2615. PMLR, 2016. URL https://proceedings.mlr.press/v48/chwialkowski16.html.

40

Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try ARC, the AI2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018. URL https://arxiv.org/abs/1803.05457.

41

Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. In arXiv preprint arXiv:2110.14168, 2021. URL https://arxiv.org/abs/2110.14168.

42

J. R. B. Cockett and G. S. H. Cruttwell. Differential structure, tangent structure, and SDG. Applied Categorical Structures, 22:331–417, 2014a. DOI https://doi.org/10.1007/s10485-013-9312-0.

43

J. R. B. Cockett and G. S. H. Cruttwell. Differential structure, tangent structure, and SDG. Applied Categorical Structures, 22(2):331–417, 2014b.

44

J. R. B. Cockett and G. S. H. Cruttwell. The Jacobi identity for tangent categories. Cahiers de Topologie et Géométrie Différentielle Catégoriques, 56(4):301–316, 2015. URL https://www.reluctantm.com/gcruttw/publications/jacobiProof.pdf.

45

J. R. B. Cockett and G. S. H. Cruttwell. Connections in tangent categories. Theory and Applications of Categories, 32(26):835–888, 2017.

46

J. R. B. Cockett and G. S. H. Cruttwell. Differential bundles and fibrations for tangent categories. Cahiers de Topologie et Géométrie Différentielle Catégoriques, 59(1):10–92, 2018.

47

Robin Cockett, Geoffrey Cruttwell, Jonathan Gallagher, Jean-Simon Pacaud Lemay, Benjamin MacAdam, Gordon Plotkin, and Dorette Pronk. Reverse derivative categories. In Computer Science Logic, volume 152 of Leibniz International Proceedings in Informatics, pages 18:1–18:16, 2020.

48

Diego Colombo, Marloes H. Maathuis, Markus Kalisch, and Thomas S. Richardson. Learning high-dimensional directed acyclic graphs with latent and selection variables. The Annals of Statistics, 40(1):294–321, 2012. DOI https://doi.org/10.1214/11-AOS940.

49

Marius Crainic and Rui Loja Fernandes. Integrability of lie brackets. Annals of Mathematics, 157(2):575–620, 2003.

50

G. S. H. Cruttwell, Bruno Gavranovic, Neil Ghani, Paul Wilson, and Fabio Zanasi. Categorical foundations of gradient-based learning. Programming Languages and Systems, pages 1–28, 2022. DOI https://doi.org/10.1007/978-3-030-99336-8_1.

51

Geoffrey Cruttwell and Jean-Simon Pacaud Lemay. Reverse tangent categories. In 32nd EACSL Annual Conference on Computer Science Logic (CSL 2024), volume 288 of Leibniz International Proceedings in Informatics (LIPIcs), pages 21:1–21:21. Schloss Dagstuhl – Leibniz-Zentrum für Informatik, 2024. DOI https://doi.org/10.4230/LIPIcs.CSL.2024.21. URL https://arxiv.org/abs/2308.01131.

52

Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. In Advances in Neural Information Processing Systems, volume 26, 2013.

53

Charles Darwin. On the Origin of Species by Means of Natural Selection. John Murray, London, 1859. URL https://darwin-online.org.uk/converted/pdf/1859_Origin_F373.pdf.

54

Charles Darwin. The Autobiography of Charles Darwin, 1809–1882. Collins, London, 1958. URL https://darwin-online.org.uk/content/contentblock?itemID=F1497viewtype=sidebasepage=1hitpage=73. Edited by Nora Barlow; autobiographical text written 1876–1881.

55

Charles Darwin and Alfred Russel Wallace. On the tendency of species to form varieties; and on the perpetuation of varieties and species by natural means of selection. Journal of the Proceedings of the Linnean Society of London. Zoology, 3(9):45–62, 1858. DOI https://doi.org/10.1111/j.1096-3642.1858.tb02500.x.

56

Artur S. d’Avila Garcez, Krysia B. Broda, and Dov M. Gabbay. Neural-Symbolic Learning Systems: Foundations and Applications. Springer, London, 2002.

57

Akash Dhasade, Anne-Marie Kermarrec, Igor Pavlovic, Diana Petrescu, Rafael Pires, Mathis Randl, and Martijn de Vos. Effective LoRA adapter routing using task representations. arXiv preprint arXiv:2601.21795, 2026.

58

Ben Dickson. How moonshot engineered a 2.8t behemoth to run in production. AlphaSignal, Sunday Deep Dive newsletter, August 2026. Production-oriented commentary on the Kimi K3 architecture.

59

Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using Real NVP. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=HkpbnH9lx.

60

Arthur Conan Doyle. The Sign of the Four. Lippincott’s Monthly Magazine, 1890. URL https://www.gutenberg.org/ebooks/2097.

61

Frank Watson Dyson, Arthur Stanley Eddington, and Charles Davidson. A determination of the deflection of light by the sun’s gravitational field, from observations made at the total eclipse of may 29, 1919. Philosophical Transactions of the Royal Society of London. Series A, 220:291–333, 1920. DOI https://doi.org/10.1098/rsta.1920.0009.

62

Stefania Ebli, Michaël Defferrard, and Gard Spreemann. Simplicial neural networks. In NeurIPS Workshop on Topological Data Analysis and Beyond, 2020.

63

Charles Ehresmann. Esquisses et types des structures algébriques. Buletinul Institutului Politehnic din Iaşi, 14:1–14, 1968.

64

Albert Einstein. Die grundlage der allgemeinen relativitätstheorie. Annalen der Physik, 354(7):769–822, 1916. DOI https://doi.org/10.1002/andp.19163540702.

65

Hayder Elesedy, Pedro M. Esperanca, Silviu Vlad Oprea, and Mete Ozay. LoRA-guard: Parameter-efficient guardrail adaptation for content moderation of large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 11746–11765, 2024. DOI https://doi.org/10.18653/v1/2024.emnlp-main.656.

66

EleutherAI. Language model evaluation harness. Software, 2026. URL https://github.com/EleutherAI/lm-evaluation-harness.

67

exo-explore. exo: Run frontier AI locally. Software, 2026. URL https://github.com/exo-explore/exo.

68

Brendan Fong, David Spivak, and Rémy Tuyéras. Backprop as functor: A compositional perspective on supervised learning. Proceedings of the 34th Annual ACM/IEEE Symposium on Logic in Computer Science, pages 1–13, 2019.

69

Tobias Fritz. A synthetic approach to Markov kernels, conditional independence and theorems on sufficient statistics. Advances in Mathematics, 370:107239, 2020. DOI https://doi.org/10.1016/j.aim.2020.107239. URL https://arxiv.org/abs/1908.07021.

70

Tobias Fritz and Andreas Klingler. The d-Separation criterion in categorical probability. Journal of Machine Learning Research, 24(46):1–49, 2023. URL https://jmlr.org/papers/v24/22-0916.html.

71

Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In Proceedings of the 40th International Conference on Machine Learning, volume 202, pages 10835–10866, 2023.

72

Bruno Gavranović, Paul Lessard, Andrew Dudzik, Tamara von Glehn, João G. M. Araújo, and Petar Veličković. Position: Categorical deep learning is an algebraic theory of all architectures, 2024. URL https://arxiv.org/abs/2402.15332.

73

Steven B. Gillispie and Michael D. Perlman. The size distribution for markov equivalence classes of acyclic digraph models. Artificial Intelligence, 141(1–2):137–155, 2002. DOI https://doi.org/10.1016/S0004-3702(02)00264-3.

74

Clark Glymour, Kun Zhang, and Peter Spirtes. Review of causal discovery methods based on graphical models. Frontiers in Genetics, 10, 2019. ISSN 1664-8021. DOI https://doi.org/10.3389/fgene.2019.00524. URL https://www.frontiersin.org/journals/genetics/articles/10.3389/fgene.2019.00524.

75

Christopher Wei Jin Goh, Cristian Bodnar, and Pietro Liò. Simplicial attention networks, 2022.

76

E Mark Gold. Language identification in the limit. Information and Control, 10(5):447–474, 1967. ISSN 0019-9958. DOI https://doi.org/https://doi.org/10.1016/S0019-9958(67)91165-5. URL https://www.sciencedirect.com/science/article/pii/S0019995867911655.

77

Jackson Gorham and Lester Mackey. Measuring sample quality with kernels. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 1292–1301. PMLR, 2017. URL https://proceedings.mlr.press/v70/gorham17a.html.

78

Peter D. Grünwald. The Minimum Description Length Principle. MIT Press, Cambridge, MA, 2007. ISBN 9780262072816.

79

Siyuan Guo, Viktor Tóth, Bernhard Schölkopf, and Ferenc Huszár. Causal de Finetti: On the identification of invariant causal structure in exchangeable data, 2022. URL https://arxiv.org/abs/2203.15756.

80

Siyuan Guo, Chi Zhang, Karthika Mohan, Ferenc Huszár, and Bernhard Schölkopf. Do Finetti: On causal effects for exchangeable data, 2024. URL https://arxiv.org/abs/2405.18836.

81

Ankita Gupta, Ethan Zuckerman, and Brendan O’Connor. Harnessing Toulmin’s theory for zero-shot argument explication. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10259–10276, Bangkok, Thailand, 2024. Association for Computational Linguistics. DOI https://doi.org/10.18653/v1/2024.acl-long.552. URL https://aclanthology.org/2024.acl-long.552/.

82

Ivan Habernal and Iryna Gurevych. Argumentation mining in user-generated web discourse. Computational Linguistics, 43(1):125–179, apr 2017. DOI https://doi.org/10.1162/COLI_a_00276. URL https://aclanthology.org/J17-1004/.

83

Elie Habib. Worldmonitor: Real-time global intelligence dashboard. GitHub repository, 2026. URL https://github.com/koala73/worldmonitor. Accessed July 21, 2026.

84

Mustafa Hajij, Ghada Zamzmi, Theodore Papamarkou, Nina Miolane, Aldo Guzmán-Sáenz, Karthikeyan Natesan Ramamurthy, Tolga Birdal, Tamal K. Dey, Soham Mukherjee, Shreyas N. Samaga, Neal Livesay, Robin Walters, Paul Rosen, and Michael T. Schaub. Topological deep learning: Going beyond graph data, 2022.

85

Mustafa Hajij, Lennart Bastian, Sarah Osentoski, Hardik Kabaria, John L. Davenport, Sheik Dawood, Balaji Cherukuri, Joseph G. Kocheemoolayil, Nastaran Shahmansouri, Adrian Lew, Theodore Papamarkou, and Tolga Birdal. Copresheaf topological neural networks: A generalized deep learning framework. arXiv, 2025. URL https://arxiv.org/abs/2505.21251.

86

Jakob Hansen and Thomas Gebhart. Sheaf neural networks. In NeurIPS Workshop on Topological Data Analysis and Beyond, 2020.

87

F. Maxwell Harper and Joseph A. Konstan. The MovieLens datasets: History and context. ACM Transactions on Interactive Intelligent Systems, 5(4):19:1–19:19, 2015.

88

Anna Harutyunyan, Will Dabney, Thomas Mesnard, Mohammad Gheshlaghi Azar, Bilal Piot, Nicolas Heess, Hado van Hasselt, Gregory Wayne, Satinder Singh, Doina Precup, and Remi Munos. Hindsight credit assignment. In Advances in Neural Information Processing Systems, volume 32, 2019.

89

Trevor Hastie, Robert Tibshirani, and Jerome Friedman. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer-Verlag, 2nd edition, 2009.

90

Alain Hauser and Peter Bühlmann. Characterization and greedy learning of interventional Markov equivalence classes of directed acyclic graphs. Journal of Machine Learning Research, 13(Aug):2409–2464, 2012.

91

Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ.

92

Ralf Hinze and Dan Marsden. Introducing String Diagrams: The Art of Category Theory. Cambridge University Press, 2023. DOI https://doi.org/10.1017/9781009317825. URL https://doi.org/10.1017/9781009317825.

93

Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, 2020.

94

Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning, 2019.

95

Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022.

96

Chengsong Huang, Qian Liu, Bill Yuchen Lin, Tianyu Pang, Chao Du, and Min Lin. LoraHub: Efficient cross-task generalization via dynamic LoRA composition. In Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=lyRpY2bpBh.

97

Yucong Huang, Xiucheng Li, Kaiqi Zhao, and Jing Li. Transitivity meets cyclicity: Explicit preference decomposition for dynamic large language model alignment, 2026. URL https://arxiv.org/abs/2605.17342.

98

Aapo Hyvärinen and Stephen M. Smith. Pairwise likelihood ratios for estimation of non-Gaussian structural equation models. Journal of Machine Learning Research, 14:111–152, 2013. URL https://www.jmlr.org/papers/v14/hyvarinen13a.html.

99

Takashi Ikeuchi, Mayumi Ide, Yan Zeng, Takashi N. Maeda, and Shohei Shimizu. Python package for causal discovery based on LiNGAM. Journal of Machine Learning Research, 24(14):1–8, 2023. URL https://jmlr.org/papers/v24/21-0321.html.

100

Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. In International Conference on Learning Representations, 2023.

101

Guido W. Imbens and Donald B. Rubin. Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press, USA, 2015. ISBN 0521885884.

102

Amin Jaber, Murat Kocaoglu, Karthikeyan Shanmugam, and Elias Bareinboim. Causal discovery from soft interventions with unknown targets: Characterization and learning. In Advances in Neural Information Processing Systems, volume 33, pages 9551–9561, 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/hash/6cd9313ed34ef58bad3fdd504355e72c-Abstract.html.

103

Bart Jacobs, Aleks Kissinger, and Fabio Zanasi. Causal inference by string diagram surgery, 2018. URL https://arxiv.org/abs/1811.08338.

104

Xiaoye Jiang, Lek-Heng Lim, Yiyu Yao, and Yinyu Ye. Statistical ranking and combinatorial Hodge theory. Mathematical Programming, 127(1):203–244, 2011.

105

Anatoli Juditsky, Arkadi Nemirovski, and Claire Tauvel. Solving variational inequalities with stochastic mirror-prox algorithm. Stochastic Systems, 1(1):17–58, 2011.

106

Sham M. Kakade. A natural policy gradient. In Advances in Neural Information Processing Systems 14. MIT Press, 2001. URL https://proceedings.neurips.cc/paper/2001/hash/4b86abe48d358ecf194c56c69108433e-Abstract.html.

107

G. M. Kelly. Basic Concepts of Enriched Category Theory, volume 64 of London Mathematical Society Lecture Note Series. Cambridge University Press, 1982.

108

Yundong Kim and Heyoung Yang. TRACE: Toulmin-based reasoning assessment through constructive elements for LLM CoT evaluation, 2026. URL https://arxiv.org/abs/2605.29656.

109

Kimi Team. Kimi K3: Open frontier intelligence. arXiv preprint arXiv:2607.24653, 2026. URL https://arxiv.org/abs/2607.24653.

110

Anders Kock. Synthetic Differential Geometry. Cambridge University Press, 2 edition, 2006. URL https://users-math.au.dk/kock/sdg99.pdf.

111

Dexter Kozen and Nicholas Ruozzi. Applications of metric coinduction. Logical Methods in Computer Science, 5(3:10):1–19, 2009. DOI https://doi.org/10.2168/LMCS-5(3:10)2009.

112

Sébastien Lachapelle, Pierre Brouillard, Tristan Deleu, and Simon Lacoste-Julien. Gradient-based neural dag learning. In International Conference on Learning Representations, 2020.

113

LangChain. LangGraph overview. Documentation, 2026. URL https://docs.langchain.com/oss/python/langgraph/overview.

114

F. William Lawvere. Functorial semantics of algebraic theories. Proceedings of the National Academy of Sciences, 50(5):869–872, 1963. DOI https://doi.org/10.1073/pnas.50.5.869. URL https://pmc.ncbi.nlm.nih.gov/articles/PMC221940/.

115

F. William Lawvere. Adjointness in foundations. Dialectica, 23(3–4):281–296, 1969.

116

F. William Lawvere. Toward the description in a smooth topos of the dynamically possible motions and deformations of a continuous body. Cahiers de Topologie et Géométrie Différentielle Catégoriques, 21(4):377–392, 1980. URL https://www.numdam.org/item/CTGDC_1980__21_4_377_0/.

117

John M. Lee. Introduction to Smooth Manifolds, volume 218 of Graduate Texts in Mathematics. Springer, 2 edition, 2012. DOI https://doi.org/10.1007/978-1-4419-9982-5.

118

Douglas B. Lenat. AM: An Artificial Intelligence Approach to Discovery in Mathematics as Heuristic Search. PhD thesis, Stanford University, 1976. URL https://en.wikisource.org/wiki/An_Artificial_Intelligence_Approach_to_Discovery_in_Mathematics_as_Heuristic_Search. Stanford Artificial Intelligence Laboratory Memo AIM-286; Computer Science Report STAN-CS-76-750.

119

Poon Leung. Classifying tangent structures using Weil algebras. Theory and Applications of Categories, 32(9):286–337, 2017a. URL http://www.tac.mta.ca/tac/volumes/32/9/32-09.pdf.

120

Poon Leung. Tangent Bundles, Monoidal Theories, and Weil Algebras. Ph.d. thesis, Macquarie University, 2017b.

121

Poon Leung. Classifying tangent structures using Weil algebras. Theory and Applications of Categories, 32(9):286–337, 2017c. URL http://www.tac.mta.ca/tac/volumes/32/9/32-09.pdf.

122

Ming Li and Paul Vitányi. An Introduction to Kolmogorov Complexity and Its Applications. Springer, Cham, 4 edition, 2019. DOI https://doi.org/10.1007/978-3-030-11298-1.

123

Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori B. Hashimoto. Diffusion-lm improves controllable text generation. In Advances in Neural Information Processing Systems, volume 35, pages 4328–4343, 2022.

124

Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=PqvMRDCJT9t.

125

Michael L. Littman, Richard S. Sutton, and Satinder Singh. Predictive representations of state. In Advances in Neural Information Processing Systems, 2001.

126

Bo Liu, Ji Liu, Mohammad Ghavamzadeh, Sridhar Mahadevan, and Marek Petrik. Finite-sample analysis of proximal gradient TD algorithms. In Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence, pages 504–513, 2015.

127

Qiang Liu, Jason D. Lee, and Michael I. Jordan. A kernelized Stein discrepancy for goodness-of-fit tests and model evaluation. In Proceedings of the 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 276–284. PMLR, 2016. URL https://proceedings.mlr.press/v48/liub16.html.

128

Marloes H Maathuis, Markus Kalisch, and Peter Bühlmann. Estimating high-dimensional intervention effects from observational data. The Annals of Statistics, 37(6A):3133–3164, 2009.

129

Saunders Mac Lane. Categories for the Working Mathematician. Springer-Verlag, New York, 1971. Graduate Texts in Mathematics, Vol. 5.

130

Saunders Mac Lane. Categories for the Working Mathematician. Springer, second edition, 1998.

131

Saunders Mac Lane and Ieke Moerdijk. Sheaves in Geometry and Logic: A First Introduction to Topos Theory. Springer New York, New York, NY, 1992. ISBN 9781461209270 1461209277. URL http://link.springer.com/book/10.1007/978-1-4612-0927-0.

132

Benjamin MacAdam. The Functorial Semantics of Lie Theory. PhD thesis, University of Calgary, 2022. URL https://arxiv.org/abs/2301.00305.

133

Kirill C. H. Mackenzie. General Theory of Lie Groupoids and Lie Algebroids, volume 213 of London Mathematical Society Lecture Note Series. Cambridge University Press, 2005.

134

Takashi N. Maeda and Shohei Shimizu. RCD: Repetitive causal discovery of linear non-Gaussian acyclic models with latent confounders. In Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 of Proceedings of Machine Learning Research, pages 735–745. PMLR, 2020. URL https://proceedings.mlr.press/v108/maeda20a.html.

135

Takashi N. Maeda and Shohei Shimizu. Causal additive models with unobserved variables. In Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence, volume 161 of Proceedings of Machine Learning Research, pages 97–106. PMLR, 2021. URL https://proceedings.mlr.press/v161/maeda21a.html.

136

Hamid Reza Maei. Gradient Temporal-Difference Learning Algorithms. PhD thesis, University of Alberta, Edmonton, Alberta, Canada, 2011.

137

Sridhar Mahadevan. Universal causality. Entropy, 25(4):574, 2023. DOI https://doi.org/10.3390/e25040574. URL https://doi.org/10.3390/e25040574.

138

Sridhar Mahadevan. Decentralized causal discovery using judo calculus, 2025a. URL https://arxiv.org/abs/2510.23942.

139

Sridhar Mahadevan. Higher algebraic K-theory of causality. Entropy, 27(5):531, 2025b. DOI https://doi.org/10.3390/e27050531. URL https://doi.org/10.3390/e27050531.

140

Sridhar Mahadevan. Intuitionistic \(j\)-do-calculus in topos causal models, 2025c. URL https://arxiv.org/abs/2510.17944.

141

Sridhar Mahadevan. Large causal models from large language models, 2025d. URL https://arxiv.org/abs/2512.07796.

142

Sridhar Mahadevan. ALLORA: A Lie-algebraic LoRA method for composable neural adapters, 2026a. Manuscript in preparation.

143

Sridhar Mahadevan. Latent confounded causal discovery via Lie bracket geometry, 2026b. URL https://arxiv.org/abs/2606.19610.

144

Sridhar Mahadevan. Categories for AGI. Book manuscript, 2026c. URL https://people.cs.umass.edu/ mahadeva/papers/catagi.pdf. Revised July 25, 2026.

145

Sridhar Mahadevan. Causal density functions, 2026d. URL https://arxiv.org/abs/2606.00754.

146

Sridhar Mahadevan. Gradient infinitesimal reinforcement learning in tangent categories. Manuscript in preparation, 2026e.

147

Sridhar Mahadevan. Infinitesimal causality. arXiv preprint arXiv:2606.24621, 2026f. URL https://arxiv.org/abs/2606.24621.

148

Sridhar Mahadevan. Kan extension transformers: A categorical unification of attention, diffusion, and predict-detach self-conditioning. arXiv preprint arXiv:2605.27259, 2026g. URL https://arxiv.org/abs/2605.27259.

149

Sridhar Mahadevan. Agentic skill optimization over Lie algebroids. arXiv preprint arXiv:2607.11493, 2026h. URL https://arxiv.org/abs/2607.11493.

150

Sridhar Mahadevan. Learning in infinitesimal non-compositional sketches. arXiv preprint arXiv:2607.15107, 2026i. URL https://arxiv.org/abs/2607.15107.

151

Sridhar Mahadevan. ODYSSEY: Constructing verifiable, local truth-preserving foundation models, 2026j. Forthcoming arXiv preprint.

152

Sridhar Mahadevan. PROMETHEUS: Automating deep causal research integrating text, data and models, 2026k. URL https://arxiv.org/abs/2605.12835.

153

Sridhar Mahadevan. Universal decision learners. arXiv preprint arXiv:2605.30694, 2026l. URL https://arxiv.org/abs/2605.30694.

154

Michael Makkai and Robert Paré. Accessible Categories: The Foundations of Categorical Model Theory, volume 104 of Contemporary Mathematics. American Mathematical Society, 1989.

155

Juan M. Maldacena. The large N limit of superconformal field theories and supergravity. Advances in Theoretical and Mathematical Physics, 2:231–252, 1998. DOI https://doi.org/10.1023/A:1026654312961. URL https://arxiv.org/abs/hep-th/9711200.

156

Thomas Robert Malthus. An Essay on the Principle of Population, as It Affects the Future Improvement of Society. J. Johnson, London, 1798. URL https://www.gutenberg.org/ebooks/4239.

157

Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. Building a large annotated corpus of english: The penn treebank. Computational Linguistics, 19(2):313–330, 1993.

158

Dan Marsden. Category theory using string diagrams. arXiv preprint arXiv:1401.7220, 2014. URL https://arxiv.org/abs/1401.7220.

159

James Martens and Roger Grosse. Optimizing neural networks with kronecker-factored approximate curvature. In Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 2408–2417. PMLR, 2015. URL https://proceedings.mlr.press/v37/martens15.html.

160

Mathematical Association of America. American Invitational Mathematics Examination (AIME) 2024. American Mathematics Competitions problem set, 2024. URL https://maa.org/maa-invitational-competitions/.

161

J.P. May. Simplicial Objects in Algebraic Topology. University of Chicago Press, 1992.

162

John McCarthy. Challenges to machine learning: Relations between reality and appearance. In Inductive Logic Programming: 16th International Conference, ILP 2006, Revised Selected Papers, volume 4455 of Lecture Notes in Computer Science, pages 2–9. Springer, 2007. DOI https://doi.org/10.1007/978-3-540-73847-3_2. URL http://jmc.stanford.edu/articles/appearance/appearance.pdf.

163

Leland McInnes, John Healy, and James Melville. UMAP: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018.

164

Rui Meng, Bhavana Dalvi Mishra, Jiefeng Chen, Chun-Liang Li, Palash Goyal, Mihir Parmar, Yiwen Song, Yale Song, Rajarishi Sinha, Parthasarathy Ranganathan, Burak Gokturk, Jinsung Yoon, and Tomas Pfister. ScientistOne: Towards human-level autonomous research via chain-of-evidence. arXiv preprint arXiv:2605.26340, 2026. DOI https://doi.org/10.48550/arXiv.2605.26340. URL https://arxiv.org/abs/2605.26340.

165

Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In International Conference on Learning Representations, 2017.

166

Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015. DOI https://doi.org/10.1038/nature14236.

167

Ieke Moerdijk and Gonzalo E. Reyes. Models for Smooth Infinitesimal Analysis. Springer, New York, 1991. DOI https://doi.org/10.1007/978-1-4757-4143-8.

168

Yoshimitsu Morinishi and Shohei Shimizu. Differentiable causal discovery of linear non-Gaussian acyclic models under unmeasured confounding. Transactions on Machine Learning Research, 2025. URL https://openreview.net/forum?id=HR7MFlW73I.

169

Kevin P. Murphy. Machine Learning: A Probabilistic Perspective. The MIT Press, 2012. ISBN 978-0262018020.

170

Ashwin Narayan, Bonnie Berger, and Hyunghoon Cho. Assessing single-cell transcriptomic variability through density-preserving data visualization. Nature Biotechnology, 39:765–774, 2021.

171

Matteo Negro, Andrea Piras, Ragib Ahsan, David Arbour, and Elena Zheleva. Relational causal discovery with latent confounders. In Proceedings of the Forty-first Conference on Uncertainty in Artificial Intelligence, volume 286 of Proceedings of Machine Learning Research, pages 3123–3154. PMLR, 2025. URL https://proceedings.mlr.press/v286/negro25a.html.

172

Arkadi Nemirovski. Prox-method with rate of convergence O(1/t) for variational inequalities with Lipschitz continuous monotone operators and smooth convex–concave saddle point problems. SIAM Journal on Optimization, 15(1):229–251, 2004.

173

Allen Newell and Herbert A. Simon. Computer science as empirical inquiry: Symbols and search. Communications of the ACM, 19(3):113–126, 1976. DOI https://doi.org/10.1145/360018.360022.

174

Juan Miguel Ogarrio, Peter Spirtes, and Joseph Ramsey. A hybrid causal search algorithm for latent variable models. In Proceedings of the Eighth International Conference on Probabilistic Graphical Models, pages 368–379, 2016.

175

George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. Normalizing flows for probabilistic modeling and inference. Journal of Machine Learning Research, 22(57):1–64, 2021. URL https://www.jmlr.org/papers/v22/19-1028.html.

176

Judea Pearl. Causality: Models, Reasoning and Inference. Cambridge University Press, USA, 2nd edition, 2009a. ISBN 052189560X.

177

Judea Pearl. Causality: Models, Reasoning, and Inference. Cambridge University Press, 2 edition, 2009b.

178

Judea Pearl. Theoretical impediments to machine learning with seven sparks from the causal revolution, 2018. URL https://arxiv.org/abs/1801.04016.

179

Judea Pearl and Dana Mackenzie. The Book of Why: The New Science of Cause and Effect. Basic Books, New York, 2018. ISBN 978-0-465-09760-9.

180

William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4195–4205, 2023.

181

Andreas Peldszus and Manfred Stede. An annotated corpus of argumentative microtexts. In Proceedings of the First European Conference on Argumentation, 2015. URL https://peldszus.github.io/files/eca2015-preprint.pdf.

182

Jan Peters and Stefan Schaal. Natural actor-critic. Neurocomputing, 71(7–9):1180–1190, 2008. DOI https://doi.org/10.1016/j.neucom.2007.11.026.

183

Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. AdapterFusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, pages 487–503, 2021. DOI https://doi.org/10.18653/v1/2021.eacl-main.39.

184

Max Planck. Ueber das gesetz der energieverteilung im normalspectrum. Annalen der Physik, 309(3):553–563, 1901. DOI https://doi.org/10.1002/andp.19013090310.

185

Clifton Poth, Hannah Sterz, Indraneil Paul, Sukannya Purkayastha, Leon Engländer, Timo Imhof, Ivan Vulić, Sebastian Ruder, Iryna Gurevych, and Jonas Pfeiffer. Adapters: A unified library for parameter-efficient and modular transfer learning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 149–160, 2023. DOI https://doi.org/10.18653/v1/2023.emnlp-demo.13.

186

Akshara Prabhakar, Yuanzhi Li, Karthik Narasimhan, Sham Kakade, Eran Malach, and Samy Jelassi. LoRA soups: Merging LoRAs for practical skill composition tasks. In Proceedings of the 31st International Conference on Computational Linguistics: Industry Track, pages 644–655, 2025.

187

Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems, volume 36, pages 53728–53741, 2023.

188

Balaraman Ravindran. An Algebraic Approach to Abstraction in Reinforcement Learning. PhD thesis, University of Massachusetts Amherst, 2004.

189

David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google-proof Q&A benchmark. arXiv preprint arXiv:2311.12022, 2023. URL https://arxiv.org/abs/2311.12022.

190

Danilo Jimenez Rezende and Shakir Mohamed. Variational inference with normalizing flows. In Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1530–1538. PMLR, 2015. URL https://proceedings.mlr.press/v37/rezende15.html.

191

B. Richter. From Categories to Homotopy Theory. Cambridge Studies in Advanced Mathematics. Cambridge University Press, 2020. ISBN 9781108479622. URL https://books.google.com/books?id=pnzUDwAAQBAJ.

192

Emily Riehl. Category Theory in Context. Dover Publications, 2017.

193

Jorma Rissanen. Modeling by shortest data description. Automatica, 14(5):465–471, 1978. DOI https://doi.org/10.1016/0005-1098(78)90005-5.

194

Graeme D. Ritchie and F. Keith Hanna. AM: A case study in AI methodology. Artificial Intelligence, 23(3):249–268, 1984. DOI https://doi.org/10.1016/0004-3702(84)90015-8.

195

James M. Robins, Andrea Rotnitzky, and Lue Ping Zhao. Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association, 89(427):846–866, 1994. DOI https://doi.org/10.1080/01621459.1994.10476818.

196

R. Tyrrell Rockafellar. Generalized directional derivatives and subgradients of nonconvex functions. Canadian Journal of Mathematics, 32(2):257–280, 1980. DOI https://doi.org/10.4153/CJM-1980-020-7.

197

Jiří Rosický. Abstract tangent functors. Diagrammes, 12:1–11, 1984.

198

Donald B. Rubin. Causal inference using potential outcomes: Design, modeling, decisions. Journal of the American Statistical Association, 100(469):322–331, 2005. DOI https://doi.org/10.1198/016214504000001880.

199

Jan J. M. M. Rutten. Universal coalgebra: a theory of systems. Theoretical Computer Science, 249(1):3–80, 2000.

200

Karen Sachs, Omar Perez, Dana Pe’er, Douglas A. Lauffenburger, and Garry P. Nolan. Causal protein-signaling networks derived from multiparameter single-cell data. Science, 308(5721):523–529, 2005. DOI https://doi.org/10.1126/science.1105809.

201

Michael Schlichtkrull, Thomas N. Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. Modeling relational data with graph convolutional networks. In The Semantic Web, pages 593–607. Springer, 2018.

202

Dominik Schmid and Allan Sly. On the number and size of markov equivalence classes of random directed acyclic graphs. arXiv:2209.04395v2 [math.PR], 2024. URL https://arxiv.org/abs/2209.04395. Submitted 2022; revised 2024. DOI https://doi.org/10.48550/arXiv.2209.04395.

203

John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz. Trust region policy optimization. In Proceedings of the 32nd International Conference on Machine Learning, volume 37, pages 1889–1897, 2015.

204

A. Sepúlveda-Jiménez. Synthetic machine learning. ResearchGate preprint, July 2025. URL https://doi.org/10.13140/RG.2.2.26283.96802.

205

Abhishek Sharma, Leo Benac, Sonali Parbhoo, and Finale Doshi-Velez. Decision-point guided safe policy improvement. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics, volume 258 of Proceedings of Machine Learning Research, pages 2935–2943. PMLR, 2025. URL https://proceedings.mlr.press/v258/sharma25a.html.

206

Shohei Shimizu, Patrik O. Hoyer, Aapo Hyvärinen, and Antti Kerminen. A linear non-Gaussian acyclic model for causal discovery. Journal of Machine Learning Research, 7:2003–2030, 2006. URL https://www.jmlr.org/papers/v7/shimizu06a.html.

207

Shohei Shimizu, Takanori Inazumi, Yasuhiro Sogawa, Aapo Hyvärinen, Yoshinobu Kawahara, Takashi Washio, Patrik O. Hoyer, and Kenneth Bollen. DirectLiNGAM: A direct method for learning a linear non-Gaussian structural equation model. Journal of Machine Learning Research, 12:1225–1248, 2011. URL https://www.jmlr.org/papers/v12/shimizu11a.html.

208

Satinder Singh, Michael R. James, and Matthew R. Rudary. Predictive state representations: A new theory for modeling dynamical systems. Proceedings of the 20th Conference on Uncertainty in Artificial Intelligence, 2004a.

209

Satinder Singh, Michael R. James, and Matthew R. Rudary. Predictive state representations: A new theory for modeling dynamical systems. In Proceedings of the 20th Conference on Uncertainty in Artificial Intelligence, pages 512–519, 2004b.

210

Paul Smolensky. On the proper treatment of connectionism. Behavioral and Brain Sciences, 11(1):1–23, 1988. DOI https://doi.org/10.1017/S0140525X00052432.

211

Vishal Soni and Satinder Singh. Abstraction in predictive state representations. In Proceedings of the Twenty-Second AAAI Conference on Artificial Intelligence, pages 639–644, 2007.

212

Elizabeth S. Spelke. Core knowledge. American Psychologist, 55(11):1233–1243, 2000. DOI https://doi.org/10.1037/0003-066X.55.11.1233. URL https://doi.org/10.1037/0003-066X.55.11.1233.

213

Elizabeth S. Spelke and Katherine D. Kinzler. Core knowledge. Developmental Science, 10(1):89–96, 2007. DOI https://doi.org/10.1111/j.1467-7687.2007.00569.x. URL https://doi.org/10.1111/j.1467-7687.2007.00569.x.

214

Peter Spirtes, Clark Glymour, and Richard Scheines. Causation, Prediction, and Search. MIT Press, 2 edition, 2000.

215

David I. Spivak. Functorial data migration. Information and Computation, 217:31–51, 2012. DOI https://doi.org/10.1016/j.ic.2012.05.001.

216

David I. Spivak. Database queries and constraints via lifting problems. Mathematical Structures in Computer Science, 24(6):e240602, 2014. DOI https://doi.org/10.1017/S0960129513000479. URL https://doi.org/10.1017/S0960129513000479.

217

Ross Street. Categorical structures. In Handbook of Algebra, volume 1, pages 529–577. Elsevier, 1996. DOI https://doi.org/10.1016/S1570-7954(96)80019-2.

218

Ross Street. Frobenius monads and pseudomonoids. Journal of Mathematical Physics, 45(10):3930–3948, 2004. DOI https://doi.org/10.1063/1.1788852.

219

Masashi Sugiyama, Taiji Suzuki, and Takafumi Kanamori. Density Ratio Estimation in Machine Learning. Cambridge University Press, 2012. DOI https://doi.org/10.1017/CBO9781139035613.

220

Richard S. Sutton. Learning to predict by the methods of temporal differences. Machine Learning, 3:9–44, 1988. DOI https://doi.org/10.1007/BF00115009.

221

Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. Adaptive Computation and Machine Learning Series. The MIT Press, second edition, 2018. ISBN 9780262039246. URL http://incompleteideas.net/book/the-book.html.

222

Richard S. Sutton, Csaba Szepesvári, and Hamid Reza Maei. A convergent O(n) temporal-difference algorithm for off-policy learning with linear function approximation. Advances in Neural Information Processing Systems, 21, 2008.

223

Gokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu, and Alekh Agarwal. A minimaximalist approach to reinforcement learning from human feedback. In Proceedings of the 41st International Conference on Machine Learning, volume 235, pages 47345–47377, 2024.

224

Dennis Tang, Prateek Yadav, Yi-Lin Sung, Jaehong Yoon, and Mohit Bansal. LoRA merging with SVD: Understanding interference and preserving performance. In ICML Workshop on Resource-Efficient Foundation Models, 2025.

225

Tatsuya Tashiro, Shohei Shimizu, Aapo Hyvärinen, and Takashi Washio. ParceLiNGAM: A causal ordering method robust against latent confounders. Neural Computation, 26(1):57–83, 2014. DOI https://doi.org/10.1162/NECO_a_00533.

226

Philip S. Thomas, William Dabney, Stephen Giguere, and Sridhar Mahadevan. Projected natural actor-critic. In Christopher J. C. Burges, Léon Bottou, Zoubin Ghahramani, and Kilian Q. Weinberger, editors, Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, United States, pages 2337–2345, 2013. URL https://proceedings.neurips.cc/paper/2013/hash/dd77279f7d325eec933f05b1672f6a1f-Abstract.html.

227

Stephen E. Toulmin. The Uses of Argument. Cambridge University Press, 1958.

228

Daniele Tramontano, Yaroslav Kivva, Saber Salehkaleybar, Mathias Drton, and Negar Kiyavash. Causal effect identification in lvLiNGAM from higher-order cumulants. arXiv preprint arXiv:2506.05202, 2025. URL https://arxiv.org/abs/2506.05202.

229

Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting, 2023. URL https://arxiv.org/abs/2305.04388.

230

Ruben van Belle. Kan Extensions in Probability Theory. PhD thesis, University of Edinburgh, 2024.

231

Alexander Van-Brunt and Matt Visser. Special cases of the baker–campbell–hausdorff formula. Journal of Mathematical Physics, 57(2), 2016. DOI https://doi.org/10.1063/1.4940799.

232

Mark J. van der Laan and Sherri Rose. Targeted Learning: Causal Inference for Observational and Experimental Data. Springer, 2011. DOI https://doi.org/10.1007/978-1-4419-9782-1.

233

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html.

234

Cédric Villani. Optimal Transport: Old and New, volume 338 of Grundlehren der mathematischen Wissenschaften. Springer, Berlin, 2009. DOI https://doi.org/10.1007/978-3-540-71050-9.

235

Eric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang, Eshaan Pathak, Matthew Ginsberg, and Dan Klein. Automated crossword solving. arXiv preprint arXiv:2205.09665, 2022.

236

Yingfan Wang, Haiyang Huang, Cynthia Rudin, and Yaron Shaposhnik. Understanding how dimension reduction tools work: An empirical approach to deciphering t-SNE, UMAP, TriMap, and PaCMAP for data visualization. Journal of Machine Learning Research, 22(201):1–73, 2021.

237

Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. MMLU-Pro: A more robust and challenging multi-task language understanding benchmark. In Advances in Neural Information Processing Systems, 2024. URL https://arxiv.org/abs/2406.01574.

238

Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2022. URL https://arxiv.org/abs/2201.11903.

239

Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. TIES-merging: Resolving interference when merging models. In Advances in Neural Information Processing Systems, 2023.

240

Yifan Yang, Ziyang Gong, Weiquan Huang, Qihao Yang, Ziwei Zhou, Zisu Huang, Yan Li, Xuemei Gao, Qi Dai, Bei Liu, Kai Qiu, Yuqing Yang, Dongdong Chen, Xue Yang, and Chong Luo. SkillOpt: Executive strategy for self-evolving agent skills, 2026. URL https://arxiv.org/abs/2605.23904.

241

Deniz Yeral. Frobenius algebras, factorization homology and the Reshetikhin–Turaev invariants, 2025. URL https://arxiv.org/abs/2508.16351.

242

Olga Zaghen, Antonio Longa, Steve Azzolin, Lev Telyatnikov, Andrea Passerini, and Pietro Liò. Sheaf diffusion goes nonlinear: Enhancing GNNs with adaptive sheaf laplacians. In Proceedings of the Geometry-grounded Representation Learning and Generative Modeling Workshop, volume 251 of Proceedings of Machine Learning Research, pages 264–276. PMLR, 2024. URL https://proceedings.mlr.press/v251/zaghen24a.html.

243

Tom Zahavy. Position: LLMs can’t jump. In Proceedings of the 43rd International Conference on Machine Learning, 2026. URL https://openreview.net/forum?id=klU4737opt. Position paper.

244

Haobo Zhang and Jiayu Zhou. Unraveling LoRA interference: Orthogonal subspaces for robust model merging. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, pages 26459–26472. Association for Computational Linguistics, 2025. DOI https://doi.org/10.18653/v1/2025.acl-long.1284. URL https://aclanthology.org/2025.acl-long.1284/.

245

Haopeng Zhang, Xiao Liu, and Jiawei Zhang. HEGEL: Hypergraph transformer for long document summarization. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10167–10176, 2022.

246

Juzheng Zhang, Jiacheng You, Ashwinee Panda, and Tom Goldstein. LoRI: Reducing cross-task interference in multi-task low-rank adaptation. In Conference on Language Modeling, 2025a. URL https://openreview.net/forum?id=b8cW86QcOD.

247

Yifan Zhang, Ge Zhang, Yue Wu, Kangping Xu, and Quanquan Gu. Beyond Bradley–Terry models: A general preference model for language model alignment. In Proceedings of the 42nd International Conference on Machine Learning, volume 267, pages 76939–76965, 2025b.

248

Ziyu Zhao, Leilei Gan, Guoyin Wang, Wangchunshu Zhou, Hongxia Yang, Kun Kuang, and Fei Wu. LoraRetriever: Input-aware LoRA retrieval and composition for mixed tasks in the wild. In Findings of the Association for Computational Linguistics: ACL 2024, pages 4447–4462, 2024. DOI https://doi.org/10.18653/v1/2024.findings-acl.263.

249

Xun Zheng, Bryon Aragam, Pradeep Ravikumar, and Eric Xing. Dags with no tears: Continuous optimization for structure learning. Advances in Neural Information Processing Systems, 2018.

250

Yujia Zheng, Biwei Huang, Wei Chen, Joseph Ramsey, Mingming Gong, Ruichu Cai, Shohei Shimizu, Peter Spirtes, and Kun Zhang. Causal-learn: Causal discovery in python. Journal of Machine Learning Research, 25(60):1–8, 2024. URL http://jmlr.org/papers/v25/23-0237.html.