references

References

Bibliography

1

Dana Angluin. Learning regular sets from queries and counterexamples. Information and Computation, 75(2):87–106, 1987. DOI https://doi.org/10.1016/0890-5401(87)90052-6. URL https://doi.org/10.1016/0890-5401(87)90052-6.

2

Anthropic. Formalizing fermat’s last theorem in Lean: A timeline and selected excerpts from Claude’s reasoning, 2026. URL https://www-cdn.anthropic.com/9e431dff043da6538d99d6c2d231b670aa3da263.pdf. Accompanying Lean development available at https://github.com/anthropics/fermats-last-theorem.

3

Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15619–15629, 2023. URL https://openaccess.thecvf.com/content/CVPR2023/html/Assran_Self-Supervised_Learning_From_Images_With_a_Joint-Embedding_Predictive_Architecture_CVPR_2023_paper.html.

4

Mido Assran et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985, 2025. DOI https://doi.org/10.48550/arXiv.2506.09985. URL https://arxiv.org/abs/2506.09985. Revised March 22, 2026.

5

Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire. The nonstochastic multiarmed bandit problem. SIAM Journal on Computing, 32(1):48–77, 2002. DOI https://doi.org/10.1137/S0097539701398375.

6

Michael Barr and Charles Wells. Category Theory for Computing Science. Centre de Recherches Mathématiques, 3 edition, 1999. URL https://www.tac.mta.ca/tac/reprints/articles/22/tr22.pdf. Reprinted in Theory and Applications of Categories, No. 22 (2012).

7

Dimitri P. Bertsekas and John N. Tsitsiklis. Parallel and Distributed Computation: Numerical Methods. Prentice-Hall, Englewood Cliffs, NJ, 1989. URL https://www.mit.edu/ dimitrib/pdc.html. Republished by Athena Scientific, 1997.

8

David Blackwell. An analog of the minimax theorem for vector payoffs. Pacific Journal of Mathematics, 6(1):1–8, 1956. DOI https://doi.org/10.2140/pjm.1956.6.1.

9

Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K. Warmuth. Learnability and the Vapnik–Chervonenkis dimension. Journal of the ACM, 36(4):929–965, 1989. DOI https://doi.org/10.1145/76359.76371.

10

John C. Bowler, Dua B. Azhar, Cambria M. Jensen, Hyun-Woo Lee, and James G. Heys. Structured experience shapes strategy learning and neural dynamics in the medial entorhinal cortex. Nature Neuroscience, September 2026. DOI https://doi.org/10.1038/s41593-026-02409-7. URL https://www.nature.com/articles/s41593-026-02409-7. Published online September 3, 2026.

11

Gunnar Carlsson and Jun Yu. A prime decomposition of probabilistic automata. arXiv preprint arXiv:1503.01502, 2015. DOI https://doi.org/10.48550/arXiv.1503.01502. URL https://arxiv.org/abs/1503.01502.

12

J. R. B. Cockett and G. S. H. Cruttwell. Differential structure, tangent structure, and SDG. Applied Categorical Structures, 22:331–417, 2014. DOI https://doi.org/10.1007/s10485-013-9312-0.

13

Robin Cockett, Jean-Simon Pacaud Lemay, and Rory B. B. Lucyshyn-Wright. Tangent categories from the coalgebras of differential categories. In 28th EACSL Annual Conference on Computer Science Logic (CSL 2020), volume 152 of Leibniz International Proceedings in Informatics (LIPIcs), pages 17:1–17:17. Schloss Dagstuhl–Leibniz-Zentrum für Informatik, 2020. DOI https://doi.org/10.4230/LIPIcs.CSL.2020.17.

14

Daice Labs. Category theory as the language of composable AI. Daice Labs research synthesis, 2026. URL https://daicelabs.com/research/category-theory. Category Theory, Machine Learning, and AI Safety.

15

Gary L. Drescher. Made-Up Minds: A Constructivist Approach to Artificial Intelligence. The MIT Press, Cambridge, MA, 1991. ISBN 9780262041201. URL https://mitpress.mit.edu/9780262517089/made-up-minds/. Paperback edition published 2003.

16

Charles Ehresmann. Esquisses et types des structures algébriques. Buletinul Institutului Politehnic din Iaşi, 14:1–14, 1968.

17

Abraham D. Flaxman, Adam Tauman Kalai, and H. Brendan McMahan. Online convex optimization in the bandit setting: Gradient descent without a gradient. In Proceedings of the Sixteenth Annual ACM-SIAM Symposium on Discrete Algorithms, pages 385–394. Society for Industrial and Applied Mathematics, 2005. URL https://arxiv.org/abs/cs/0408007.

18

Brendan Fong, David I. Spivak, and Rémy Tuyéras. Backprop as functor: A compositional perspective on supervised learning. In Proceedings of the 34th Annual ACM/IEEE Symposium on Logic in Computer Science, LICS 2019, pages 1–13, 2019. DOI https://doi.org/10.1109/LICS.2019.8785665. URL https://arxiv.org/abs/1711.10455.

19

E. Mark Gold. Limiting recursion. The Journal of Symbolic Logic, 30(1):28–48, 1965. DOI https://doi.org/10.2307/2270580. URL https://doi.org/10.2307/2270580.

20

E. Mark Gold. Language identification in the limit. Information and Control, 10(5):447–474, 1967. DOI https://doi.org/10.1016/S0019-9958(67)91165-5. URL https://doi.org/10.1016/S0019-9958(67)91165-5.

21

Ryan Greenblatt, Ajeya Cotra, and Hjalmar Wijk. Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI/Hugging Face hacking incident. METR research report, August 2026. URL https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/. Published August 26, 2026.

22

David Ha and Jürgen Schmidhuber. World models. arXiv preprint arXiv:1803.10122, 2018. DOI https://doi.org/10.48550/arXiv.1803.10122. URL https://arxiv.org/abs/1803.10122.

23

Redi Haderi, Cihan Okay, and Walker H. Stern. The operadic theory of convexity. CoRR, abs/2403.18102, 2024. URL https://arxiv.org/abs/2403.18102.

24

Elad Hazan. Introduction to online convex optimization. Foundations and Trends in Optimization, 2(3–4):157–325, 2016. DOI https://doi.org/10.1561/2400000013.

25

Elad Hazan. Introduction to Online Convex Optimization. arXiv, 2 edition, 2023. URL https://arxiv.org/abs/1909.05207. arXiv:1909.05207v3.

26

Alec Helbling, Andrey Bryutkin, Mauro Martino, Duen Horng Chau, Nima Dehmamy, and Hendrik Strobelt. Flow reasoning models: Turning flows into efficient recurrent reasoners. CoRR, abs/2606.29150, 2026. DOI https://doi.org/10.48550/arXiv.2606.29150. URL https://arxiv.org/abs/2606.29150. Version 3.

27

William James. The Principles of Psychology, volume 1. Henry Holt and Company, New York, 1890. URL https://psychclassics.yorku.ca/James/Principles/index.htm.

28

Daniel M. Kan. Adjoint functors. Transactions of the American Mathematical Society, 87(2):294–329, 1958. DOI https://doi.org/10.2307/1993102.

29

G. M. Kelly. Basic Concepts of Enriched Category Theory, volume 64 of London Mathematical Society Lecture Note Series. Cambridge University Press, 1982. URL https://www.tac.mta.ca/tac/reprints/articles/10/tr10abs.html. Reprinted in Theory and Applications of Categories, No. 10 (2005).

30

Anders Kock. Synthetic Differential Geometry. Cambridge University Press, 2 edition, 2006. ISBN 978-0-521-68738-6. DOI https://doi.org/10.1017/CBO9780511550812. URL https://www.cambridge.org/core/books/synthetic-differential-geometry/4933FF5EDF69BD13E8D8FCA346BB1C06.

31

Kenneth Krohn and John Rhodes. Algebraic theory of machines. I. prime decomposition theorem for finite semigroups and machines. Transactions of the American Mathematical Society, 116:450–464, 1965. DOI https://doi.org/10.1090/S0002-9947-1965-0188316-1. URL https://www.ams.org/journals/tran/1965-116-00/S0002-9947-1965-0188316-1/.

32

Stephen Lack. Composing PROPs. Theory and Applications of Categories, 13(9):147–163, 2004. URL https://tac.mta.ca/tac/volumes/13/9/13-09abs.html.

33

Brenden M. Lake and Marco Baroni. Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 2873–2882, 2018. URL https://proceedings.mlr.press/v80/lake18a.html.

34

Brenden M. Lake and Marco Baroni. Human-like systematic generalization through a meta-learning neural network. Nature, 623(7985):115–121, 2023. DOI https://doi.org/10.1038/s41586-023-06668-3.

35

Brenden M. Lake, Tomer D. Ullman, Joshua B. Tenenbaum, and Samuel J. Gershman. Building machines that learn and think like people. Behavioral and Brain Sciences, 40:e253, 2017. DOI https://doi.org/10.1017/S0140525X16001837.

36

Tor Lattimore and Csaba Szepesvári. Bandit Algorithms. Cambridge University Press, 2020. ISBN 9781108486828. DOI https://doi.org/10.1017/9781108571401. URL https://www.cambridge.org/core/books/bandit-algorithms/8E39FD004E6CE036680F90DD0C6F09FC.

37

F. William Lawvere. Functorial semantics of algebraic theories. Proceedings of the National Academy of Sciences, 50(5):869–872, 1963. DOI https://doi.org/10.1073/pnas.50.5.869. URL https://pmc.ncbi.nlm.nih.gov/articles/PMC221940/.

38

Yann LeCun. A path towards autonomous machine intelligence. OpenReview position paper, version 0.9.2, June 2022. URL https://openreview.net/forum?id=BZ5a1r-kVsf.

39

Michael L. Littman, Richard S. Sutton, and Satinder Singh. Predictive representations of state. In Advances in Neural Information Processing Systems 14, pages 1555–1561, 2001. URL https://proceedings.neurips.cc/paper/2001/hash/1e4d36177d71bbb3558e43af9577d70e-Abstract.html.

40

Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. Transformers learn shortcuts to automata. In International Conference on Learning Representations, 2023. DOI https://doi.org/10.48550/arXiv.2210.10749. URL https://openreview.net/forum?id=De4FYqjFueZ. Oral presentation.

41

Saunders Mac Lane. Categories for the Working Mathematician, volume 5 of Graduate Texts in Mathematics. Springer, New York, 1971. ISBN 978-1-4612-9839-7. DOI https://doi.org/10.1007/978-1-4612-9839-7. URL https://link.springer.com/book/10.1007/978-1-4612-9839-7.

42

Sridhar Mahadevan. Universal decision models. CoRR, abs/2110.15431, 2021. URL https://arxiv.org/abs/2110.15431.

43

Sridhar Mahadevan. Universal imitation games. arXiv preprint arXiv:2405.01540, 2024. URL https://arxiv.org/abs/2405.01540.

44

Sridhar Mahadevan. Categories for Artificial General Intelligence. Categorical AI Book Project, 2026a. Book manuscript.

45

Sridhar Mahadevan. Infinitesimal Creativity: Learning by Double Involution. The Infinitesimal Creativity Book Project, 2026b.

46

Sridhar Mahadevan. Machine Learning from Enforcing Compositionality. The LINCS Book Project, 2026c.

47

Sridhar Mahadevan. ODYSSEY: Constructing verifiable local truth-preserving foundation models, 2026d. URL https://arxiv.org/abs/2606.27593.

48

Sridhar Mahadevan. PROMETHEUS: Automating deep causal research integrating text, data and models, 2026e. URL https://arxiv.org/abs/2605.12835.

49

H. Brendan McMahan. Follow-the-regularized-leader and mirror descent: Equivalence theorems and L1 regularization. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, volume 15 of Proceedings of Machine Learning Research, pages 525–533. PMLR, 2011. URL https://proceedings.mlr.press/v15/mcmahan11b.html.

50

METR. Task-completion time horizons of frontier AI models. Online research dashboard, 2026. URL https://metr.org/time-horizons/. Last updated May 8, 2026.

51

Andrea Moro. Impossible Languages. The MIT Press, Cambridge, MA, 2023. ISBN 9780262549233. URL https://mitpress.mit.edu/9780262549233/impossible-languages/. Paperback edition, published September 19, 2023.

52

OpenAI. Path to Astra: Critical capabilities and frontier safeguards, September 2026a. URL https://openai.com/index/path-to-astra/. Published September 1, 2026.

53

OpenAI. Separating signal from noise in coding evaluations. Research report, July 2026b. URL https://openai.com/index/separating-signal-from-noise-coding-evaluations/. Published July 8, 2026.

54

Judea Pearl. Causality: Models, Reasoning, and Inference. Cambridge University Press, 2 edition, 2009. ISBN 978-0-521-89560-6. DOI https://doi.org/10.1017/CBO9780511803161. URL https://www.cambridge.org/core/books/causality/B0046844FAE10CBF274D4ACBDAEB5F5B.

55

Jonas Peters, Dominik Janzing, and Bernhard Schölkopf. Elements of Causal Inference: Foundations and Learning Algorithms. Adaptive Computation and Machine Learning. MIT Press, 2017. ISBN 978-0-262-03731-0. URL https://mitpress.mit.edu/9780262037310/elements-of-causal-inference/.

56

Jean Piaget. The Origins of Intelligence in Children. International Universities Press, New York, 1952. URL https://iucat.iu.edu/iub/3827762.

57

Top Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller, Marta Kryven, and Kevin Ellis. PoE-World: Compositional world modeling with products of programmatic experts. In Advances in Neural Information Processing Systems, volume 38, 2025. DOI https://doi.org/10.52202/085713-0896. URL https://proceedings.neurips.cc/paper_files/paper/2025/hash/262dd62fd1bbb30d6a6b4d578f5e65ff-Abstract-Conference.html. Code: https://github.com/topwasu/poe-world.

58

Boris Plotkin and Tatjana Plotkin. Decompositions and complexity of linear automata. arXiv preprint arXiv:1506.06017, 2015. DOI https://doi.org/10.48550/arXiv.1506.06017. URL https://arxiv.org/abs/1506.06017.

59

Martin L. Puterman. Markov Decision Processes: Discrete Stochastic Dynamic Programming. John Wiley & Sons, 1994. ISBN 978-0-471-61977-2. DOI https://doi.org/10.1002/9780470316887. URL https://onlinelibrary.wiley.com/doi/book/10.1002/9780470316887.

60

Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, S. M. Ali Eslami, and Matthew Botvinick. Machine theory of mind. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 4218–4227, 2018. URL https://proceedings.mlr.press/v80/rabinowitz18a.html.

61

Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar. On the convergence of Adam and beyond. CoRR, abs/1904.09237, 2019. URL https://arxiv.org/abs/1904.09237.

62

Birgit Richter. From Categories to Homotopy Theory, volume 188 of Cambridge Studies in Advanced Mathematics. Cambridge University Press, 2020. DOI https://doi.org/10.1017/9781108855891. URL https://www.cambridge.org/core/books/from-categories-to-homotopy-theory/A109E2C4B720337DE19A15EB4FA8C9A6.

63

Emily Riehl. Categorical Homotopy Theory, volume 24 of New Mathematical Monographs. Cambridge University Press, 2014. DOI https://doi.org/10.1017/CBO9781107261457.

64

Emily Riehl. Category Theory in Context. Dover Publications, 2016. URL https://emilyriehl.github.io/books/.

65

Emily Riehl and Dominic Verity. Elements of \(\infty \)-Category Theory. Cambridge University Press, 2022. DOI https://doi.org/10.1017/9781108936880. URL https://www.cambridge.org/core/books/elements-of-infinity-category-theory/DAC48C449AB8C2C1B1E528A49D27FC6D.

66

Ronald L. Rivest and Robert E. Schapire. Diversity-based inference of finite automata. Journal of the ACM, 41(3):555–589, 1994. DOI https://doi.org/10.1145/176584.176589. URL https://people.csail.mit.edu/rivest/pubs/RS94a.pdf.

67

R. Tyrrell Rockafellar. Monotone operators and the proximal point algorithm. SIAM Journal on Control and Optimization, 14(5):877–898, 1976. DOI https://doi.org/10.1137/0314056.

68

R. Tyrrell Rockafellar. Proto-differentiability of set-valued mappings and its applications in optimization. Annales de l’Institut Henri Poincaré C, Analyse non linéaire, 6:449–482, 1989. DOI https://doi.org/10.1016/S0294-1449(17)30034-3.

69

Alessandro Ronca, Nadezda Alexandrovna Knorozova, and Giuseppe De Giacomo. Automata cascades: Expressivity and sample complexity. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence, volume 37, pages 9588–9595, 2023. DOI https://doi.org/10.1609/aaai.v37i8.26147. URL https://ojs.aaai.org/index.php/AAAI/article/view/26147. Issue 8.

70

Jan J. M. M. Rutten. Universal coalgebra: A theory of systems. Theoretical Computer Science, 249(1):3–80, 2000. DOI https://doi.org/10.1016/S0304-3975(00)00056-6. URL https://doi.org/10.1016/S0304-3975(00)00056-6.

71

Jenny R. Saffran, Richard N. Aslin, and Elissa L. Newport. Statistical learning by 8-month-old infants. Science, 274(5294):1926–1928, 1996. DOI https://doi.org/10.1126/science.274.5294.1926.

72

Jerome H. Saltzer and Michael D. Schroeder. The protection of information in computer systems. Proceedings of the IEEE, 63(9):1278–1308, 1975. DOI https://doi.org/10.1109/PROC.1975.9939.

73

Patrick Schultz and David I. Spivak. Temporal Type Theory: A Topos-Theoretic Approach to Systems and Behavior. arXiv, 2017. DOI https://doi.org/10.48550/arXiv.1710.10258. URL https://arxiv.org/abs/1710.10258. arXiv:1710.10258v3, 224 pages.

74

Shai Shalev-Shwartz. Online learning and online convex optimization. Foundations and Trends in Machine Learning, 4(2):107–194, 2012. DOI https://doi.org/10.1561/2200000018.

75

Elizabeth S. Spelke. What Babies Know: Core Knowledge and Composition, volume 1. Oxford University Press, New York, 2022. ISBN 9780190618247. DOI https://doi.org/10.1093/oso/9780190618247.001.0001. URL https://academic.oup.com/book/43912.

76

Elizabeth S. Spelke. Précis of What Babies Know. Behavioral and Brain Sciences, 47:e120, 2023. DOI https://doi.org/10.1017/S0140525X23002443. URL https://pubmed.ncbi.nlm.nih.gov/37248696/.

77

Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, 2 edition, 2018. URL http://incompleteideas.net/book/the-book-2nd.html.

78

Richard S. Sutton, Doina Precup, and Satinder Singh. Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning. Artificial Intelligence, 112(1–2):181–211, 1999. DOI https://doi.org/10.1016/S0004-3702(99)00052-1. URL https://www.sciencedirect.com/science/article/pii/S0004370299000521.

79

Denis Thérien. Two-sided wreath product of categories. Journal of Pure and Applied Algebra, 74(3):307–315, 1991. DOI https://doi.org/10.1016/0022-4049(91)90119-M.

80

John N. Tsitsiklis. Asynchronous stochastic approximation and Q-learning. Machine Learning, 16(3):185–202, 1994. DOI https://doi.org/10.1023/A:1022689125041. URL https://www.mit.edu/ jnt/Papers/J052-94-jnt-q.pdf.

81

Leslie G. Valiant. A theory of the learnable. Communications of the ACM, 27(11):1134–1142, 1984. DOI https://doi.org/10.1145/1968.1972.

82

Leslie G. Valiant. Evolvability. Journal of the ACM, 56(1):1–21, 2009. DOI https://doi.org/10.1145/1462153.1462156. URL https://dl.acm.org/doi/10.1145/1462153.1462156.

83

V. N. Vapnik and A. Ya. Chervonenkis. On the uniform convergence of relative frequencies of events to their probabilities. Theory of Probability and Its Applications, 16(2):264–280, 1971. DOI https://doi.org/10.1137/1116025.

84

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, pages 5998–6008, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html.

85

Steven Weinberg. A model of leptons. Physical Review Letters, 19(21):1264–1266, 1967. DOI https://doi.org/10.1103/PhysRevLett.19.1264. URL https://journals.aps.org/prl/abstract/10.1103/PhysRevLett.19.1264.

86

Charles Wells. A Krohn–Rhodes theorem for categories. Journal of Algebra, 64(1):37–45, 1980. DOI https://doi.org/10.1016/0021-8693(80)90130-1.

87

Charles Wells. Wreath product decomposition of categories. II. Acta Scientiarum Mathematicarum, 52:321–324, 1988. URL https://acta.bibl.u-szeged.hu/15191/1/math_052_fasc_003_004_321-324.pdf.

88

Eugene P. Wigner. On unitary representations of the inhomogeneous lorentz group. Annals of Mathematics, 40(1):149–204, 1939. DOI https://doi.org/10.2307/1968551. URL https://cds.cern.ch/record/405887/.

89

Hans S. Witsenhausen. On information structures, feedback and causality. SIAM Journal on Control, 9(2):149–160, 1971. DOI https://doi.org/10.1137/0309013.

90

Hans S. Witsenhausen. The intrinsic model for discrete stochastic control: Some open problems. In Control Theory, Numerical Methods and Computer Systems Modelling, volume 107 of Lecture Notes in Economics and Mathematical Systems, pages 322–335. Springer, Berlin, 1975.

91

Chen Ning Yang and Robert L. Mills. Conservation of isotopic spin and isotopic gauge invariance. Physical Review, 96(1):191–195, 1954. DOI https://doi.org/10.1103/PhysRev.96.191.

92

Hongxin Zhang, Zeyuan Wang, Qiushi Lyu, Zheyuan Zhang, Sunli Chen, Tianmin Shu, Behzad Dariush, Kwonjoon Lee, Yilun Du, and Chuang Gan. COMBO: Compositional world models for embodied multi-agent cooperation. In International Conference on Learning Representations, 2025. URL https://proceedings.iclr.cc/paper_files/paper/2025/hash/7d03c6bf9f07acb4038eea96c63db52d-Abstract-Conference.html. Project page: https://umass-embodied-agi.github.io/COMBO/.

93

Siyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li, Dit-Yan Yeung, and Chuang Gan. RoboDreamer: Learning compositional world models for robot imagination. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 61885–61896. PMLR, 2024. URL https://proceedings.mlr.press/v235/zhou24f.html. Project page: https://robovideo.github.io; code: https://github.com/rainbow979/robodreamer.