lin-0165
Further reading
Kan extensions and their universal properties can be approached through Mac Lane [ 1998 ] or Riehl [ 2017 ] ; their use in probability is developed by van Belle [ 2024 ] . The transformer reference point is the original attention architecture [ Vaswani et al. , 2017 ] . Kimi K3 provides a current primary account of hybrid attention, latent expert routing, depth-wise residual retrieval, quantization-aware post-training, and large-scale model execution [ Kimi Team , 2026 ] . For an accessible production-oriented overview, see Dickson [ 2026 ] ; the architectural claims in this chapter follow the primary report. Diffusion models [ Ho et al. , 2020 ] , diffusion language models [ Li et al. , 2022 ] , and flow matching [ Lipman et al. , 2023 ] provide neighboring examples in which transport and alternative computational routes matter.
For the topology of local commutators, May [ 1992 ] supplies the simplicial background and Jiang et al. [ 2011 ] gives a statistical Hodge decomposition on graphs. The KET paper [ Mahadevan , 2026g ] develops the broader categorical architecture from which the chapter’s post-training LINCS audit is extracted. These readings help separate three claims that are easy to conflate: a Kan-extension specification, an implemented neural architecture, and a successful blockwise repair experiment.