Low-dimensional topology of deep neural networks
Junyu Ren, Lek-Heng Lim
Abstract
We study layered models, including feedforward networks, ResNets, and transformers, by limiting each layer to a width of , i.e., as representation space. This allows us to track how a neural network changes low-dimensional topological invariants through its layers. Just about any topological structure may be simplified or even trivialized by simply increasing dimension; e.g., any knot is equivalent to an unknot in . By restricting to , we not only isolate the effects of activation and depth from that of width, we work in a space that lends itself to easy visualization. We focus on linking number here, deferring other invariants like link groups, Milnor's -invariants, knot types, ambient cobordisms, to a sequel. We provide full proofs and empirical experiments to justify the following insights: When measured by their power to effect changes in linking numbers, the layer-skipping feature in ResNets is as powerful as the attention mechanism in transformers; both ResNets and transformers are strictly more powerful than feedforward neural networks with monotonic activations, which are in turn more powerful than invertible and flow-based models; but replacing monotonic activation with a nonmonotonic one elevates a feedforward network into the same expressivity class as ResNets and transformers. These results suggest that low-dimensional topology can be a useful tool to guide designs of AI architectures. We also generalize our results from to arbitrary .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- E(n) Equivariant Graph Neural NetworksVictor Garcia Satorras, Emiel Hoogeboom, Max WellingICML 2021 · 1,432 citations
- Mean Flows for One-step Generative ModelingZhengyang Geng, Mingyang Deng, Xingjian Bai, Zico Kolter et al.NeurIPS 2025 · 628 citations
- Minimum Width for Universal ApproximationSejun Park, Chulhee Yun, Jaeho Lee, Jinwoo ShinICLR 2021 · 148 citations
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel et al.ICLR 2023 · 87 citations
- Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal ApproximationLi'ang Li, Yifei Duan, Guanghua Ji, Yongqiang CaiICML 2023 · 20 citations
Related papers
- Separation Power of Equivariant Neural NetworksMarco Pacini, Xiaowen Dong, Bruno Lepri, Gabriele SantinICLR 2025
- Transformative or Conservative? Conservation laws for ResNets and TransformersSibylle Marcotte, Rémi Gribonval, Gabriel PeyréICML 2025
- Rank Diminishing in Deep Neural NetworksRuili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao et al.NeurIPS 2022 · 64 citations
- Experimental Observations of the Topology of Convolutional Neural Network ActivationsEmilie Purvine, Davis Brown, Brett A. Jefferson, Cliff A. Joslyn et al.AAAI 2023 · 21 citations
- Alignment of CNN and Human Judgments of Geometric and Topological ConceptsNeha Upadhyay, Vijay Marupudi, Kamala Varma, Sashank VarmaAAAI 2025 · 2 citations
