Exploiting Space Folding by Neural Networks
Michal Lewandowski, Raphael Pisoni, Bernhard Heinzl, Bernhard Alois Moser
Abstract
Recent findings suggest that consecutive layers of neural networks with the ReLU activation function fold the input space during the learning process. While many works hint at this phenomenon, an approach to quantify the folding was only recently proposed by means of a space folding measure based on the Hamming distance in the ReLU activation space. Moreover, it has been observed that space folding values increase with network depth when the generalization error is low, but decrease when the error increases, thus underpinning that learned symmetries in the data manifold (visible in terms of space folds) contribute to the network's generalization capacity. Inspired by these findings, we propose a novel regularization scheme that enforces folding early during the training process. Further, we generalize the space folding measure to a wider class of activation functions through the introduction of equivalence classes of input data. We then analyze its mathematical and computational properties and propose an efficient sampling strategy for its implementation. Lastly, we outline the connection between learning with increased folding and contrastive learning, hinting that the former is a generalization of the latter. We underpin our claims with an experimental evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fcef9c7-1f1b-4bc4-8201-b8664ce24791Builds on4
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Functional vs. parametric equivalence of ReLU networksMary Phuong, Christoph H. LampertICLR 2020 · 53 citations
- Empirical Studies on the Properties of Linear Regions in Deep Neural NetworksXiao Zhang, Dongrui WuICLR 2020 · 44 citations
- Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural languageEghbal A. Hosseini, Evelina FedorenkoNeurIPS 2023 · 41 citations
Related papers
- Task structure and nonlinearity jointly determine learned representational geometryMatteo Alleman, Jack W. Lindsey, Stefano FusiICLR 2024 · 11 citations
- Hidden Symmetries of ReLU NetworksJ. Elisenda Grigsby, Kathryn Lindsey, David RolnickICML 2023 · 35 citations
- Activation-Descent Regularization for Input Optimization of ReLU NetworksHongzhan Yu, Sicun GaoICML 2024 · 1 citation
- Separation Power of Equivariant Neural NetworksMarco Pacini, Xiaowen Dong, Bruno Lepri, Gabriele SantinICLR 2025
- Learning-Based Image Registration With Meta-RegularizationEbrahim Al Safadi, Xubo SongCVPR 2021
