Exploiting Space Folding by Neural Networks
Michal Lewandowski, Raphael Pisoni, Bernhard Heinzl, Bernhard Alois Moser
摘要
Recent findings suggest that consecutive layers of neural networks with the ReLU activation function fold the input space during the learning process. While many works hint at this phenomenon, an approach to quantify the folding was only recently proposed by means of a space folding measure based on the Hamming distance in the ReLU activation space. Moreover, it has been observed that space folding values increase with network depth when the generalization error is low, but decrease when the error increases, thus underpinning that learned symmetries in the data manifold (visible in terms of space folds) contribute to the network's generalization capacity. Inspired by these findings, we propose a novel regularization scheme that enforces folding early during the training process. Further, we generalize the space folding measure to a wider class of activation functions through the introduction of equivalence classes of input data. We then analyze its mathematical and computational properties and propose an efficient sampling strategy for its implementation. Lastly, we outline the connection between learning with increased folding and contrastive learning, hinting that the former is a generalization of the latter. We underpin our claims with an experimental evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Functional vs. parametric equivalence of ReLU networksMary Phuong, Christoph H. LampertICLR 2020 · 被引用 53 次
- Empirical Studies on the Properties of Linear Regions in Deep Neural NetworksXiao Zhang, Dongrui WuICLR 2020 · 被引用 44 次
- Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural languageEghbal A. Hosseini, Evelina FedorenkoNeurIPS 2023 · 被引用 41 次
相关 Paper
- Task structure and nonlinearity jointly determine learned representational geometryMatteo Alleman, Jack W. Lindsey, Stefano FusiICLR 2024 · 被引用 11 次
- Hidden Symmetries of ReLU NetworksJ. Elisenda Grigsby, Kathryn Lindsey, David RolnickICML 2023 · 被引用 35 次
- Activation-Descent Regularization for Input Optimization of ReLU NetworksHongzhan Yu, Sicun GaoICML 2024 · 被引用 1 次
- Separation Power of Equivariant Neural NetworksMarco Pacini, Xiaowen Dong, Bruno Lepri, Gabriele SantinICLR 2025
- Learning-Based Image Registration With Meta-RegularizationEbrahim Al Safadi, Xubo SongCVPR 2021
