Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
Sajad Movahedi, Antonio Orvieto, Seyed-Mohsen Moosavi-Dezfooli
Abstract
In this paper, we propose the geometric invariance hypothesis (GIH), which argues that the input space curvature of a neural network remains invariant under transformation in certain architecture-dependent directions during training. We investigate a simple, non-linear binary classification problem residing on a plane in a high dimensional space and observe that-unlike MLPs-ResNets fail to generalize depending on the orientation of the plane. Motivated by this example, we define a neural network's average geometry and average geometry evolution as compact architecture-dependent summaries of the model's input-output geometry and its evolution during training. By investigating the average geometry evolution at initialization, we discover that the geometry of a neural network evolves according to the data covariance projected onto its average geometry. This means that the geometry only changes in a subset of the input space when the average geometry is low-rank, such as in ResNets. This causes an architecture-dependent invariance property in the input space curvature, which we dub GIH. Finally, we present extensive experimental results to observe the consequences of GIH and how it relates to generalization in neural networks. The code for this paper is available at https://github.com/dr-faustus/GIH .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 276b3801-d475-42ab-bb29-a5dc0de6bb37Cited by top-tier papers2
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang et al.ACL 2026 · 5 citations
- On the Anisotropy of Score-Based Generative ModelsAndreas Floros, Seyed-Mohsen Moosavi-Dezfooli, Pier Luigi DragottiICML 2026 · 1 citation
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 235 citations
- Neural Networks as Kernel Learners: The Silent Alignment EffectAlexander B. Atanasov, Blake Bordelon, Cengiz PehlevanICLR 2022 · 110 citations
Related papers
- Data Representations' Study of Latent Image ManifoldsIlya Kaufman, Omri AzencotICML 2023 · 11 citations
- Neural (Tangent Kernel) CollapseMariia Seleznova, Dana Weitzner, Raja Giryes, Gitta Kutyniok et al.NeurIPS 2023 · 23 citations
- Task structure and nonlinearity jointly determine learned representational geometryMatteo Alleman, Jack W. Lindsey, Stefano FusiICLR 2024 · 11 citations
- Improving Neural Network Surface Processing with Principal CurvaturesJosquin Harrison, James Benn, Maxime SermesantNeurIPS 2024 · 3 citations
- A Classification of -invariant Shallow Neural NetworksDevanshu Agrawal, James OstrowskiNeurIPS 2022 · 11 citations
