Uniform Convergence, Adversarial Spheres and a Simple Remedy
Gregor Bachmann, Seyed-Mohsen Moosavi-Dezfooli, Thomas Hofmann
Abstract
Previous work has cast doubt on the general framework of uniform convergence and its ability to explain generalization in neural networks. By considering a specific dataset, it was observed that a neural network completely misclassifies a projection of the training data (adversarial set), rendering any existing generalization bound based on uniform convergence vacuous. We provide an extensive theoretical investigation of the previously studied data setting through the lens of infinitely-wide models. We prove that the Neural Tangent Kernel (NTK) also suffers from the same phenomenon and we uncover its origin. We highlight the important role of the output bias and show theoretically as well as empirically how a sensible choice completely mitigates the problem. We identify sharp phase transitions in the accuracy on the adversarial set and study its dependency on the training sample size. As a result, we are able to characterize critical sample sizes beyond which the effect disappears. Moreover, we study decompositions of a neural network into a clean and noisy part by considering its canonical decomposition into its different eigenfunctions and show empirically that for too small bias the adversarial phenomenon still persists.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cabde88c-6a9a-4da7-b78b-c90bcf9570efCited by top-tier papers6
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 272 citations
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 62 citations
- Integral Probability Metrics PAC-Bayes BoundsRon Amit, Baruch Epstein, Shay Moran, Ron MeirNeurIPS 2022 · 25 citations
- From Tempered to Benign Overfitting in ReLU Neural NetworksGuy Kornowski, Gilad Yehudai, Ohad ShamirNeurIPS 2023 · 18 citations
- Generalization Through the Lens of Leave-One-Out ErrorGregor Bachmann, Thomas Hofmann, Aurélien LucchiICLR 2022 · 9 citations
Builds on4
- Frequency Bias in Neural Networks for Input of Non-Uniform DensityRonen Basri, Meirav Galun, Amnon Geifman, David W. Jacobs et al.ICML 2020 · 229 citations
- On the Similarity between the Laplace and Neural Tangent KernelsAmnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun et al.NeurIPS 2020 · 118 citations
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 107 citations
- In Defense of Uniform Convergence: Generalization via Derandomization with an Application to Interpolating PredictorsJeffrey Negrea, Gintare Karolina Dziugaite, Daniel M. RoyICML 2020 · 66 citations
Related papers
- Disentangling Trainability and Generalization in Deep Neural NetworksLechao Xiao, Jeffrey Pennington, Samuel Stern SchoenholzICML 2020 · 91 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
- Tuning Frequency Bias in Neural Network Training with Nonuniform DataAnnan Yu, Yunan Yang, Alex TownsendICLR 2023 · 2 citations
- The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich RegimesAlexander B. Atanasov, Blake Bordelon, Sabarish Sainathan, Cengiz PehlevanICLR 2023 · 4 citations
- Theoretical Analysis of Robust Overfitting for Wide DNNs: An NTK ApproachShaopeng Fu, Di WangICLR 2024 · 9 citations
