Adversarial Examples Exist in Two-Layer ReLU Networks for Low Dimensional Linear Subspaces
Odelia Melamed, Gilad Yehudai, Gal Vardi
Abstract
Despite a great deal of research, it is still not well-understood why trained neural networks are highly vulnerable to adversarial examples. In this work we focus on two-layer neural networks trained using data which lie on a low dimensional linear subspace. We show that standard gradient methods lead to non-robust neural networks, namely, networks which have large gradients in directions orthogonal to the data subspace, and are susceptible to small adversarial -perturbations in these directions. Moreover, we show that decreasing the initialization scale of the training algorithm, or adding regularization, can make the trained network more robust to adversarial perturbations orthogonal to the data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Learning a Neuron by a Shallow ReLU Network: Dynamics and Implicit Bias for Correlated InputsDmitry Chistikov, Matthias Englert, Ranko LazicNeurIPS 2023 · 22 citations
- When Flatness Does (Not) Guarantee Adversarial RobustnessNils Philipp Walter, Linara Adilova, Jilles Vreeken, Michael KampICLR 2026 · 7 citations
- From Curiosity to Caution: Mitigating Reward Hacking for Best-of- with PessimismZhuohao Yu, Steven Z. Wu, Adam BlockICLR 2026 · 3 citations
- Feature compression is the root cause of adversarial fragility in neural networksJingchao Gao, Ziqing Lu, Raghu Mudumbai, Xiaodong Wu et al.ICLR 2026 · 3 citations
- H-SPLID: HSIC-based Saliency Preserving Latent Information DecompositionLukas Miklautz, Chengzhi Shi, Andrii Shkabrii, Theodoros-Thirimachos Davarakis et al.NeurIPS 2025 · 2 citations
Builds on11
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 260 citations
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir et al.NeurIPS 2022 · 196 citations
Related papers
- Gradient Methods Provably Converge to Non-Robust NetworksGal Vardi, Gilad Yehudai, Ohad ShamirNeurIPS 2022 · 32 citations
- Hold me tight! Influence of discriminative features on deep network boundariesGuillermo Ortiz-Jiménez, Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2020 · 53 citations
- Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured DataBinghui Li, Yuanzhi LiICLR 2025
- What Can the Neural Tangent Kernel Tell Us About Adversarial Robustness?Nikolaos Tsilivis, Julia KempeNeurIPS 2022 · 28 citations
- How Benign is Benign Overfitting ?Amartya Sanyal, Puneet K. Dokania, Varun Kanade, Philip H. S. TorrICLR 2021 · 61 citations
