Adversarial Examples Exist in Two-Layer ReLU Networks for Low Dimensional Linear Subspaces
Odelia Melamed, Gilad Yehudai, Gal Vardi
摘要
Despite a great deal of research, it is still not well-understood why trained neural networks are highly vulnerable to adversarial examples. In this work we focus on two-layer neural networks trained using data which lie on a low dimensional linear subspace. We show that standard gradient methods lead to non-robust neural networks, namely, networks which have large gradients in directions orthogonal to the data subspace, and are susceptible to small adversarial -perturbations in these directions. Moreover, we show that decreasing the initialization scale of the training algorithm, or adding regularization, can make the trained network more robust to adversarial perturbations orthogonal to the data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning a Neuron by a Shallow ReLU Network: Dynamics and Implicit Bias for Correlated InputsDmitry Chistikov, Matthias Englert, Ranko LazicNeurIPS 2023 · 被引用 22 次
- When Flatness Does (Not) Guarantee Adversarial RobustnessNils Philipp Walter, Linara Adilova, Jilles Vreeken, Michael KampICLR 2026 · 被引用 7 次
- From Curiosity to Caution: Mitigating Reward Hacking for Best-of- with PessimismZhuohao Yu, Steven Z. Wu, Adam BlockICLR 2026 · 被引用 3 次
- Feature compression is the root cause of adversarial fragility in neural networksJingchao Gao, Ziqing Lu, Raghu Mudumbai, Xiaodong Wu 等ICLR 2026 · 被引用 3 次
- H-SPLID: HSIC-based Saliency Preserving Latent Information DecompositionLukas Miklautz, Chengzhi Shi, Andrii Shkabrii, Theodoros-Thirimachos Davarakis 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper11
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain 等NeurIPS 2020 · 被引用 503 次
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 被引用 260 次
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir 等NeurIPS 2022 · 被引用 196 次
相关 Paper
- Gradient Methods Provably Converge to Non-Robust NetworksGal Vardi, Gilad Yehudai, Ohad ShamirNeurIPS 2022 · 被引用 32 次
- Hold me tight! Influence of discriminative features on deep network boundariesGuillermo Ortiz-Jiménez, Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2020 · 被引用 53 次
- Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured DataBinghui Li, Yuanzhi LiICLR 2025
- What Can the Neural Tangent Kernel Tell Us About Adversarial Robustness?Nikolaos Tsilivis, Julia KempeNeurIPS 2022 · 被引用 28 次
- How Benign is Benign Overfitting ?Amartya Sanyal, Puneet K. Dokania, Varun Kanade, Philip H. S. TorrICLR 2021 · 被引用 61 次
