Can Implicit Bias Imply Adversarial Robustness?
Hancheng Min, René Vidal
Abstract
The implicit bias of gradient-based training algorithms has been considered mostly beneficial as it leads to trained networks that often generalize well. However, Frei et al. (2023) show that such implicit bias can harm adversarial robustness. Specifically, they show that if the data consists of clusters with small inter-cluster correlation, a shallow (two-layer) ReLU network trained by gradient flow generalizes well, but it is not robust to adversarial attacks of small radius. Moreover, this phenomenon occurs despite the existence of a much more robust classifier that can be explicitly constructed from a shallow network. In this paper, we extend recent analyses of neuron alignment to show that a shallow network with a polynomial ReLU activation (pReLU) trained by gradient flow not only generalizes well but is also robust to adversarial attacks. Our results highlight the importance of the interplay between data structure and architecture design in the implicit bias and robustness of trained networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a36f3bb6-f52d-4d68-8f07-1eae0d224770Cited by top-tier papers3
- A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language ModelsMujtaba Hussain Mirza, Antonio D’Orazio, Odelia Melamed, Iacopo MasiCVPR 2026 · 2 citations
- Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural NetworksBinghui Li, Zhixuan Pan, Kaifeng Lyu, Jian LiICLR 2025
- AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example DefensesNicholas Carlini, Edoardo Debenedetti, Javier Rando, Milad Nasr et al.ICML 2025
Builds on23
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Randomized Smoothing of All Shapes and SizesGreg Yang, Tony Duan, J. Edward Hu, Hadi Salman et al.ICML 2020 · 237 citations
Related papers
- The Double-Edged Sword of Implicit Bias: Generalization vs. Robustness in ReLU NetworksSpencer Frei, Gal Vardi, Peter L. Bartlett, Nati SrebroNeurIPS 2023 · 25 citations
- Gradient Methods Provably Converge to Non-Robust NetworksGal Vardi, Gilad Yehudai, Ohad ShamirNeurIPS 2022 · 32 citations
- Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal DataYiwen Kou, Zixiang Chen, Quanquan GuNeurIPS 2023 · 24 citations
- Training invariances and the low-rank phenomenon: beyond linear networksThien Le, Stefanie JegelkaICLR 2022 · 39 citations
- On the Effective Number of Linear Regions in Shallow Univariate ReLU Networks: Convergence Guarantees and Implicit BiasItay Safran, Gal Vardi, Jason D. LeeNeurIPS 2022 · 26 citations
