When Flatness Does (Not) Guarantee Adversarial Robustness
Nils Philipp Walter, Linara Adilova, Jilles Vreeken, Michael Kamp
Abstract
Despite their empirical success, neural networks remain vulnerable to small, adversarial perturbations. A longstanding hypothesis suggests that flat minima, regions of low curvature in the loss landscape, offer increased robustness. While intuitive, this connection has remained largely informal and incomplete. By rigorously formalizing the relationship, we show this intuition is only partially correct: flatness implies local but not global adversarial robustness. To arrive at this result, we first derive a closed-form expression for relative flatness in the penultimate layer, and then show we can use this to constrain the variation of the loss in input space. This allows us to formally analyze the adversarial robustness of the entire network. We then show that to maintain robustness beyond a local neighborhood, the loss needs to curve sharply away from the data manifold. We validate our theoretical predictions empirically across architectures and datasets, uncovering the geometric structure that governs adversarial vulnerability, and linking flatness to model confidence: adversarial examples often lie in large, flat regions where the model is confidently wrong. Our results challenge simplified views of flatness and provide a nuanced understanding of its role in robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69fe1892-07fe-4225-9e05-8113f798b6a0Cited by top-tier papers3
- Unveiling the Basin-Like Loss Landscape in Large Language ModelsHuanran Chen, Zeming Wei, Yao Huang, Yichi Zhang et al.ICLR 2026 · 14 citations
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 12 citations
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via GrokkingTing Han, Linara Adilova, Henning Petzka, Jens Kleesiek et al.NeurIPS 2025 · 9 citations
Builds on19
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- Relating Adversarially Robust Generalization to Flat MinimaDavid Stutz, Matthias Hein, Bernt SchieleICCV 2021 · 80 citations
- Optimizing Mode Connectivity via Neuron AlignmentN. Joseph Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk et al.NeurIPS 2020 · 104 citations
- On the Local Complexity of Linear Regions in Deep ReLU NetworksNiket Patel, Guido MontúfarICML 2025
- On Linear Stability of SGD and Input-Smoothness of Neural NetworksChao Ma, Lexing YingNeurIPS 2021 · 73 citations
- Prompting Adversarial Transferability via Path Flatness AttackZeze Tao, Jinjia Peng, Huibing WangAAAI 2026
