Why Does Little Robustness Help? A Further Step Towards Understanding Adversarial Transferability
Yechao Zhang, Shengshan Hu, Leo Yu Zhang, Junyu Shi, Minghui Li, Xiaogeng Liu, Wei Wan, Hai Jin
Abstract
Adversarial examples for deep neural networks (DNNs) are transferable: examples that successfully fool one white-box surrogate model can also deceive other black-box models with different architectures. Although a bunch of empirical studies have provided guidance on generating highly transferable adversarial examples, many of these findings fail to be well explained and even lead to confusing or inconsistent advice for practical use.In this paper, we take a further step towards understanding adversarial transferability, with a particular focus on surrogate aspects. Starting from the intriguing "little robustness" phenomenon, where models adversarially trained with mildly perturbed adversarial samples can serve as better surrogates for transfer attacks, we attribute it to a trade-off between two dominant factors: model smoothness and gradient similarity. Our research focuses on their joint effects on transferability, rather than demonstrating the separate relationships alone. Through a combination of theoretical and empirical analyses, we hypothesize that the data distribution shift induced by off-manifold samples in adversarial training is the reason that impairs gradient similarity.Building on these insights, we further explore the impacts of prevalent data augmentation and gradient regularization on transferability and analyze how the trade-off manifests in various training methods, thus building a comprehensive blueprint for the regulation mechanisms behind transferability. Finally, we provide a general route for constructing superior surrogates to boost transferability, which optimizes both model smoothness and gradient similarity simultaneously, e.g., the combination of input gradient regularization and sharpness-aware minimization (SAM), validated by extensive experiments. In summary, we call for attention to the united impacts of these two factors for launching effective transfer attacks, rather than optimizing one while ignoring the other, and emphasize the crucial role of manipulating surrogate models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ef3d85f-9510-4039-b68f-5bb127c8f0ccCited by top-tier papers16
- Downstream-agnostic Adversarial ExamplesZiqi Zhou, Shengshan Hu, Ruizhi Zhao, Qian Wang et al.ICCV 2023 · 45 citations
- DarkSAM: Fooling Segment Anything Model to Segment NothingZiqi Zhou, Yufei Song, Minghui Li, Shengshan Hu et al.NeurIPS 2024 · 44 citations
- Securely Fine-tuning Pre-trained Encoders Against Adversarial ExamplesZiqi Zhou, Minghui Li, Wei Liu, Shengshan Hu et al.S&P 2024 · 23 citations
- NumbOD: A Spatial-Frequency Fusion Attack Against Object DetectorsZiqi Zhou, Bowen Li, Yufei Song, Zhifei Yu et al.AAAI 2025 · 20 citations
- Enhancing Adversarial Transferability with Adversarial Weight TuningJiahao Chen, Zhou Feng, Rui Zeng, Yuwen Pu et al.AAAI 2025 · 11 citations
Builds on50
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
Related papers
- LRS: Enhancing Adversarial Transferability through Lipschitz Regularized SurrogateTao Wu, Tie Luo, Donald C. Wunsch IIAAAI 2024 · 11 citations
- StyLess: Boosting the Transferability of Adversarial ExamplesKaisheng Liang, Bin XiaoCVPR 2023
- Transferability Bound Theory: Exploring Relationship between Adversarial Transferability and FlatnessMingyuan Fan, Xiaodan Li, Cen Chen, Wenmeng Zhou et al.NeurIPS 2024 · 13 citations
- Training Meta-Surrogate Model for Transferable Adversarial AttackYunxiao Qin, Yuanhao Xiong, Jinfeng Yi, Cho-Jui HsiehAAAI 2023 · 31 citations
- Boosting Black-Box Attack with Partially Transferred Conditional Adversarial DistributionYan Feng, Baoyuan Wu, Yanbo Fan, Li Liu et al.CVPR 2022 · 34 citations
