Boosting the Transferability of Adversarial Samples via Attention
Weibin Wu, Yuxin Su, Xixian Chen, Shenglin Zhao, Irwin King, Michael R. Lyu, Yu-Wing Tai
Abstract
The widespread deployment of deep models necessitates the assessment of model vulnerability in practice, especially for safety-and security-sensitive domains such as autonomous driving and medical diagnosis. Transfer-based attacks against image classifiers thus elicit mounting interest, where attackers are required to craft adversarial images based on local proxy models without the feedback information from remote target ones. However, under such a challenging but practical setup, the synthesized adversarial samples often achieve limited success due to overfitting to the local model employed. In this work, we propose a novel mechanism to alleviate the overfitting issue. It computes model attention over extracted features to regularize the search of adversarial examples, which prioritizes the corruption of critical features that are likely to be adopted by diverse architectures. Consequently, it can promote the transferability of resultant adversarial instances. Extensive experiments on ImageNet classifiers confirm the effectiveness of our strategy and its superiority to state-of-the-art benchmarks in both white-box and black-box settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3952e86-c239-4313-a1f0-d804368a0afeCited by top-tier papers45
- On Success and Simplicity: A Second Look at Transferable Targeted AttacksZhengyu Zhao, Zhuoran Liu, Martha A. LarsonNeurIPS 2021 · 173 citations
- Towards Transferable Adversarial Attacks on Vision TransformersZhipeng Wei, Jingjing Chen, Micah Goldblum, Zuxuan Wu et al.AAAI 2022 · 156 citations
- Shadows can be Dangerous: Stealthy and Effective Physical-world Adversarial Attack by Natural PhenomenonYiqi Zhong, Xianming Liu, Deming Zhai, Junjun Jiang et al.CVPR 2022 · 148 citations
- Improving Adversarial Transferability via Neuron Attribution-based AttacksJianping Zhang, Weibin Wu, Jen-tse Huang, Yizhan Huang et al.CVPR 2022 · 140 citations
- Structure Invariant Transformation for better Adversarial TransferabilityXiaosen Wang, Zeliang Zhang, Jianping ZhangICCV 2023 · 130 citations
Builds on3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
Related papers
- Improving the Transferability of Adversarial Samples With Adversarial TransformationsWeibin Wu, Yuxin Su, Michael R. Lyu, Irwin KingCVPR 2021
- Feature Importance-aware Transferable Adversarial AttacksZhibo Wang, Hengchang Guo, Zhifei Zhang, Wenxin Liu et al.ICCV 2021 · 306 citations
- Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box DomainsQilong Zhang, Xiaodan Li, Yuefeng Chen, Jingkuan Song et al.ICLR 2022 · 85 citations
- Blurred-Dilated Method for Adversarial AttacksYang Deng, Weibin Wu, Jianping Zhang, Zibin ZhengNeurIPS 2023 · 10 citations
- Generating Transferable Adversarial Examples against Vision TransformersYuxuan Wang, Jiakai Wang, Zixin Yin, Ruihao Gong et al.ACM MM 2022 · 25 citations
