Improving Adversarial Transferability via Neuron Attribution-based Attacks
Jianping Zhang, Weibin Wu, Jen-tse Huang, Yizhan Huang, Wenxuan Wang, Yuxin Su, Michael R. Lyu
摘要
Deep neural networks (DNNs) are known to be vulnerable to adversarial examples. It is thus imperative to devise effective attack algorithms to identify the deficiencies of DNNs beforehand in security-sensitive applications. To efficiently tackle the black-box setting where the target model's particulars are unknown, feature-level transfer-based attacks propose to contaminate the intermediate feature outputs of local models, and then directly employ the crafted adversarial samples to attack the target model. Due to the transferability of features, feature-level attacks have shown promise in synthesizing more transferable adversarial samples. However, existing feature-level attacks generally employ inaccurate neuron importance estimations, which deteriorates their transferability. To overcome such pitfalls, in this paper, we propose the Neuron Attribution-based Attack (NAA), which conducts feature-level attacks with more accurate neuron importance estimations. Specifically, we first completely attribute a model's output to each neuron in a middle layer. We then derive an approximation scheme of neuron attribution to tremendously reduce the computation overhead. Finally, we weight neurons based on their attribution results and launch feature-level attacks. Extensive experiments confirm the superiority of our approach to the state-of-the-art benchmarks. Our code is available at: https://github.com/jpzhang1810/NAA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- An Adaptive Model Ensemble Adversarial Attack for Boosting Adversarial TransferabilityBin Chen, Jia-Li Yin, Shukai Chen, Bohao Chen 等ICCV 2023 · 被引用 96 次
- BiasAsker: Measuring the Bias in Conversational AI SystemYuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu 等FSE 2023 · 被引用 50 次
- Improving Adversarial Transferability via Intermediate-level Perturbation DecayQizhang Li, Yiwen Guo, Wangmeng Zuo, Hao ChenNeurIPS 2023 · 被引用 43 次
- Towards Reasonable Budget Allocation in Untargeted Graph Structure Attacks via Gradient DebiasZihan Liu, Yun Luo, Lirong Wu, Zicheng Liu 等NeurIPS 2022 · 被引用 41 次
- Detecting Adversarial Data by Probing Multiple Perturbations Using Expected Perturbation ScoreShuhai Zhang, Feng Liu, Jiahao Yang, Yifan Yang 等ICML 2023 · 被引用 39 次
它引用的顶会 Paper9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang 等ICLR 2020 · 被引用 765 次
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor 等NeurIPS 2020 · 被引用 506 次
- Feature Importance-aware Transferable Adversarial AttacksZhibo Wang, Hengchang Guo, Zhifei Zhang, Wenxin Liu 等ICCV 2021 · 被引用 306 次
相关 Paper
- Pixel2Feature Attack (P2FA): Rethinking the Perturbed Space to Enhance Adversarial TransferabilityRenpu Liu, Hao Wu, Jiawei Zhang, Xin Cheng 等ICML 2025
- Learning Transferable Adversarial PerturbationsKrishna Kanth Nakka, Mathieu SalzmannNeurIPS 2021 · 被引用 75 次
- Improving the Transferability of Adversarial Samples With Adversarial TransformationsWeibin Wu, Yuxin Su, Michael R. Lyu, Irwin KingCVPR 2021
- Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack TransferabilityNathan Inkawhich, Kevin J. Liang, Binghui Wang, Matthew Inkawhich 等NeurIPS 2020 · 被引用 105 次
- Enhancing Adversarial Example Transferability With an Intermediate Level AttackQian Huang, Isay Katsman, Zeqi Gu, Horace He 等ICCV 2019 · 被引用 293 次
