Minimizing Maximum Model Discrepancy for Transferable Black-box Targeted Attacks
Anqi Zhao, Tong Chu, Yahao Liu, Wen Li, Jingjing Li, Lixin Duan
Abstract
In this work, we study the black-box targeted attack problem from the model discrepancy perspective. On the theoretical side, we present a generalization error bound for black-box targeted attacks, which gives a rigorous theoretical analysis for guaranteeing the success of the attack. We reveal that the attack error on a target model mainly depends on empirical attack error on the substitute model and the maximum model discrepancy among substitute models. On the algorithmic side, we derive a new algorithm for black-box targeted attack based on our theoretical analysis, in which we additionally minimize the maximum model discrepancy (M3D) of the substitute models when training the generator to generate adversarial examples. In this way, our model is capable of crafting highly transferable adversarial examples that are robust to the model variation, thus improving the success rate for attacking the black-box model. We conduct extensive experiments on the ImageNet dataset with different classification models, and our proposed approach outperforms existing state-of-the-art methods by a significant margin. Our codes will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset UsageWenyi Zhang, Ju Jia, Xiaojun Jia, Yihao Huang et al.SIGIR 2025 · 3 citations
- Optimal transport based adversarial patch to leverage large scale attack transferabilityPol Labarbarie, Adrien Chan-Hon-Tong, Stéphane Herbin, Milad Leyli-AbadiICLR 2024 · 1 citation
- Adversarial Perturbations Are Formed by Iteratively Learning Linear Combinations of the Right Singular Vectors of the Adversarial JacobianThomas Paniagua, Chinmay Savadikar, Tianfu WuICML 2025
- Understanding Model Ensemble in Transferable Adversarial AttackWei Yao, Zeliang Zhang, Huayi Tang, Yong LiuICML 2025
- Automated Mass Malware Factory: The Convergence of Piggybacking and Adversarial Example in Android Malicious Software GenerationHeng Li, Zhiyuan Yao, Bang Wu, Cuiying Gao et al.NDSS 2025
Builds on16
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor et al.NeurIPS 2020 · 506 citations
- Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNetsDongxian Wu, Yisen Wang, Shu-Tao Xia, James Bailey et al.ICLR 2020 · 357 citations
Related papers
- Training Meta-Surrogate Model for Transferable Adversarial AttackYunxiao Qin, Yuanhao Xiong, Jinfeng Yi, Cho-Jui HsiehAAAI 2023 · 31 citations
- Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box DomainsQilong Zhang, Xiaodan Li, Yuefeng Chen, Jingkuan Song et al.ICLR 2022 · 85 citations
- On Generating Transferable Targeted PerturbationsMuzammal Naseer, Salman H. Khan, Munawar Hayat, Fahad Shahbaz Khan et al.ICCV 2021 · 93 citations
- Towards Multiple Black-boxes Attack via Adversarial Example Generation NetworkMingxing Duan, Kenli Li, Lingxi Xie, Qi Tian et al.ACM MM 2021 · 21 citations
- Blurred-Dilated Method for Adversarial AttacksYang Deng, Weibin Wu, Jianping Zhang, Zibin ZhengNeurIPS 2023 · 10 citations
