A Theory of Transfer-Based Black-Box Attacks: Explanation and Implications
Yanbo Chen, Weiwei Liu
Abstract
Transfer-based attacks [1] are a practical method of black-box adversarial attacks in which the attacker aims to craft adversarial examples from a source model that is transferable to the target model. Many empirical works [2–6] have tried to explain the transferability of adversarial examples from different angles. However, these works only provide ad hoc explanations without quantitative analyses. The theory behind transfer-based attacks remains a mystery. This paper studies transfer-based attacks under a unified theoretical framework. We propose an explanatory model, called the manifold attack model , that formalizes popular beliefs and explains the existing empirical results. Our model explains why adversarial examples are transferable even when the source model is inaccurate as observed in Papernot et al. [7]. Moreover, our model implies that the existence of transferable adversarial examples depends on the “curvature” of the data manifold, which further explains why the success rates of transfer-based attacks are hard to improve. We also discuss our model’s expressive power and applicability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2b8b1da-a2b4-4b7f-9267-32469fff3244Cited by top-tier papers14
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language ModelsJiaxin Song, Yixu Wang, Jie Li, Xuan Tong et al.NeurIPS 2025 · 14 citations
- A Closer Look at Curriculum Adversarial Training: From an Online PerspectiveLianghe Shi, Weiwei LiuAAAI 2024 · 7 citations
- DRF: Improving Certified Robustness via Distributional Robustness FrameworkZekai Wang, Zhengyu Zhou, Weiwei LiuAAAI 2024 · 7 citations
- A Provable Decision Rule for Out-of-Distribution DetectionXinsong Ma, Xin Zou, Weiwei LiuICML 2024 · 4 citations
- Vulnerable Data-Aware Adversarial TrainingYuqi Feng, Jiahao Fan, Yanan SunNeurIPS 2025 · 2 citations
Builds on14
- A Closer Look at Accuracy vs. RobustnessYao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov et al.NeurIPS 2020 · 336 citations
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin et al.ICML 2023 · 300 citations
- Perceptual Adversarial Robustness: Defense Against Unseen Threat ModelsCassidy Laidlaw, Sahil Singla, Soheil FeiziICLR 2021 · 217 citations
- Towards Transferable Adversarial Attacks on Vision TransformersZhipeng Wei, Jingjing Chen, Micah Goldblum, Zuxuan Wu et al.AAAI 2022 · 156 citations
- Black-Box Adversarial Attack with Transferable Model-based EmbeddingZhichao Huang, Tong ZhangICLR 2020 · 131 citations
Related papers
- Enhancing Adversarial Example Transferability With an Intermediate Level AttackQian Huang, Isay Katsman, Zeqi Gu, Horace He et al.ICCV 2019 · 293 citations
- Transferability Bound Theory: Exploring Relationship between Adversarial Transferability and FlatnessMingyuan Fan, Xiaodan Li, Cen Chen, Wenmeng Zhou et al.NeurIPS 2024 · 13 citations
- Blurred-Dilated Method for Adversarial AttacksYang Deng, Weibin Wu, Jianping Zhang, Zibin ZhengNeurIPS 2023 · 10 citations
- Global-Local Characteristic Excited Cross-Modal Attacks from Images to VideosRuikui Wang, Yuanfang Guo, Yunhong WangAAAI 2023 · 15 citations
- Minimizing Maximum Model Discrepancy for Transferable Black-box Targeted AttacksAnqi Zhao, Tong Chu, Yahao Liu, Wen Li et al.CVPR 2023
