A Theory of Transfer-Based Black-Box Attacks: Explanation and Implications
Yanbo Chen, Weiwei Liu
摘要
Transfer-based attacks [1] are a practical method of black-box adversarial attacks in which the attacker aims to craft adversarial examples from a source model that is transferable to the target model. Many empirical works [2–6] have tried to explain the transferability of adversarial examples from different angles. However, these works only provide ad hoc explanations without quantitative analyses. The theory behind transfer-based attacks remains a mystery. This paper studies transfer-based attacks under a unified theoretical framework. We propose an explanatory model, called the manifold attack model , that formalizes popular beliefs and explains the existing empirical results. Our model explains why adversarial examples are transferable even when the source model is inaccurate as observed in Papernot et al. [7]. Moreover, our model implies that the existence of transferable adversarial examples depends on the “curvature” of the data manifold, which further explains why the success rates of transfer-based attacks are hard to improve. We also discuss our model’s expressive power and applicability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language ModelsJiaxin Song, Yixu Wang, Jie Li, Xuan Tong 等NeurIPS 2025 · 被引用 14 次
- A Closer Look at Curriculum Adversarial Training: From an Online PerspectiveLianghe Shi, Weiwei LiuAAAI 2024 · 被引用 7 次
- DRF: Improving Certified Robustness via Distributional Robustness FrameworkZekai Wang, Zhengyu Zhou, Weiwei LiuAAAI 2024 · 被引用 7 次
- A Provable Decision Rule for Out-of-Distribution DetectionXinsong Ma, Xin Zou, Weiwei LiuICML 2024 · 被引用 4 次
- Vulnerable Data-Aware Adversarial TrainingYuqi Feng, Jiahao Fan, Yanan SunNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper14
- A Closer Look at Accuracy vs. RobustnessYao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov 等NeurIPS 2020 · 被引用 336 次
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin 等ICML 2023 · 被引用 300 次
- Perceptual Adversarial Robustness: Defense Against Unseen Threat ModelsCassidy Laidlaw, Sahil Singla, Soheil FeiziICLR 2021 · 被引用 217 次
- Towards Transferable Adversarial Attacks on Vision TransformersZhipeng Wei, Jingjing Chen, Micah Goldblum, Zuxuan Wu 等AAAI 2022 · 被引用 156 次
- Black-Box Adversarial Attack with Transferable Model-based EmbeddingZhichao Huang, Tong ZhangICLR 2020 · 被引用 131 次
相关 Paper
- Enhancing Adversarial Example Transferability With an Intermediate Level AttackQian Huang, Isay Katsman, Zeqi Gu, Horace He 等ICCV 2019 · 被引用 293 次
- Transferability Bound Theory: Exploring Relationship between Adversarial Transferability and FlatnessMingyuan Fan, Xiaodan Li, Cen Chen, Wenmeng Zhou 等NeurIPS 2024 · 被引用 13 次
- Blurred-Dilated Method for Adversarial AttacksYang Deng, Weibin Wu, Jianping Zhang, Zibin ZhengNeurIPS 2023 · 被引用 10 次
- Global-Local Characteristic Excited Cross-Modal Attacks from Images to VideosRuikui Wang, Yuanfang Guo, Yunhong WangAAAI 2023 · 被引用 15 次
- Minimizing Maximum Model Discrepancy for Transferable Black-box Targeted AttacksAnqi Zhao, Tong Chu, Yahao Liu, Wen Li 等CVPR 2023
