Improving Transferable Targeted Adversarial Attacks with Model Self-Enhancement
Han Wu, Guanyan Ou, Weibin Wu, Zibin Zheng
Abstract
Various transfer attack methods have been proposed to evaluate the robustness of deep neural networks (DNNs). Although manifesting remarkable performance in generating untargeted adversarial perturbations, existing proposals still fail to achieve high targeted transferability. In this work, we discover that the adversarial perturbations' over-fitting towards source models of mediocre generalization capability can hurt their targeted transferability. To address this issue, we focus on enhancing the source model's gener-alization capability to improve its ability to conduct trans-ferable targeted adversarial attacks. In pursuit of this goal, we propose a novel model self-enhancement method that in-corporates two major components: Sharpness-Aware Self-Distillation (SASD) and Weight Scaling (WS). Specifically, SASD distills a fine-tuned auxiliary model, which mirrors the source model's structure, into the source model while flattening the source model's loss landscape. WS obtains an approximate ensemble of numerous pruned models to per-form model augmentation, which can be conveniently syn-ergized with SASD to elevate the source model's generalization capability and thus improve the resultant targeted per-turbations' transferability. Extensive experiments corrobo-rate the effectiveness of the proposed method. Notably, under the black-box setting, our approach can outperform the state-of-the-art baselines by a significant margin of 12.2% on average in terms of the obtained targeted transferability. Code is available at https://github.com/g4alllf/SASD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-ExpertsLi Bai, Qingqing Ye, Xinwei Zhang, Sen Zhang et al.NeurIPS 2025 · 6 citations
- Learning Robust Vision-Language Models from Natural Latent SpacesZhangyun Wang, Ni Ding, Aniket MahantiNeurIPS 2025 · 3 citations
- RaPA: Enhancing Transferable Targeted Attacks via Random Parameter PruningTongrui Su, Qingbin Li, Shengyu Zhu, Wei Chen et al.CVPR 2026 · 1 citation
- Understanding Model Ensemble in Transferable Adversarial AttackWei Yao, Zeliang Zhang, Huayi Tang, Yong LiuICML 2025
- WhisperSplat: Lossless Steganography in 3D Gaussian SplattingNicole Meng, Ronak Sahu, Miao Yin, Faysal Hossain Shezan et al.ICML 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
Related papers
- Blurred-Dilated Method for Adversarial AttacksYang Deng, Weibin Wu, Jianping Zhang, Zibin ZhengNeurIPS 2023 · 10 citations
- StyLess: Boosting the Transferability of Adversarial ExamplesKaisheng Liang, Bin XiaoCVPR 2023
- Towards Transferable Targeted Adversarial ExamplesZhibo Wang, Hongshan Yang, Yunhe Feng, Peng Sun et al.CVPR 2023
- Enhancing Adversarial Transferability with Adversarial Weight TuningJiahao Chen, Zhou Feng, Rui Zeng, Yuwen Pu et al.AAAI 2025 · 11 citations
- How to choose your best allies for a transferable attack?Thibault Maho, Seyed-Mohsen Moosavi-Dezfooli, Teddy FuronICCV 2023 · 1 citation
