A Little Robustness Goes a Long Way: Leveraging Robust Features for Targeted Transfer Attacks
Jacob M. Springer, Melanie Mitchell, Garrett T. Kenyon
Abstract
Adversarial examples for neural network image classifiers are known to be transferable: examples optimized to be misclassified by a source classifier are often misclassified as well by classifiers with different architectures. However, targeted adversarial examples -- optimized to be classified as a chosen target class -- tend to be less transferable between architectures. While prior research on constructing transferable targeted attacks has focused on improving the optimization procedure, in this work we examine the role of the source classifier. Here, we show that training the source classifier to be"slightly robust"-- that is, robust to small-magnitude adversarial examples -- substantially improves the transferability of class-targeted and representation-targeted adversarial attacks, even between architectures as different as convolutional neural networks and transformers. The results we present provide insight into the nature of adversarial examples as well as the mechanisms underlying so-called"robust"classifiers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 247c8b17-9d12-429c-a991-e0722b4e8f30Cited by top-tier papers16
- Why Does Little Robustness Help? A Further Step Towards Understanding Adversarial TransferabilityYechao Zhang, Shengshan Hu, Leo Yu Zhang, Junyu Shi et al.S&P 2024 · 36 citations
- What Can the Neural Tangent Kernel Tell Us About Adversarial Robustness?Nikolaos Tsilivis, Julia KempeNeurIPS 2022 · 28 citations
- Can Adversarial Training Be Manipulated By Non-Robust Features?Lue Tao, Lei Feng, Hongxin Wei, Jinfeng Yi et al.NeurIPS 2022 · 20 citations
- AGS: Affordable and Generalizable Substitute Training for Transferable Adversarial AttackRuikui Wang, Yuanfang Guo, Yunhong WangAAAI 2024 · 17 citations
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action ModelsHui Lu, Yi Yu, Yiming Yang, Chenyu Yi et al.CVPR 2026 · 12 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster InferenceBenjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock et al.ICCV 2021 · 1,009 citations
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor et al.NeurIPS 2020 · 506 citations
Related papers
- Towards Transferable Targeted Adversarial ExamplesZhibo Wang, Hongshan Yang, Yunhe Feng, Peng Sun et al.CVPR 2023
- On Adversarial Training without Perturbing all ExamplesMax Maria Losch, Mohamed Omran, David Stutz, Mario Fritz et al.ICLR 2024 · 5 citations
- Enhancing Adversarial Example Transferability With an Intermediate Level AttackQian Huang, Isay Katsman, Zeqi Gu, Horace He et al.ICCV 2019 · 293 citations
- On Success and Simplicity: A Second Look at Transferable Targeted AttacksZhengyu Zhao, Zhuoran Liu, Martha A. LarsonNeurIPS 2021 · 173 citations
- Uncovering the Connections Between Adversarial Transferability and Knowledge TransferabilityKaizhao Liang, Jacky Y. Zhang, Boxin Wang, Zhuolin Yang et al.ICML 2021 · 33 citations
