Robust Feature-Level Adversaries are Interpretability Tools
Stephen Casper, Max Nadeau, Dylan Hadfield-Menell, Gabriel Kreiman
Abstract
The literature on adversarial attacks in computer vision typically focuses on pixel-level perturbations. These tend to be very difficult to interpret. Recent work that manipulates the latent representations of image generators to create"feature-level"adversarial perturbations gives us an opportunity to explore perceptible, interpretable adversarial attacks. We make three contributions. First, we observe that feature-level attacks provide useful classes of inputs for studying representations in models. Second, we show that these adversaries are uniquely versatile and highly robust. We demonstrate that they can be used to produce targeted, universal, disguised, physically-realizable, and black-box attacks at the ImageNet scale. Third, we show how these adversarial images can be used as a practical interpretability tool for identifying bugs in networks. We use these adversaries to make predictions about spurious associations between features and classes which we then test by designing"copy/paste"attacks in which one natural image is pasted into another to cause a targeted misclassification. Our results suggest that feature-level attacks are a promising approach for rigorous interpretability research. They support the design of tools to better understand what a model has learned and diagnose brittle feature associations. Code is available at https://github.com/thestephencasper/feature_level_adv
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f7ac91c-9136-45b9-8acf-a0e88b468b85Cited by top-tier papers9
- Adversarial training for high-stakes reliabilityDaniel M. Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman et al.NeurIPS 2022 · 79 citations
- Red Teaming Deep Neural Networks with Feature Synthesis ToolsStephen Casper, Tong Bu, Yuxiao Li, Jiawei Li et al.NeurIPS 2023 · 23 citations
- Breaking Barriers in Physical-World Adversarial Examples: Improving Robustness and Transferability via Robust FeatureYichen Wang, Yuxuan Chou, Ziqi Zhou, Hangtao Zhang et al.AAAI 2025 · 20 citations
- Corrupting Neuron Explanations of Deep Visual FeaturesDivyansh Srivastava, Tuomas P. Oikarinen, Tsui-Wei WengICCV 2023 · 3 citations
- SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV DetectorsAixuan Li, Mochu Xiang, Bosen Hou, Zhexiong Wan et al.CVPR 2026 · 1 citation
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor et al.NeurIPS 2020 · 506 citations
- Simulating a Primary Visual Cortex at the Front of CNNs Improves Robustness to Image PerturbationsJoel Dapello, Tiago Marques, Martin Schrimpf, Franziska Geiger et al.NeurIPS 2020 · 250 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
Related papers
- Transferable Perturbations of Deep Feature DistributionsNathan Inkawhich, Kevin J. Liang, Lawrence Carin, Yiran ChenICLR 2020 · 100 citations
- Towards Feature Space Adversarial Attack by Style PerturbationQiuling Xu, Guanhong Tao, Siyuan Cheng, Xiangyu ZhangAAAI 2021 · 33 citations
- Fooling Network Interpretation in Image ClassificationAkshayvarun Subramanya, Vipin Pillai, Hamed PirsiavashICCV 2019 · 68 citations
- Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack TransferabilityNathan Inkawhich, Kevin J. Liang, Binghui Wang, Matthew Inkawhich et al.NeurIPS 2020 · 105 citations
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
