Adversarial Attacks on the Interpretation of Neuron Activation Maximization
Géraldin Nanfack, Alexander Fulleringer, Jonathan Marty, Michael Eickenberg, Eugene Belilovsky
Abstract
The internal functional behavior of trained Deep Neural Networks is notoriously difficult to interpret. Activation-maximization approaches are one set of techniques used to interpret and analyze trained deep-learning models. These consist in finding inputs that maximally activate a given neuron or feature map. These inputs can be selected from a data set or obtained by optimization. However, interpretability methods may be subject to being deceived. In this work, we consider the concept of an adversary manipulating a model for the purpose of deceiving the interpretation. We propose an optimization framework for performing this manipulation and demonstrate a number of ways that popular activation-maximization interpretation techniques associated with CNNs can be manipulated to change the interpretations, shedding light on the reliability of these methods. * Equal contribution. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Don't trust your eyes: on the (un)reliability of feature visualizationsRobert Geirhos, Roland S. Zimmermann, Blair L. Bilodeau, Wieland Brendel et al.ICML 2024 · 38 citations
- Linear Explanations for Individual NeuronsTuomas P. Oikarinen, Tsui-Wei WengICML 2024 · 18 citations
- Manipulating Feature Visualizations with Gradient SlingshotsDilyara Bareeva, Marina M.-C. Höhne, Alexander Warnecke, Lukas Pirch et al.NeurIPS 2025 · 8 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- On Feature Learning in the Presence of Spurious CorrelationsPavel Izmailov, Polina Kirichenko, Nate Gruver, Andrew Gordon WilsonNeurIPS 2022 · 208 citations
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 163 citations
- Natural Language Descriptions of Deep Visual FeaturesEvan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili et al.ICLR 2022 · 160 citations
Related papers
- Fooling Network Interpretation in Image ClassificationAkshayvarun Subramanya, Vipin Pillai, Hamed PirsiavashICCV 2019 · 68 citations
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji et al.USENIX Security 2020
- DANCE: Enhancing saliency maps using decoysYang Young Lu, Wenbo Guo, Xinyu Xing, William Stafford NobleICML 2021 · 14 citations
- Proper Network Interpretability Helps Adversarial Robustness in ClassificationAkhilan Boopathy, Sijia Liu, Gaoyuan Zhang, Cynthia Liu et al.ICML 2020 · 74 citations
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
