How Well do Feature Visualizations Support Causal Understanding of CNN Activations?
Roland S. Zimmermann, Judy Borowski, Robert Geirhos, Matthias Bethge, Thomas S. A. Wallis, Wieland Brendel
摘要
A precise understanding of why units in an artificial network respond to certain stimuli would constitute a big step towards explainable artificial intelligence. One widely used approach towards this goal is to visualize unit responses via activation maximization. These synthetic feature visualizations are purported to provide humans with precise information about the image features that cause a unit to be activated -an advantage over other alternatives like strongly activating natural dataset samples. If humans indeed gain causal insight from visualizations, this should enable them to predict the effect of an intervention, such as how occluding a certain patch of the image (say, a dog's head) changes a unit's activation. Here, we test this hypothesis by asking humans to decide which of two square occlusions causes a larger change to a unit's activation. Both a large-scale crowdsourced experiment and measurements with experts show that on average the extremely activating feature visualizations by Olah et al. [40] indeed help humans on this task (68 ± 4 % accuracy; baseline performance without any visualizations is 60 ± 3 %). However, they do not provide any substantial advantage over other visualizations (such as e.g. dataset samples), which yield similar performance (66±3 % to 67±3 % accuracy). Taken together, we propose an objective psychophysical task to quantify the benefit of unit-level interpretability methods for humans, and find no evidence that a widely-used feature visualization method provides humans with better "causal understanding" of unit activations than simple alternative visualizations. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Don't trust your eyes: on the (un)reliability of feature visualizationsRobert Geirhos, Roland S. Zimmermann, Blair L. Bilodeau, Wieland Brendel 等ICML 2024 · 被引用 38 次
- Scale Alone Does not Improve Mechanistic Interpretability in Vision ModelsRoland S. Zimmermann, Thomas Klein, Wieland BrendelNeurIPS 2023 · 被引用 32 次
- BrainSCUBA: Fine-Grained Natural Language Captions of Visual Cortex SelectivityAndrew F. Luo, Margaret M. Henderson, Michael J. Tarr, Leila WehbeICLR 2024 · 被引用 31 次
- Adversarial Attacks on the Interpretation of Neuron Activation MaximizationGéraldin Nanfack, Alexander Fulleringer, Jonathan Marty, Michael Eickenberg 等AAAI 2024 · 被引用 13 次
- Unlocking Feature Visualization for Deep Network with MAgnitude Constrained OptimizationThomas Fel, Thibaut Boissin, Victor Boutin, Agustin M. Picard 等NeurIPS 2023 · 被引用 12 次
它引用的顶会 Paper8
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 被引用 216 次
- Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistencyRobert Geirhos, Kristof Meding, Felix A. WichmannNeurIPS 2020 · 被引用 154 次
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas 等ICML 2020 · 被引用 146 次
相关 Paper
- Exemplary Natural Images Explain CNN Activations Better than State-of-the-Art Feature VisualizationJudy Borowski, Roland Simon Zimmermann, Judith Schepers, Robert Geirhos 等ICLR 2021 · 被引用 5 次
- Stretching Beyond the Obvious: A Gradient-Free Framework to Unveil the Hidden Landscape of Visual InvarianceLorenzo Tausani, Paolo Muratore, Morgan Bruce Talbot, Giacomo Amerio 等ICLR 2026
- Do Users Benefit From Interpretable Vision? A User Study, Baseline, And DatasetLeon Sixt, Martin Schuessler, Oana-Iuliana Popescu, Philipp Weiß 等ICLR 2022 · 被引用 21 次
- Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingAleksandar Makelov, Georg Lange, Atticus Geiger, Neel NandaICLR 2024 · 被引用 48 次
- Measuring Per-Unit Interpretability at Scale Without HumansRoland S. Zimmermann, David A. Klindt, Wieland BrendelNeurIPS 2024 · 被引用 5 次
