Labeling Neural Representations with Inverse Recognition
Kirill Bykov, Laura Kopf, Shinichi Nakajima, Marius Kloft, Marina M.-C. Höhne
摘要
Deep Neural Networks (DNNs) demonstrate remarkable capabilities in learning complex hierarchical data representations, but the nature of these representations remains largely unknown. Existing global explainability methods, such as Network Dissection, face limitations such as reliance on segmentation masks, lack of statistical significance testing, and high computational demands. We propose Inverse Recognition (INVERT), a scalable approach for connecting learned representations with human-understandable concepts by leveraging their capacity to discriminate between these concepts. In contrast to prior work, INVERT is capable of handling diverse types of neurons, exhibits less computational complexity, and does not rely on the availability of segmentation masks. Moreover, INVERT provides an interpretable metric assessing the alignment between the representation and its corresponding explanation and delivering a measure of statistical significance. We demonstrate the applicability of INVERT in various scenarios, including the identification of representations affected by spurious correlations, and the interpretation of the hierarchical structure of decision-making within the models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- CoSy: Evaluating Textual Explanations of NeuronsLaura Kopf, Philine Lou Bommer, Anna Hedström, Sebastian Lapuschkin 等NeurIPS 2024 · 被引用 23 次
- Linear Explanations for Individual NeuronsTuomas P. Oikarinen, Tsui-Wei WengICML 2024 · 被引用 18 次
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description FrameworkLaura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer 等NeurIPS 2025 · 被引用 12 次
- LaVCa: LLM-assisted Visual Cortex CaptioningTakuya Matsuyama, Shinji Nishimoto, Yu TakagiICLR 2026 · 被引用 8 次
- Manipulating Feature Visualizations with Gradient SlingshotsDilyara Bareeva, Marina M.-C. Höhne, Alexander Warnecke, Lukas Pirch 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
- Natural Language Descriptions of Deep Visual FeaturesEvan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili 等ICLR 2022 · 被引用 160 次
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas 等ICLR 2023 · 被引用 60 次
相关 Paper
- A Disentangling Invertible Interpretation Network for Explaining Latent RepresentationsPatrick Esser, Robin Rombach, Björn OmmerCVPR 2020
- Brain Dissection: fMRI-trained Networks Reveal Spatial Selectivity in the Processing of Natural ImagesGabriel Sarch, Michael J. Tarr, Katerina Fragkiadaki, Leila WehbeNeurIPS 2023 · 被引用 20 次
- Interpreting Deep Neural Networks with Relative Sectional Propagation by Analyzing Comparative Gradients and Hostile ActivationsWoo-Jeoung Nam, Jaesik Choi, Seong-Whan LeeAAAI 2021 · 被引用 19 次
- DISCOVER: Making Vision Networks Interpretable via Competition and DissectionKonstantinos P. Panousis, Sotirios ChatzisNeurIPS 2023 · 被引用 9 次
- Comparing the Decision-Making Mechanisms by Transformers and CNNs via Explanation MethodsMingqi Jiang, Saeed Khorram, Fuxin LiCVPR 2024
