Towards a fuller understanding of neurons with Clustered Compositional Explanations
Biagio La Rosa, Leilani Gilpin, Roberto Capobianco
摘要
Compositional Explanations [30] is a method for identifying logical formulas of concepts that approximate the neurons' behavior. However, these explanations are linked to the small spectrum of neuron activations (i.e., the highest ones) used to check the alignment, thus lacking completeness. In this paper, we propose a generalization, called Clustered Compositional Explanations, that combines Compositional Explanations with clustering and a novel search heuristic to approximate a broader spectrum of the neuron behavior. We define and address the problems connected to the application of these methods to multiple ranges of activations, analyze the insights retrievable by using our algorithm, and propose desiderata qualities that can be used to study the explanations returned by different algorithms. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Linear Explanations for Individual NeuronsTuomas P. Oikarinen, Tsui-Wei WengICML 2024 · 被引用 18 次
- Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing Their ContributionsNamitha Padmanabhan, Matthew Gwilliam, Pulkit Kumar, Shishira R. Maiya 等CVPR 2024 · 被引用 2 次
- What is Missing? Explaining Neurons Activated by Absent ConceptsRobin Hesse, Simone Schaub-Meyer, Janina Hesse, Bernt Schiele 等ICML 2026 · 被引用 1 次
- Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated InterpretabilityTuomas Oikarinen, Ge Yan, Akshay Kulkarni, Tsui-Wei WengCVPR 2026 · 被引用 1 次
- Select, Hypothesize and Verify: Towards Verified Neuron Concept InterpretationZeBin Ji, Yang Hu, Xiuli Bi, Bo Liu 等CVPR 2026
它引用的顶会 Paper7
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
- On the Pitfalls of Analyzing Individual Neurons in Language ModelsOmer Antverg, Yonatan BelinkovICLR 2022 · 被引用 64 次
- Finding Skill Neurons in Pre-trained Transformer-based Language ModelsXiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou 等EMNLP 2022 · 被引用 19 次
- CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision NetworksTuomas P. Oikarinen, Tsui-Wei WengICLR 2023 · 被引用 9 次
- Exemplary Natural Images Explain CNN Activations Better than State-of-the-Art Feature VisualizationJudy Borowski, Roland Simon Zimmermann, Judith Schepers, Robert Geirhos 等ICLR 2021 · 被引用 5 次
相关 Paper
- Guaranteed Optimal Compositional Explanations for NeuronsBiagio La Rosa, Leilani GilpinICML 2026
- Global Explainability of GNNs via Logic Combination of Learned ConceptsSteve Azzolin, Antonio Longa, Pietro Barbiero, Pietro Liò 等ICLR 2023 · 被引用 11 次
- Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?Maxime Méloux, Silviu Maniu, François Portet, Maxime PeyrardICLR 2025
- Global Concept-Based Interpretability for Graph Neural Networks via Neuron AnalysisHan Xuanyuan, Pietro Barbiero, Dobrik Georgiev, Lucie Charlotte Magister 等AAAI 2023 · 被引用 62 次
- What's in the Box? Exploring the Inner Life of Neural Networks with Robust RulesJonas Fischer, Anna Oláh, Jilles VreekenICML 2021 · 被引用 11 次
