Towards a fuller understanding of neurons with Clustered Compositional Explanations
Biagio La Rosa, Leilani Gilpin, Roberto Capobianco
Abstract
Compositional Explanations [30] is a method for identifying logical formulas of concepts that approximate the neurons' behavior. However, these explanations are linked to the small spectrum of neuron activations (i.e., the highest ones) used to check the alignment, thus lacking completeness. In this paper, we propose a generalization, called Clustered Compositional Explanations, that combines Compositional Explanations with clustering and a novel search heuristic to approximate a broader spectrum of the neuron behavior. We define and address the problems connected to the application of these methods to multiple ranges of activations, analyze the insights retrievable by using our algorithm, and propose desiderata qualities that can be used to study the explanations returned by different algorithms. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab8a696a-648b-40ea-8148-bf9caef786bfCited by top-tier papers9
- Linear Explanations for Individual NeuronsTuomas P. Oikarinen, Tsui-Wei WengICML 2024 · 18 citations
- Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing Their ContributionsNamitha Padmanabhan, Matthew Gwilliam, Pulkit Kumar, Shishira R. Maiya et al.CVPR 2024 · 2 citations
- What is Missing? Explaining Neurons Activated by Absent ConceptsRobin Hesse, Simone Schaub-Meyer, Janina Hesse, Bernt Schiele et al.ICML 2026 · 1 citation
- Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated InterpretabilityTuomas Oikarinen, Ge Yan, Akshay Kulkarni, Tsui-Wei WengCVPR 2026 · 1 citation
- Select, Hypothesize and Verify: Towards Verified Neuron Concept InterpretationZeBin Ji, Yang Hu, Xiuli Bi, Bo Liu et al.CVPR 2026
Builds on7
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- On the Pitfalls of Analyzing Individual Neurons in Language ModelsOmer Antverg, Yonatan BelinkovICLR 2022 · 64 citations
- Finding Skill Neurons in Pre-trained Transformer-based Language ModelsXiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou et al.EMNLP 2022 · 19 citations
- CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision NetworksTuomas P. Oikarinen, Tsui-Wei WengICLR 2023 · 9 citations
- Exemplary Natural Images Explain CNN Activations Better than State-of-the-Art Feature VisualizationJudy Borowski, Roland Simon Zimmermann, Judith Schepers, Robert Geirhos et al.ICLR 2021 · 5 citations
Related papers
- Guaranteed Optimal Compositional Explanations for NeuronsBiagio La Rosa, Leilani GilpinICML 2026
- Global Explainability of GNNs via Logic Combination of Learned ConceptsSteve Azzolin, Antonio Longa, Pietro Barbiero, Pietro Liò et al.ICLR 2023 · 11 citations
- Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?Maxime Méloux, Silviu Maniu, François Portet, Maxime PeyrardICLR 2025
- Global Concept-Based Interpretability for Graph Neural Networks via Neuron AnalysisHan Xuanyuan, Pietro Barbiero, Dobrik Georgiev, Lucie Charlotte Magister et al.AAAI 2023 · 62 citations
- What's in the Box? Exploring the Inner Life of Neural Networks with Robust RulesJonas Fischer, Anna Oláh, Jilles VreekenICML 2021 · 11 citations
