Guaranteed Optimal Compositional Explanations for Neurons
Biagio La Rosa, Leilani Gilpin
Abstract
Compositional explanations are a family of methods that aim to describe the spatial alignment between neurons' receptive field activations and concepts through logical rules, typically computed via a search over all possible concept combinations. Since computing the spatial alignment over the entire state space is computationally infeasible, the literature commonly adopts assumptions related to the structure of the combinations and beam search to restrict the state space. However, beam search cannot provide any theoretical guarantees of optimality, and it remains unclear how close current explanations are to the true optimum. In this theoretical paper, we address this gap by introducing the first framework for computing guaranteed optimal compositional explanations over the entire state space spanned by the adopted assumptions. Specifically, we propose: (i) a decomposition that identifies the factors influencing the spatial alignment, (ii) a heuristic to estimate the alignment at any stage of the search, and (iii) the first algorithm that can compute optimal compositional explanations in a time comparable to exhaustive beam search. Using this framework, we demonstrate that 10-40% of explanations previously obtained with beam search are suboptimal when overlapping concepts are involved. Finally, we evaluate a beam-search variant guided by our proposed decomposition and heuristic, showing that it matches or improves runtime over prior methods while offering greater flexibility in hyperparameters and computational resources. Optimal Compositional Explanations Preliminaries and Terminology Let L 1 be a concept set including properties of interest for a given model and task (e.g., names of colors, objects, categories, shapes). Let D = x 1 , x 2 , ..., x n be a dataset, where each input x i has
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Identifying Interpretable Subspaces in Image RepresentationsNeha Mukund Kalibhat, Shweta Bhardwaj, C. Bayan Bruss, Hamed Firooz et al.ICML 2023 · 41 citations
- Labeling Neural Representations with Inverse RecognitionKirill Bykov, Laura Kopf, Shinichi Nakajima, Marius Kloft et al.NeurIPS 2023 · 36 citations
- Linear Explanations for Individual NeuronsTuomas P. Oikarinen, Tsui-Wei WengICML 2024 · 18 citations
Related papers
- Towards a fuller understanding of neurons with Clustered Compositional ExplanationsBiagio La Rosa, Leilani Gilpin, Roberto CapobiancoNeurIPS 2023 · 17 citations
- Towards Compositionality in Concept LearningAdam Stein, Aaditya Naik, Yinjun Wu, Mayur Naik et al.ICML 2024 · 11 citations
- Decomposition of Concept-Level Rules in Visual ScenesFan Shi, Yuxuan Liang, Xiaolei Chen, Haiyang Yu et al.ICLR 2026
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue et al.ICLR 2020 · 55 citations
- Neural Compositional Rule Learning for Knowledge Graph ReasoningKewei Cheng, Nesreen K. Ahmed, Yizhou SunICLR 2023 · 15 citations
