Guaranteed Optimal Compositional Explanations for Neurons
Biagio La Rosa, Leilani Gilpin
摘要
Compositional explanations are a family of methods that aim to describe the spatial alignment between neurons' receptive field activations and concepts through logical rules, typically computed via a search over all possible concept combinations. Since computing the spatial alignment over the entire state space is computationally infeasible, the literature commonly adopts assumptions related to the structure of the combinations and beam search to restrict the state space. However, beam search cannot provide any theoretical guarantees of optimality, and it remains unclear how close current explanations are to the true optimum. In this theoretical paper, we address this gap by introducing the first framework for computing guaranteed optimal compositional explanations over the entire state space spanned by the adopted assumptions. Specifically, we propose: (i) a decomposition that identifies the factors influencing the spatial alignment, (ii) a heuristic to estimate the alignment at any stage of the search, and (iii) the first algorithm that can compute optimal compositional explanations in a time comparable to exhaustive beam search. Using this framework, we demonstrate that 10-40% of explanations previously obtained with beam search are suboptimal when overlapping concepts are involved. Finally, we evaluate a beam-search variant guided by our proposed decomposition and heuristic, showing that it matches or improves runtime over prior methods while offering greater flexibility in hyperparameters and computational resources. Optimal Compositional Explanations Preliminaries and Terminology Let L 1 be a concept set including properties of interest for a given model and task (e.g., names of colors, objects, categories, shapes). Let D = x 1 , x 2 , ..., x n be a dataset, where each input x i has
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
- Identifying Interpretable Subspaces in Image RepresentationsNeha Mukund Kalibhat, Shweta Bhardwaj, C. Bayan Bruss, Hamed Firooz 等ICML 2023 · 被引用 41 次
- Labeling Neural Representations with Inverse RecognitionKirill Bykov, Laura Kopf, Shinichi Nakajima, Marius Kloft 等NeurIPS 2023 · 被引用 36 次
- Linear Explanations for Individual NeuronsTuomas P. Oikarinen, Tsui-Wei WengICML 2024 · 被引用 18 次
相关 Paper
- Towards a fuller understanding of neurons with Clustered Compositional ExplanationsBiagio La Rosa, Leilani Gilpin, Roberto CapobiancoNeurIPS 2023 · 被引用 17 次
- Towards Compositionality in Concept LearningAdam Stein, Aaditya Naik, Yinjun Wu, Mayur Naik 等ICML 2024 · 被引用 11 次
- Decomposition of Concept-Level Rules in Visual ScenesFan Shi, Yuxuan Liang, Xiaolei Chen, Haiyang Yu 等ICLR 2026
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue 等ICLR 2020 · 被引用 55 次
- Neural Compositional Rule Learning for Knowledge Graph ReasoningKewei Cheng, Nesreen K. Ahmed, Yizhou SunICLR 2023 · 被引用 15 次
