Towards Compositionality in Concept Learning
Adam Stein, Aaditya Naik, Yinjun Wu, Mayur Naik, Eric Wong
Abstract
Concept-based interpretability methods offer a lens into the internals of foundation models by decomposing their embeddings into high-level concepts. These concept representations are most useful when they are compositional, meaning that the individual concepts compose to explain the full sample. We show that existing unsupervised concept extraction methods find concepts which are not compositional. To automatically discover compositional concept representations, we identify two salient properties of such representations, and propose Compositional Concept Extraction (CCE) for finding concepts which obey these properties. We evaluate CCE on five different datasets over image and text data. Our evaluation shows that CCE finds more compositional concept representations than baselines and yields better accuracy on four downstream classification tasks. Code and data are available at https://github.com/adaminsky/compositional_concepts .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c1c5e70-d3fe-43fd-beb1-4ea0195fee18Cited by top-tier papers6
- The Geometry of Reasoning: Flowing Logics in Representation SpaceYufa Zhou, Yixiao Wang, Xunjian Yin, Shuyan Zhou et al.ICLR 2026 · 29 citations
- FACE: Faithful Automatic Concept ExtractionDipkamal Bhusal, Michael Clifford, Sara Rampazzi, Nidhi RastogiNeurIPS 2025 · 11 citations
- Intrinsic Concept Extraction Based on Compositional InterpretabilityHanyu Shi, Hong Tao, Guoheng Huang, Jianbin Jiang et al.CVPR 2026
- Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual AbductionShanshan Huang, Haoxuan Li, Chunyuan Zheng, Mingyuan Ge et al.CVPR 2025
- SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph EmbeddingYuan Zang, Zitian Tang, Junho Cho, Jaewook Yoo et al.NeurIPS 2025
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 796 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
Related papers
- Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)Usha Bhalla, Alex Oesterling, Suraj Srinivas, Flávio P. Calmon et al.NeurIPS 2024 · 146 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Overlooked Factors in Concept-Based Explanations: Dataset Choice, Concept Learnability, and Human CapabilityVikram V. Ramaswamy, Sunnie S. Y. Kim, Ruth Fong, Olga RussakovskyCVPR 2023
- Identifying Interpretable Subspaces in Image RepresentationsNeha Mukund Kalibhat, Shweta Bhardwaj, C. Bayan Bruss, Hamed Firooz et al.ICML 2023 · 41 citations
- Learning by Analogy: A Causal Framework for Compositional GeneralizationLingjing Kong, Shaoan Xie, Yang Jiao, Yetian Chen et al.CVPR 2026
