Learning Concept Bottleneck Models from Mechanistic Explanations
Antonio De Santis, Schrasing Tong, Marco Brambilla, Lalana Kagal
摘要
Concept Bottleneck Models (CBMs) aim for ante-hoc interpretability by learning a bottleneck layer that predicts interpretable concepts before the decision. State-of-the-art approaches typically select which concepts to learn via human specification, open knowledge graphs, prompting an LLM, or using general CLIP concepts. However, concepts defined a-priori may not have sufficient predictive power for the task or even be learnable from the available data. As a result, these CBMs often significantly trail their black-box counterpart when controlling for information leakage. To address this, we introduce a novel CBM pipeline named Mechanistic CBM (M-CBM), which builds the bottleneck directly from a black-box model’s own learned concepts. These concepts are extracted via Sparse Autoencoders (SAEs) and subsequently named and annotated on a selected subset of images using a Multimodal LLM. For fair comparison and leakage control, we also introduce the Number of Contributing Concepts (NCC), a decision-level sparsity metric that extends the recently proposed NEC metric. Across diverse datasets, we show that M-CBMs consistently surpass prior CBMs at matched sparsity, while improving concept predictions and providing concise explanations. Our code is available at https://github.com/Antonio-Dee/M-CBM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li 等NeurIPS 2020 · 被引用 390 次
相关 Paper
- Interpretable and Steerable Concept Bottleneck Sparse AutoencodersAkshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy, Shusen Liu 等CVPR 2026 · 被引用 6 次
- VLG-CBM: Training Concept Bottleneck Models with Vision-Language GuidanceDivyansh Srivastava, Ge Yan, Lily WengNeurIPS 2024 · 被引用 87 次
- Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image ClassificationYue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin 等CVPR 2023
- V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept TokenizerHangzhou He, Lei Zhu, Xinliang Zhang, Shuang Zeng 等AAAI 2025 · 被引用 11 次
- Post-hoc Concept Bottleneck ModelsMert Yüksekgönül, Maggie Wang, James ZouICLR 2023 · 被引用 37 次
