Editable Concept Bottleneck Models
Lijie Hu, Chenyang Ren, Zhengyu Hu, Hongbin Lin, Cheng-Long Wang, Zhen Tan, Weimin Lyu, Jingfeng Zhang, Hui Xiong, Di Wang
Abstract
Concept Bottleneck Models (CBMs) have garnered much attention for their ability to elucidate the prediction process through a humanunderstandable concept layer. However, most previous studies focused on cases where the data, including concepts, are clean. In many scenarios, we often need to remove/insert some training data or new concepts from trained CBMs for reasons such as privacy concerns, data mislabelling, spurious concepts, and concept annotation errors. Thus, deriving efficient editable CBMs without retraining from scratch remains a challenge, particularly in large-scale applications. To address these challenges, we propose Editable Concept Bottleneck Models (ECBMs). Specifically, ECBMs support three different levels of data removal: concept-label-level, concept-level, and data-level. ECBMs enjoy mathematically rigorous closedform approximations derived from influence functions that obviate the need for retraining. Experimental results demonstrate the efficiency and adaptability of our ECBMs, affirming their practical value in CBMs. Code is available on https: //github.com/kaustpradalab/ECBM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Towards Multi-dimensional Explanation Alignment for Medical ClassificationLijie Hu, Songning Lai, Wenshuo Chen, Hongru Xiao et al.NeurIPS 2024 · 8 citations
- Semi-Supervised Concept Bottleneck ModelsLijie Hu, Tianhao Huang, Huanyi Xie, Xilin Gong et al.ICCV 2025 · 4 citations
- Editable XAI: Toward Bidirectional Human-AI Alignment with Co-Editable Explanations of Interpretable AttributesHaoyang Chen, Jingwen Bai, Fang Tian, Brian Y. LimCHI 2026 · 2 citations
- Debugging Concept Bottleneck Models through Removal and RetrainingEric Enouen, Sainyam GalhotraICLR 2026 · 2 citations
- Partially Shared Concept Bottleneck ModelsDelong Zhao, Qiang Huang, Di Yan, Yiqun Sun et al.AAAI 2026 · 2 citations
Builds on16
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
- Machine Unlearning for Random ForestsJonathan Brophy, Daniel LowdICML 2021 · 222 citations
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui et al.ICLR 2024 · 146 citations
Related papers
- Post-hoc Concept Bottleneck ModelsMert Yüksekgönül, Maggie Wang, James ZouICLR 2023 · 37 citations
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 163 citations
- Flexible Concept Bottleneck ModelXingbo Du, Qiantong Dou, Lei Fan, Rui ZhangAAAI 2026
- Learning to Receive Help: Intervention-Aware Concept Embedding ModelsMateo Espinosa Zarlenga, Katie Collins, Krishnamurthy Dvijotham, Adrian Weller et al.NeurIPS 2023 · 56 citations
- A Closer Look at the Intervention Procedure of Concept Bottleneck ModelsSungbin Shin, Yohan Jo, Sungsoo Ahn, Namhoon LeeICML 2023 · 59 citations
