A Closer Look at the Intervention Procedure of Concept Bottleneck Models
Sungbin Shin, Yohan Jo, Sungsoo Ahn, Namhoon Lee
Abstract
Concept bottleneck models (CBMs) are a class of interpretable neural network models that predict the target response of a given input based on its high-level concepts. Unlike the standard end-to-end models, CBMs enable domain experts to intervene on the predicted concepts and rectify any mistakes at test time, so that more accurate task predictions can be made at the end. While such intervenability provides a powerful avenue of control, many aspects of the intervention procedure remain rather unexplored. In this work, we develop various ways of selecting intervening concepts to improve the intervention effectiveness and conduct an array of in-depth analyses as to how they evolve under different circumstances. Specifically, we find that an informed intervention strategy can reduce the task error more than ten times compared to the current baseline under the same amount of intervention counts in realistic settings, and yet, this can vary quite significantly when taking into account different intervention granularity. We verify our findings through comprehensive evaluations, not only on the standard real datasets, but also on synthetic datasets that we generate based on a set of different causal graphs. We further discover some major pitfalls of the current practices which, without a proper addressing, raise concerns on reliability and fairness of the intervention procedure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2316d6fc-6a17-4106-98df-319cfd9dda36Cited by top-tier papers21
- Probabilistic Concept Bottleneck ModelsEunji Kim, Dahuin Jung, Sangha Park, Siwon Kim et al.ICML 2023 · 108 citations
- Learning to Receive Help: Intervention-Aware Concept Embedding ModelsMateo Espinosa Zarlenga, Katie Collins, Krishnamurthy Dvijotham, Adrian Weller et al.NeurIPS 2023 · 56 citations
- Stochastic Concept Bottleneck ModelsMoritz Vandenhirtz, Sonia Laguna, Ricards Marcinkevics, Julia E. VogtNeurIPS 2024 · 56 citations
- Auxiliary Losses for Learning Generalizable Concept-based ModelsIvaxi Sheth, Samira Ebrahimi KahouNeurIPS 2023 · 52 citations
- Beyond Concept Bottleneck Models: How to Make Black Boxes Intervenable?Sonia Laguna, Ricards Marcinkevics, Moritz Vandenhirtz, Julia E. VogtNeurIPS 2024 · 39 citations
Builds on4
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 163 citations
- Interactive Concept Bottleneck ModelsKushal Chauhan, Rishabh Tiwari, Jan Freyberg, Pradeep Shenoy et al.AAAI 2023 · 91 citations
- Post-hoc Concept Bottleneck ModelsMert Yüksekgönül, Maggie Wang, James ZouICLR 2023 · 37 citations
- Debiasing Concept-based Explanations with Causal AnalysisMohammad Taha Bahadori, David HeckermanICLR 2021 · 8 citations
Related papers
- Learning to Intervene on Concept BottlenecksDavid Steinmann, Wolfgang Stammer, Felix Friedrich, Kristian KerstingICML 2024 · 32 citations
- Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate ExpertsAndrea Pugnana, Riccardo Massidda, Francesco Giannini, Pietro Barbiero et al.NeurIPS 2025 · 11 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Debugging Concept Bottleneck Models through Removal and RetrainingEric Enouen, Sainyam GalhotraICLR 2026 · 2 citations
- Prototype-Grounded Concept Models for Verifiable Concept AlignmentStefano Colamonaco, David Debot, Pietro Barbiero, Giuseppe MarraICML 2026
