Learning to Receive Help: Intervention-Aware Concept Embedding Models
Mateo Espinosa Zarlenga, Katie Collins, Krishnamurthy Dvijotham, Adrian Weller, Zohreh Shams, Mateja Jamnik
Abstract
Concept Bottleneck Models (CBMs) tackle the opacity of neural architectures by constructing and explaining their predictions using a set of high-level concepts. A special property of these models is that they permit concept interventions, wherein users can correct mispredicted concepts and thus improve the model's performance. Recent work, however, has shown that intervention efficacy can be highly dependent on the order in which concepts are intervened on and on the model's architecture and training hyperparameters. We argue that this is rooted in a CBM's lack of train-time incentives for the model to be appropriately receptive to concept interventions. To address this, we propose Intervention-aware Concept Embedding models (IntCEMs), a novel CBM-based architecture and training paradigm that improves a model's receptiveness to test-time interventions. Our model learns a concept intervention policy in an end-to-end fashion from where it can sample meaningful intervention trajectories at train-time. This conditions IntCEMs to effectively select and receive concept interventions when deployed at test-time. Our experiments show that IntCEMs significantly outperform state-of-the-art concept-interpretable models when provided with test-time concept interventions, demonstrating the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30292d9f-0568-436d-9934-933e4993fc03Cited by top-tier papers12
- Stochastic Concept Bottleneck ModelsMoritz Vandenhirtz, Sonia Laguna, Ricards Marcinkevics, Julia E. VogtNeurIPS 2024 · 56 citations
- Causally Reliable Concept Bottleneck ModelsGiovanni de Felice, Arianna Casanova Flores, Francesco De Santis, Silvia Santini et al.NeurIPS 2025 · 20 citations
- Shortcuts and Identifiability in Concept-based Models from a Neuro-Symbolic LensSamuele Bortolotti, Emanuele Marconato, Paolo Morettin, Andrea Passerini et al.NeurIPS 2025 · 17 citations
- Understanding Inter-Concept Relationships in Concept-Based ModelsNaveen Raman, Mateo Espinosa Zarlenga, Mateja JamnikICML 2024 · 12 citations
- Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate ExpertsAndrea Pugnana, Riccardo Massidda, Francesco Giannini, Pietro Barbiero et al.NeurIPS 2025 · 11 citations
Builds on12
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 163 citations
- Probabilistic Concept Bottleneck ModelsEunji Kim, Dahuin Jung, Sangha Park, Siwon Kim et al.ICML 2023 · 108 citations
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell et al.NeurIPS 2020 · 83 citations
Related papers
- A Closer Look at the Intervention Procedure of Concept Bottleneck ModelsSungbin Shin, Yohan Jo, Sungsoo Ahn, Namhoon LeeICML 2023 · 59 citations
- Concept Embedding Models: Beyond the Accuracy-Explainability Trade-OffMateo Espinosa Zarlenga, Pietro Barbiero, Gabriele Ciravegna, Giuseppe Marra et al.NeurIPS 2022
- Learning to Intervene on Concept BottlenecksDavid Steinmann, Wolfgang Stammer, Felix Friedrich, Kristian KerstingICML 2024 · 32 citations
- Avoiding Leakage Poisoning: Concept Interventions Under Distribution ShiftsMateo Espinosa Zarlenga, Gabriele Dominici, Pietro Barbiero, Zohreh Shams et al.ICML 2025
- Interactive Concept Bottleneck ModelsKushal Chauhan, Rishabh Tiwari, Jan Freyberg, Pradeep Shenoy et al.AAAI 2023 · 91 citations
