Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts
Mateo Espinosa Zarlenga, Gabriele Dominici, Pietro Barbiero, Zohreh Shams, Mateja Jamnik
Abstract
In this paper, we investigate how concept-based models (CMs) respond to out-of-distribution (OOD) inputs. CMs are interpretable neural architectures that first predict a set of high-level concepts (e.g., stripes, black) and then predict a task label from those concepts. In particular, we study the impact of concept interventions (i.e., operations where a human expert corrects a CM's mispredicted concepts at test time) on CMs' task predictions when inputs are OOD. Our analysis reveals a weakness in current state-ofthe-art CMs, which we term leakage poisoning, that prevents them from properly improving their accuracy when intervened on for OOD inputs. To address this, we introduce MixCEM, a new CM that learns to dynamically exploit leaked information missing from its concepts only when this information is in-distribution. Our results across tasks with and without complete sets of concept annotations demonstrate that MixCEMs outperform strong baselines by significantly improving their accuracy for both in-distribution and OOD samples in the presence and absence of concept interventions. Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts C o n c e p t In te rv e n ti o n Concept Bottleneck Label Predictor Concept Predictor Input Human Expert
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Hierarchical Concept-based Interpretable ModelsOscar Hill, Mateo Espinosa Zarlenga, Mateja JamnikICLR 2026 · 3 citations
- Mixture of Concept Bottleneck ExpertsFrancesco De Santis, Gabriele Ciravegna, Giovanni De Felice, Arianna Casanova et al.ICML 2026 · 2 citations
- Partially Shared Concept Bottleneck ModelsDelong Zhao, Qiang Huang, Di Yan, Yiqun Sun et al.AAAI 2026 · 2 citations
- CB-SLICE: Concept-Based Interpretable Error Slice DiscoveryYael Konforti, Mateo Espinosa Zarlenga, Elaf Almahmoud, Mateja JamnikICML 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 163 citations
- Probabilistic Concept Bottleneck ModelsEunji Kim, Dahuin Jung, Sangha Park, Siwon Kim et al.ICML 2023 · 108 citations
Related papers
- Learning to Receive Help: Intervention-Aware Concept Embedding ModelsMateo Espinosa Zarlenga, Katie Collins, Krishnamurthy Dvijotham, Adrian Weller et al.NeurIPS 2023 · 56 citations
- A Closer Look at the Intervention Procedure of Concept Bottleneck ModelsSungbin Shin, Yohan Jo, Sungsoo Ahn, Namhoon LeeICML 2023 · 59 citations
- Interactive Concept Bottleneck ModelsKushal Chauhan, Rishabh Tiwari, Jan Freyberg, Pradeep Shenoy et al.AAAI 2023 · 91 citations
- Learning to Intervene on Concept BottlenecksDavid Steinmann, Wolfgang Stammer, Felix Friedrich, Kristian KerstingICML 2024 · 32 citations
- Concept Embedding Models: Beyond the Accuracy-Explainability Trade-OffMateo Espinosa Zarlenga, Pietro Barbiero, Gabriele Ciravegna, Giuseppe Marra et al.NeurIPS 2022
