Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts
Mateo Espinosa Zarlenga, Gabriele Dominici, Pietro Barbiero, Zohreh Shams, Mateja Jamnik
摘要
In this paper, we investigate how concept-based models (CMs) respond to out-of-distribution (OOD) inputs. CMs are interpretable neural architectures that first predict a set of high-level concepts (e.g., stripes, black) and then predict a task label from those concepts. In particular, we study the impact of concept interventions (i.e., operations where a human expert corrects a CM's mispredicted concepts at test time) on CMs' task predictions when inputs are OOD. Our analysis reveals a weakness in current state-ofthe-art CMs, which we term leakage poisoning, that prevents them from properly improving their accuracy when intervened on for OOD inputs. To address this, we introduce MixCEM, a new CM that learns to dynamically exploit leaked information missing from its concepts only when this information is in-distribution. Our results across tasks with and without complete sets of concept annotations demonstrate that MixCEMs outperform strong baselines by significantly improving their accuracy for both in-distribution and OOD samples in the presence and absence of concept interventions. Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts C o n c e p t In te rv e n ti o n Concept Bottleneck Label Predictor Concept Predictor Input Human Expert
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Hierarchical Concept-based Interpretable ModelsOscar Hill, Mateo Espinosa Zarlenga, Mateja JamnikICLR 2026 · 被引用 3 次
- Mixture of Concept Bottleneck ExpertsFrancesco De Santis, Gabriele Ciravegna, Giovanni De Felice, Arianna Casanova 等ICML 2026 · 被引用 2 次
- Partially Shared Concept Bottleneck ModelsDelong Zhao, Qiang Huang, Di Yan, Yiqun Sun 等AAAI 2026 · 被引用 2 次
- CB-SLICE: Concept-Based Interpretable Error Slice DiscoveryYael Konforti, Mateo Espinosa Zarlenga, Elaf Almahmoud, Mateja JamnikICML 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li 等NeurIPS 2020 · 被引用 390 次
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 被引用 163 次
- Probabilistic Concept Bottleneck ModelsEunji Kim, Dahuin Jung, Sangha Park, Siwon Kim 等ICML 2023 · 被引用 108 次
相关 Paper
- Learning to Receive Help: Intervention-Aware Concept Embedding ModelsMateo Espinosa Zarlenga, Katie Collins, Krishnamurthy Dvijotham, Adrian Weller 等NeurIPS 2023 · 被引用 56 次
- A Closer Look at the Intervention Procedure of Concept Bottleneck ModelsSungbin Shin, Yohan Jo, Sungsoo Ahn, Namhoon LeeICML 2023 · 被引用 59 次
- Interactive Concept Bottleneck ModelsKushal Chauhan, Rishabh Tiwari, Jan Freyberg, Pradeep Shenoy 等AAAI 2023 · 被引用 91 次
- Learning to Intervene on Concept BottlenecksDavid Steinmann, Wolfgang Stammer, Felix Friedrich, Kristian KerstingICML 2024 · 被引用 32 次
- Concept Embedding Models: Beyond the Accuracy-Explainability Trade-OffMateo Espinosa Zarlenga, Pietro Barbiero, Gabriele Ciravegna, Giuseppe Marra 等NeurIPS 2022
