Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization
Emiliano Penaloza, Tianyue H. Zhang, Laurent Charlin, Mateo Espinosa Zarlenga
摘要
Concept Bottleneck Models (CBMs) propose to enhance the trustworthiness of AI systems by constraining their decisions on a set of humanunderstandable concepts. However, CBMs typically assume that datasets contains accurate concept labels-an assumption often violated in practice, which we show can significantly degrade performance (by 25% in some cases). To address this, we introduce the Concept Preference Optimization (CPO) objective, a new loss function based on Direct Preference Optimization, which effectively mitigates the negative impact of concept mislabeling on CBM performance. We provide an analysis on some key properties of the CPO objective showing it directly optimizes for the concept's posterior distribution, and contrast it against Binary Cross Entropy (BCE) where we show CPO is inherently less sensitive to concept noise. We empirically confirm our analysis finding that CPO consistently outperforms BCE in three real-world datasets with and without added label noise. We make our code available on Github 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Partially Shared Concept Bottleneck ModelsDelong Zhao, Qiang Huang, Di Yan, Yiqun Sun 等AAAI 2026 · 被引用 2 次
- CB-SLICE: Concept-Based Interpretable Error Slice DiscoveryYael Konforti, Mateo Espinosa Zarlenga, Elaf Almahmoud, Mateja JamnikICML 2026
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li 等NeurIPS 2020 · 被引用 390 次
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
相关 Paper
- Auxiliary Losses for Learning Generalizable Concept-based ModelsIvaxi Sheth, Samira Ebrahimi KahouNeurIPS 2023 · 被引用 52 次
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 被引用 163 次
- Credal Concept Bottleneck Models for Epistemic-Aleatoric Uncertainty DecompositionTanmoy Mukherjee, Thomas Bailleux, Pierre Marquis, Zied BouraouiACL 2026
- Semi-Supervised Concept Bottleneck ModelsLijie Hu, Tianhao Huang, Huanyi Xie, Xilin Gong 等ICCV 2025 · 被引用 4 次
- Concept Embedding Models: Beyond the Accuracy-Explainability Trade-OffMateo Espinosa Zarlenga, Pietro Barbiero, Gabriele Ciravegna, Giuseppe Marra 等NeurIPS 2022
