An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations
Seonghwan Park, Jueun Mun, Donghyun Oh, Namhoon Lee
Abstract
Concept bottleneck models (CBMs) ensure interpretability by decomposing predictions into human interpretable concepts. Yet the annotations used for training CBMs that enable this transparency are often noisy, and the impact of such corruption is not well understood. In this study, we present the first systematic study of noise in CBMs and show that even moderate corruption simultaneously impairs prediction performance, interpretability, and the intervention effectiveness. Our analysis identifies a susceptible subset of concepts whose accuracy declines far more than the average gap between noisy and clean supervision and whose corruption accounts for most performance loss. To mitigate this vulnerability we propose a two-stage framework. During training, sharpness-aware minimization stabilizes the learning of noise-sensitive concepts. During inference, where clean labels are unavailable, we rank concepts by predictive entropy and correct only the most uncertain ones, using uncertainty as a proxy for susceptibility. Theoretical analysis and extensive ablations elucidate why sharpness-aware training confers robustness and why uncertainty reliably identifies susceptible concepts, providing a principled basis that preserves both interpretability and resilience in the presence of noise. 0.0 0.1 0.2 0.3 0.4 Noise Rate 5 31 57 83 Task Acc. (%) 0.0 0.1 0.2 0.3 0.4 Noise Rate 5 31 57 83 Task Acc. (%) Concept Target Concept + Target 0.0 0.1 0.2 0.3 0.4 Noise Rate 85 89 93 97 Concept Acc. (%) 0.0 0.1 0.2 0.3 0.4 Noise Rate 59 67 75 83 Concept Align. (%) 0.0 0.1 0.2 0.3 0.4 Noise Rate 43 58 73 88 Task Acc. (%) (a) Task accuracy 0.0 0.1 0.2 0.3 0.4 Noise Rate 37 53 69 85 Task Acc. (%) Concept Target Concept + Target (b) Source of degradation 0.0 0.1 0.2 0.3 0.4 Noise Rate 74 76 78 80 Concept Acc. (%) (c) Concept accuracy 0.0 0.1 0.2 0.3 0.4 Noise Rate 68 71 74 77 Concept Align. (%) (d) Concept alignment Blue Upperparts ( = 0.0) Concept Inactivated Concept Activated (a) γ = 0.0 Blue Upperparts ( = 0.2) Concept Inactivated Concept Activated (b) γ = 0.2 Blue Upperparts ( = 0.4) Concept Inactivated Concept Activated
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9b6760d-38ae-4ba7-85a3-38aab01432bdCited by top-tier papers1
Ask how each one uses itBuilds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
Related papers
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 163 citations
- There Was Never a Bottleneck in Concept Bottleneck ModelsAntonio Almudévar, José Miguel Hernández-Lobato, Alfonso OrtegaICLR 2026 · 9 citations
- A Closer Look at the Intervention Procedure of Concept Bottleneck ModelsSungbin Shin, Yohan Jo, Sungsoo Ahn, Namhoon LeeICML 2023 · 59 citations
- Stochastic Concept Bottleneck ModelsMoritz Vandenhirtz, Sonia Laguna, Ricards Marcinkevics, Julia E. VogtNeurIPS 2024 · 56 citations
- Concepts' Information Bottleneck ModelsKarim Galliamov, Syed Muhammad Ahsan Raza Kazmi, Adil Khan, Adín Ramírez RiveraICLR 2026 · 2 citations
