Lune

NeurIPS2025顶会

An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations

Seonghwan Park, Jueun Mun, Donghyun Oh, Namhoon Lee

2025年份
10被引次数
1顶会引用

摘要

Concept bottleneck models (CBMs) ensure interpretability by decomposing predictions into human interpretable concepts. Yet the annotations used for training CBMs that enable this transparency are often noisy, and the impact of such corruption is not well understood. In this study, we present the first systematic study of noise in CBMs and show that even moderate corruption simultaneously impairs prediction performance, interpretability, and the intervention effectiveness. Our analysis identifies a susceptible subset of concepts whose accuracy declines far more than the average gap between noisy and clean supervision and whose corruption accounts for most performance loss. To mitigate this vulnerability we propose a two-stage framework. During training, sharpness-aware minimization stabilizes the learning of noise-sensitive concepts. During inference, where clean labels are unavailable, we rank concepts by predictive entropy and correct only the most uncertain ones, using uncertainty as a proxy for susceptibility. Theoretical analysis and extensive ablations elucidate why sharpness-aware training confers robustness and why uncertainty reliably identifies susceptible concepts, providing a principled basis that preserves both interpretability and resilience in the presence of noise. 0.0 0.1 0.2 0.3 0.4 Noise Rate 5 31 57 83 Task Acc. (%) 0.0 0.1 0.2 0.3 0.4 Noise Rate 5 31 57 83 Task Acc. (%) Concept Target Concept + Target 0.0 0.1 0.2 0.3 0.4 Noise Rate 85 89 93 97 Concept Acc. (%) 0.0 0.1 0.2 0.3 0.4 Noise Rate 59 67 75 83 Concept Align. (%) 0.0 0.1 0.2 0.3 0.4 Noise Rate 43 58 73 88 Task Acc. (%) (a) Task accuracy 0.0 0.1 0.2 0.3 0.4 Noise Rate 37 53 69 85 Task Acc. (%) Concept Target Concept + Target (b) Source of degradation 0.0 0.1 0.2 0.3 0.4 Noise Rate 74 76 78 80 Concept Acc. (%) (c) Concept accuracy 0.0 0.1 0.2 0.3 0.4 Noise Rate 68 71 74 77 Concept Align. (%) (d) Concept alignment Blue Upperparts ( = 0.0) Concept Inactivated Concept Activated (a) γ = 0.0 Blue Upperparts ( = 0.2) Concept Inactivated Concept Activated (b) γ = 0.2 Blue Upperparts ( = 0.4) Concept Inactivated Concept Activated

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖