Classification with Conceptual Safeguards
Hailey Joren, Charles T. Marx, Berk Ustun
摘要
We propose a new approach to promote safety in classification tasks with established concepts. Our approach -- called a conceptual safeguard -- acts as a verification layer for models that predict a target outcome by first predicting the presence of intermediate concepts. Given this architecture, a safeguard ensures that a model meets a minimal level of accuracy by abstaining from uncertain predictions. In contrast to a standard selective classifier, a safeguard provides an avenue to improve coverage by allowing a human to confirm the presence of uncertain concepts on instances on which it abstains. We develop methods to build safeguards that maximize coverage without compromising safety, namely techniques to propagate the uncertainty in concept predictions and to flag salient concepts for human review. We benchmark our approach on a collection of real-world and synthetic datasets, showing that it can improve performance and coverage in deep learning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Concept Bottleneck Large Language ModelsChung-En Sun, Tuomas P. Oikarinen, Berk Ustun, Tsui-Wei WengICLR 2025
- Sufficient Context: A New Lens on Retrieval Augmented Generation SystemsHailey Joren, Jianyi Zhang, Chun-Sung Ferng, Da-Cheng Juan 等ICLR 2025
- Selective Preference AggregationShreyas Kadekodi, Hayden McTavish, Berk UstunICML 2025
- Feature Responsiveness Scores: Model-Agnostic Explanations for RecourseSeung Hyun Cheon, Anneke Wernerfelt, Sorelle A. Friedler, Berk UstunICLR 2025
它引用的顶会 Paper7
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li 等NeurIPS 2020 · 被引用 390 次
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 被引用 163 次
- Classification with Rejection Based on Cost-sensitive ClassificationNontawat Charoenphakdee, Zhenghang Cui, Yivan Zhang, Masashi SugiyamaICML 2021 · 被引用 78 次
- Selective Classification Can Magnify Disparities Across GroupsErik Jones, Shiori Sagawa, Pang Wei Koh, Ananya Kumar 等ICLR 2021 · 被引用 50 次
相关 Paper
- Confidence-aware Contrastive Learning for Selective ClassificationYu-Chang Wu, Shen-Huan Lyu, Haopu Shang, Xiangyu Wang 等ICML 2024 · 被引用 9 次
- Training Private Models That Know What They Don't KnowStephan Rabanser, Anvith Thudi, Abhradeep Guha Thakurta, Krishnamurthy Dvijotham 等NeurIPS 2023 · 被引用 10 次
- DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image DiagnosticsHong Hai Nguyen, Duong Bach, Nam Phan, Cuong V. Nguyen 等AAAI 2026
- Hierarchical Selective ClassificationShani Goren, Ido Galil, Ran El-YanivNeurIPS 2024 · 被引用 16 次
- Training Uncertainty-Aware Classifiers with Conformalized Deep LearningBat-Sheva Einbinder, Yaniv Romano, Matteo Sesia, Yanfei ZhouNeurIPS 2022 · 被引用 84 次
