Classification with Conceptual Safeguards
Hailey Joren, Charles T. Marx, Berk Ustun
Abstract
We propose a new approach to promote safety in classification tasks with established concepts. Our approach -- called a conceptual safeguard -- acts as a verification layer for models that predict a target outcome by first predicting the presence of intermediate concepts. Given this architecture, a safeguard ensures that a model meets a minimal level of accuracy by abstaining from uncertain predictions. In contrast to a standard selective classifier, a safeguard provides an avenue to improve coverage by allowing a human to confirm the presence of uncertain concepts on instances on which it abstains. We develop methods to build safeguards that maximize coverage without compromising safety, namely techniques to propagate the uncertainty in concept predictions and to flag salient concepts for human review. We benchmark our approach on a collection of real-world and synthetic datasets, showing that it can improve performance and coverage in deep learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5da310a5-2e48-4f9f-aea6-2d9fd6a200b0Cited by top-tier papers4
- Concept Bottleneck Large Language ModelsChung-En Sun, Tuomas P. Oikarinen, Berk Ustun, Tsui-Wei WengICLR 2025
- Sufficient Context: A New Lens on Retrieval Augmented Generation SystemsHailey Joren, Jianyi Zhang, Chun-Sung Ferng, Da-Cheng Juan et al.ICLR 2025
- Selective Preference AggregationShreyas Kadekodi, Hayden McTavish, Berk UstunICML 2025
- Feature Responsiveness Scores: Model-Agnostic Explanations for RecourseSeung Hyun Cheon, Anneke Wernerfelt, Sorelle A. Friedler, Berk UstunICLR 2025
Builds on7
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 163 citations
- Classification with Rejection Based on Cost-sensitive ClassificationNontawat Charoenphakdee, Zhenghang Cui, Yivan Zhang, Masashi SugiyamaICML 2021 · 78 citations
- Selective Classification Can Magnify Disparities Across GroupsErik Jones, Shiori Sagawa, Pang Wei Koh, Ananya Kumar et al.ICLR 2021 · 50 citations
Related papers
- Confidence-aware Contrastive Learning for Selective ClassificationYu-Chang Wu, Shen-Huan Lyu, Haopu Shang, Xiangyu Wang et al.ICML 2024 · 9 citations
- Training Private Models That Know What They Don't KnowStephan Rabanser, Anvith Thudi, Abhradeep Guha Thakurta, Krishnamurthy Dvijotham et al.NeurIPS 2023 · 10 citations
- DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image DiagnosticsHong Hai Nguyen, Duong Bach, Nam Phan, Cuong V. Nguyen et al.AAAI 2026
- Hierarchical Selective ClassificationShani Goren, Ido Galil, Ran El-YanivNeurIPS 2024 · 16 citations
- Training Uncertainty-Aware Classifiers with Conformalized Deep LearningBat-Sheva Einbinder, Yaniv Romano, Matteo Sesia, Yanfei ZhouNeurIPS 2022 · 84 citations
