Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
Stephan Rabanser, Ali Shahin Shamsabadi, Olive Franzese, Xiao Wang, Adrian Weller, Nicolas Papernot
Abstract
Cautious predictions -where a machine learning model abstains when uncertain -are crucial for limiting harmful errors in safety-critical applications. In this work, we identify a novel threat: a dishonest institution can exploit these mechanisms to discriminate or unjustly deny services under the guise of uncertainty. We demonstrate the practicality of this threat by introducing an uncertainty-inducing attack called Mirage, which deliberately reduces confidence in targeted input regions, thereby covertly disadvantaging specific individuals. At the same time, Mirage maintains high predictive performance across all data points. To counter this threat, we propose Confidential Guardian, a framework that analyzes calibration metrics on a reference dataset to detect artificially suppressed confidence. Additionally, it employs zero-knowledge proofs of verified inference to ensure that reported confidence scores genuinely originate from the deployed model. This prevents the provider from fabricating arbitrary model confidence values while protecting the model's proprietary details. Our results confirm that Confidential Guardian effectively prevents the misuse of cautious predictions, providing verifiable assurances that abstention reflects genuine model uncertainty rather than malicious intent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dbb0ced-bd22-4171-b88b-3b1fb7c4d403Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 789 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Scaling Out-of-Distribution Detection for Real-World SettingsDan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou et al.ICML 2022 · 653 citations
- Wolverine: Fast, Scalable, and Communication-Efficient Zero-Knowledge Proofs for Boolean and Arithmetic CircuitsChenkai Weng, Kang Yang, Jonathan Katz, Xiao WangS&P 2021 · 205 citations
Related papers
- On Calibration of LLM-based Guard Models for Reliable Content ModerationHongfu Liu, Hengguan Huang, Xiangming Gu, Hao Wang et al.ICLR 2025 · 1 citation
- With False Friends Like These, Who Can Notice Mistakes?Lue Tao, Lei Feng, Jinfeng Yi, Songcan ChenAAAI 2022 · 6 citations
- Not All Features Are Equal: Discovering Essential Features for Preserving Prediction PrivacyFatemehsadat Mireshghallah, Mohammadkazem Taram, Ali Jalali, Ahmed Taha Elthakeb et al.WWW 2021 · 59 citations
- A Method to Facilitate Membership Inference Attacks in Deep Learning ModelsZitao Chen, Karthik PattabiramanNDSS 2025
- The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal GuardrailsShuo Shi, Rui Yin, Naen Xu, Jiahao Chen et al.KDD 2026 · 1 citation
