Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
Stephan Rabanser, Ali Shahin Shamsabadi, Olive Franzese, Xiao Wang, Adrian Weller, Nicolas Papernot
摘要
Cautious predictions -where a machine learning model abstains when uncertain -are crucial for limiting harmful errors in safety-critical applications. In this work, we identify a novel threat: a dishonest institution can exploit these mechanisms to discriminate or unjustly deny services under the guise of uncertainty. We demonstrate the practicality of this threat by introducing an uncertainty-inducing attack called Mirage, which deliberately reduces confidence in targeted input regions, thereby covertly disadvantaging specific individuals. At the same time, Mirage maintains high predictive performance across all data points. To counter this threat, we propose Confidential Guardian, a framework that analyzes calibration metrics on a reference dataset to detect artificially suppressed confidence. Additionally, it employs zero-knowledge proofs of verified inference to ensure that reported confidence scores genuinely originate from the deployed model. This prevents the provider from fabricating arbitrary model confidence values while protecting the model's proprietary details. Our results confirm that Confidential Guardian effectively prevents the misuse of cautious predictions, providing verifiable assurances that abstention reflects genuine model uncertainty rather than malicious intent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 被引用 789 次
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- Scaling Out-of-Distribution Detection for Real-World SettingsDan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou 等ICML 2022 · 被引用 653 次
- Wolverine: Fast, Scalable, and Communication-Efficient Zero-Knowledge Proofs for Boolean and Arithmetic CircuitsChenkai Weng, Kang Yang, Jonathan Katz, Xiao WangS&P 2021 · 被引用 205 次
相关 Paper
- On Calibration of LLM-based Guard Models for Reliable Content ModerationHongfu Liu, Hengguan Huang, Xiangming Gu, Hao Wang 等ICLR 2025 · 被引用 1 次
- With False Friends Like These, Who Can Notice Mistakes?Lue Tao, Lei Feng, Jinfeng Yi, Songcan ChenAAAI 2022 · 被引用 6 次
- Not All Features Are Equal: Discovering Essential Features for Preserving Prediction PrivacyFatemehsadat Mireshghallah, Mohammadkazem Taram, Ali Jalali, Ahmed Taha Elthakeb 等WWW 2021 · 被引用 59 次
- A Method to Facilitate Membership Inference Attacks in Deep Learning ModelsZitao Chen, Karthik PattabiramanNDSS 2025
- The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal GuardrailsShuo Shi, Rui Yin, Naen Xu, Jiahao Chen 等KDD 2026 · 被引用 1 次
