Enhancing Adversarial Robustness with Conformal Prediction: A Framework for Guaranteed Model Reliability
Jie Bao, Chuangyin Dang, Rui Luo, Hanwei Zhang, Zhixin Zhou
Abstract
As deep learning models are increasingly deployed in high-risk applications, robust defenses against adversarial attacks and reliable performance guarantees become paramount. Moreover, accuracy alone does not provide sufficient assurance or reliable uncertainty estimates for these models. This study advances adversarial training by leveraging principles from Conformal Prediction. Specifically, we develop an adversarial attack method, termed OPSA (OPtimal Size Attack), designed to reduce the efficiency of conformal prediction at any significance level by maximizing model uncertainty without requiring coverage guarantees. Correspondingly, we introduce OPSA-AT (Adversarial Training), a defense strategy that integrates OPSA within a novel conformal training paradigm. Experimental evaluations demonstrate that our OPSA attack method induces greater uncertainty compared to baseline approaches for various defenses. Conversely, our OPSA-AT defensive model significantly enhances robustness not only against OPSA but also other adversarial attacks, and maintains reliable prediction. Our findings highlight the effectiveness of this integrated approach for developing trustworthy and resilient deep learning models for safety-critical domains. Our code is available at https://github.com/bjbbbb/Enha ncing-Adversarial-Robustness-wit h-Conformal-Prediction .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9a3ed5e-e531-4c4a-83b9-d8611738d0b7Cited by top-tier papers6
- A Minimum Variance Path Principle for Accurate and Stable Score-Based Density Ratio EstimationWei Chen, Jiacheng Li, Shigui Li, Zhiqi Lin et al.ICLR 2026 · 4 citations
- Fast Conformal Prediction Using Conditional Interquantile IntervalsNaixin Guo, Rui Luo, Zhixin ZhouAAAI 2026 · 4 citations
- Enhancing Image-Conditional Coverage in Segmentation: Adaptive Thresholding via Differentiable Miscoverage LossRui Luo, Jie Bao, Xiaoyi Su, Wen Li et al.ICLR 2026
- Cost-Sensitive Conformal Training with Provably Controllable Learning BoundsXuesong Jia, Yuanjie Shi, Ziquan Liu, Yi Xu et al.AAAI 2026
- Conformity Score Averaging for ClassificationRui Luo, Zhixin ZhouICML 2025
Builds on17
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Classification with Valid and Adaptive CoverageYaniv Romano, Matteo Sesia, Emmanuel J. CandèsNeurIPS 2020 · 586 citations
- Attacks Which Do Not Kill Training Make Adversarial Learning StrongerJingfeng Zhang, Xilie Xu, Bo Han, Gang Niu et al.ICML 2020 · 452 citations
- Learning Optimal Conformal ClassifiersDavid Stutz, Krishnamurthy Dvijotham, Ali Taylan Cemgil, Arnaud DoucetICLR 2022 · 123 citations
Related papers
- The Pitfalls and Promise of Conformal Inference Under Adversarial AttacksZiquan Liu, Yufei Cui, Yan Yan, Yi Xu et al.ICML 2024 · 9 citations
- Direct Prediction Set Minimization via Bilevel Conformal Classifier TrainingYuanjie Shi, Hooman Shahrokhi, Xuesong Jia, Xiongzhi Chen et al.ICML 2025
- Provably Adversarially Robust Detection of Out-of-Distribution Data (Almost) for FreeAlexander Meinke, Julian Bitterwolf, Matthias HeinNeurIPS 2022 · 23 citations
- Data Poisoning Attacks against Conformal PredictionYangyi Li, Aobo Chen, Wei Qian, Chenxu Zhao et al.ICML 2024 · 10 citations
- Training Uncertainty-Aware Classifiers with Conformalized Deep LearningBat-Sheva Einbinder, Yaniv Romano, Matteo Sesia, Yanfei ZhouNeurIPS 2022 · 84 citations
