PROSAC: Provably Safe Certification for Machine Learning Models under Adversarial Attacks
Chen Feng, Ziquan Liu, Zhuo Zhi, Ilija Bogunovic, Carsten Gerner-Beuerle, Miguel Rodrigues
Abstract
It is widely known that state-of-the-art machine learning models, including vision and language models, can be seriously compromised by adversarial perturbations. It is therefore increasingly relevant to develop capabilities to certify their performance in the presence of the most effective adversarial attacks. Our paper offers a new approach to certify the performance of machine learning models in the presence of adversarial attacks with population level risk guarantees. In particular, we introduce the notion of (α,ζ)-safe machine learning model. We propose a hypothesis testing procedure, based on the availability of a calibration set, to derive statistical guarantees providing that the probability of declaring that the adversarial (population) risk of a machine learning model is less than α (i.e. the model is safe), while the model is in fact unsafe (i.e. the model adversarial population risk is higher than α), is less than ζ. We also propose Bayesian optimization algorithms to determine efficiently whether a machine learning model is (α,ζ)-safe in the presence of an adversarial attack, along with statistical guarantees. We apply our framework to a range of machine learning models - including various sizes of vision Transformer (ViT) and ResNet models - impaired by a variety of adversarial attacks, such as PGDAttack, MomentumAttack, GenAttack and BanditAttack, to illustrate the operation of our approach. Importantly, we show that ViT's are generally more robust to adversarial attacks than ResNets, and large models are generally more robust than smaller models. Our approach goes beyond existing empirical adversarial risk-based certification guarantees. It formulates rigorous (and provable) performance guarantees that can be used to satisfy regulatory requirements mandating the use of state-of-the-art technical tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 409e15c8-5138-4a0c-a5ea-411212227eb1Cited by top-tier papers8
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle et al.ICLR 2026 · 24 citations
- Rethinking LoRA for Privacy-Preserving Federated Learning in Large ModelsJin Liu, Yinbin Miao, Ning Xi, Junkang LiuICLR 2026 · 9 citations
- GenSR: Symbolic regression based on equation generative spaceQian Li, Yuxiao Hu, Juncheng Liu, Yuntian ChenICLR 2026 · 7 citations
- Attribution-Guided Model Rectification of Unreliable Neural Network BehaviorsPeiyu Yang, Naveed Akhtar, Jiantong Jiang, Ajmal MianCVPR 2026 · 4 citations
- Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar DiagnosisChen Feng, Zhuo Zhi, Zhao Huang, Jiawei Ge et al.CVPR 2026 · 4 citations
Builds on14
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Globally-Robust Neural NetworksKlas Leino, Zifan Wang, Matt FredriksonICML 2021 · 150 citations
- Orthogonalizing Convolutional Layers with the Cayley TransformAsher Trockman, J. Zico KolterICLR 2021 · 137 citations
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala et al.ICLR 2024 · 132 citations
Related papers
- CC-CERT: A Probabilistic Approach to Certify General Robustness of Neural NetworksMikhail Pautov, Nurislam Tursynbek, Marina Munkhoeva, Nikita Muravev et al.AAAI 2022 · 27 citations
- Improving Robust Generalization by Direct PAC-Bayesian Bound MinimizationZifan Wang, Nan Ding, Tomer Levinboim, Xi Chen et al.CVPR 2023
- Support is All You Need for Certified VAE TrainingChangming Xu, Debangshu Banerjee, Deepak Vasisht, Gagandeep SinghICLR 2025
- Towards Practical Certifiable Patch Defense with Vision TransformerZhaoyu Chen, Bo Li, Jianghe Xu, Shuang Wu et al.CVPR 2022 · 60 citations
- Probabilistic Robustness Certificates against Adversarial AttacksSara Taheri, Majid ZamaniICML 2026
