PROSAC: Provably Safe Certification for Machine Learning Models under Adversarial Attacks
Chen Feng, Ziquan Liu, Zhuo Zhi, Ilija Bogunovic, Carsten Gerner-Beuerle, Miguel Rodrigues
摘要
It is widely known that state-of-the-art machine learning models, including vision and language models, can be seriously compromised by adversarial perturbations. It is therefore increasingly relevant to develop capabilities to certify their performance in the presence of the most effective adversarial attacks. Our paper offers a new approach to certify the performance of machine learning models in the presence of adversarial attacks with population level risk guarantees. In particular, we introduce the notion of (α,ζ)-safe machine learning model. We propose a hypothesis testing procedure, based on the availability of a calibration set, to derive statistical guarantees providing that the probability of declaring that the adversarial (population) risk of a machine learning model is less than α (i.e. the model is safe), while the model is in fact unsafe (i.e. the model adversarial population risk is higher than α), is less than ζ. We also propose Bayesian optimization algorithms to determine efficiently whether a machine learning model is (α,ζ)-safe in the presence of an adversarial attack, along with statistical guarantees. We apply our framework to a range of machine learning models - including various sizes of vision Transformer (ViT) and ResNet models - impaired by a variety of adversarial attacks, such as PGDAttack, MomentumAttack, GenAttack and BanditAttack, to illustrate the operation of our approach. Importantly, we show that ViT's are generally more robust to adversarial attacks than ResNets, and large models are generally more robust than smaller models. Our approach goes beyond existing empirical adversarial risk-based certification guarantees. It formulates rigorous (and provable) performance guarantees that can be used to satisfy regulatory requirements mandating the use of state-of-the-art technical tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle 等ICLR 2026 · 被引用 24 次
- Rethinking LoRA for Privacy-Preserving Federated Learning in Large ModelsJin Liu, Yinbin Miao, Ning Xi, Junkang LiuICLR 2026 · 被引用 9 次
- GenSR: Symbolic regression based on equation generative spaceQian Li, Yuxiao Hu, Juncheng Liu, Yuntian ChenICLR 2026 · 被引用 7 次
- Attribution-Guided Model Rectification of Unreliable Neural Network BehaviorsPeiyu Yang, Naveed Akhtar, Jiantong Jiang, Ajmal MianCVPR 2026 · 被引用 4 次
- Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar DiagnosisChen Feng, Zhuo Zhi, Zhao Huang, Jiawei Ge 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper14
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu 等S&P 2019 · 被引用 1,022 次
- Globally-Robust Neural NetworksKlas Leino, Zifan Wang, Matt FredriksonICML 2021 · 被引用 150 次
- Orthogonalizing Convolutional Layers with the Cayley TransformAsher Trockman, J. Zico KolterICLR 2021 · 被引用 137 次
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
相关 Paper
- CC-CERT: A Probabilistic Approach to Certify General Robustness of Neural NetworksMikhail Pautov, Nurislam Tursynbek, Marina Munkhoeva, Nikita Muravev 等AAAI 2022 · 被引用 27 次
- Improving Robust Generalization by Direct PAC-Bayesian Bound MinimizationZifan Wang, Nan Ding, Tomer Levinboim, Xi Chen 等CVPR 2023
- Support is All You Need for Certified VAE TrainingChangming Xu, Debangshu Banerjee, Deepak Vasisht, Gagandeep SinghICLR 2025
- Towards Practical Certifiable Patch Defense with Vision TransformerZhaoyu Chen, Bo Li, Jianghe Xu, Shuang Wu 等CVPR 2022 · 被引用 60 次
- Probabilistic Robustness Certificates against Adversarial AttacksSara Taheri, Majid ZamaniICML 2026
