Lune

ICML2026Top-tier venue

Probabilistic Robustness Certificates against Adversarial Attacks

Sara Taheri, Majid Zamani

2026Year

Abstract

The growing use of machine learning in safetycritical settings increases vulnerability to adversarial attacks. Existing defense mechanisms typically either lack formal guarantees or depend on restrictive assumptions about the model family, the threat model, or the perturbation budget, and many only offer point-wise certification. Importantly, they often overlook the inherent stochasticity of modern training pipelines, which undermines their practical reliability. In this work, we introduce a probabilistic framework that views gradient-based training as a discrete-time stochastic dynamical system and formulates adversarial robustness as a safety verification task. Using barrier certificates (BC), we derive sufficient conditions to probabilistically certify a robust radius against worst-case ℓ p -bounded perturbation, guaranteeing that the final model parameters remain within a safe set probabilistically. For tractable computation, we represent BCs with neural networks and obtain probably approximately correct (PAC) guarantees through a scenario convex problem. Our approach determines the maximum certified radius within which the trained model achieves probabilistic accuracy at a pre-specified confidence level. Experiments on MNIST, SVHN, and CIFAR-10 show that our framework offers robustness guarantees under stochastic training, while being model-agnostic and not requiring knowledge of the attack strategy.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6021f578-5536-4c0f-b8cd-24477380807f

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines