Disentangling Safe and Unsafe Image Corruptions via Anisotropy and Locality
Ramchandran Muthukumar, Ambar Pal, Jeremias Sulam, René Vidal
Abstract
State-of-the-art machine learning systems are vulnerable to small perturbations to their input, where "small" is defined according to a threat model that assigns a positive threat to each perturbation. Most prior works define a task-agnostic, isotropic, and global threat, like the ℓ p norm, where the magnitude of the perturbation fully determines the degree of the threat and neither the direction of the attack nor its position in space matter. However, common corruptions in computer vision, such as blur, compression, or occlusions, are not well captured by such threat models. This paper proposes a novel threat model called Projected Displacement (PD) to study robustness beyond existing isotropic and global threat models. The proposed threat model measures the threat of a perturbation via its alignment with unsafe directions, defined as directions in the input space along which a perturbation of sufficient magnitude changes the ground truth class label. Unsafe directions are identified locally for each input based on observed training data. In this way, the PD-threat model exhibits anisotropy and locality. Experiments on Imagenet-1k data indicate that, for any input, the set of perturbations with small PD threat includes safe perturbations of large ℓ p norm that preserve the true label, such as noise, blur and compression, while simultaneously excluding unsafe perturbations that alter the true label. Unlike perceptual threat models based on embeddings of large-vision models, the PD-threat model can be readily computed for arbitrary classification tasks without pre-training or finetuning. Further additional task information such as sensitivity to image regions or concept hierarchies can be easily integrated into the assessment of threat and thus the PD threat model presents practitioners with a flexible, task-driven threat specification that alleviates the limitations of ℓ p -threat models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86e284c5-4767-4f67-bad9-21ee6d1fbcf5Builds on16
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Adversarial Sensor Attack on LiDAR-based Perception in Autonomous DrivingYulong Cao, Chaowei Xiao, Benjamin Cyr, Yimeng Zhou et al.CCS 2019 · 626 citations
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 451 citations
- DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic DataStephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai et al.NeurIPS 2023 · 413 citations
Related papers
- Towards Verifying Robustness of Neural Networks Against A Family of Semantic PerturbationsJeet Mohapatra, Tsui-Wei Weng, Pin-Yu Chen, Sijia Liu et al.CVPR 2020
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 597 citations
- Perceptual Adversarial Robustness: Defense Against Unseen Threat ModelsCassidy Laidlaw, Sahil Singla, Soheil FeiziICLR 2021 · 217 citations
- GSmooth: Certified Robustness against Semantic Transformations via Generalized Randomized SmoothingZhongkai Hao, Chengyang Ying, Yinpeng Dong, Hang Su et al.ICML 2022 · 27 citations
- Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold RobustnessSuraj Srinivas, Sebastian Bordt, Himabindu LakkarajuNeurIPS 2023 · 24 citations
