Reverse Engineering of Imperceptible Adversarial Image Perturbations
Yifan Gong, Yuguang Yao, Yize Li, Yimeng Zhang, Xiaoming Liu, Xue Lin, Sijia Liu
Abstract
It has been well recognized that neural network based image classifiers are easily fooled by images with tiny perturbations crafted by an adversary. There has been a vast volume of research to generate and defend such adversarial attacks. However, the following problem is left unexplored: How to reverse-engineer adversarial perturbations from an adversarial image? This leads to a new adversarial learning paradigm-Reverse Engineering of Deceptions (RED). If successful, RED allows us to estimate adversarial perturbations and recover the original images. However, carefully crafted, tiny adversarial perturbations are difficult to recover by optimizing a unilateral RED objective. For example, the pure image denoising method may overfit to minimizing the reconstruction error but hardly preserves the classification properties of the true adversarial perturbations. To tackle this challenge, we formalize the RED problem and identify a set of principles crucial to the RED approach design. Particularly, we find that prediction alignment and proper data augmentation (in terms of spatial transformations) are two criteria to achieve a generalizable RED approach. By integrating these RED principles with image denoising, we propose a new Class-Discriminative Denoising based RED framework, termed CDD-RED. Extensive experiments demonstrate the effectiveness of CDD-RED under different evaluation metrics (ranging from the pixel-level, prediction-level to the attribution-level alignment) and a variety of attack generation methods (e.g., FGSM, PGD, CW, AutoAttack, and adaptive attacks). Codes are available at link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang et al.NeurIPS 2024 · 200 citations
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu et al.ICLR 2026 · 15 citations
- Tracing Hyperparameter Dependencies for Model Parsing via Learnable Graph Pooling NetworkXiao Guo, Vishal Asnani, Sijia Liu, Xiaoming LiuNeurIPS 2024 · 13 citations
Builds on15
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
Related papers
- Post-Training Detection of Backdoor Attacks for Two-Class and Multi-Attack ScenariosZhen Xiang, David J. Miller, George KesidisICLR 2022 · 50 citations
- Red-Teaming Privacy-Protective Perturbations: Blind Face Restoration as an Attack StrategyZelin Li, Yifan Liu, Huimin Zeng, Yaokun Liu et al.WWW 2026
- Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion ModelsYixin Liu, Ruoxi Chen, Xun Chen, Lichao SunKDD 2026 · 3 citations
- RevINN: An End-to-End Invertible Neural Network for Reversible Adversarial Examples GenerationJielun Huang, Chi-Man Pun, Guoheng HuangCVPR 2026
- Delving into Deep Image Prior for Adversarial Defense: A Novel Reconstruction-based Defense FrameworkLi Ding, Yongwei Wang, Xin Ding, Kaiwen Yuan et al.ACM MM 2021 · 7 citations
