Reverse Engineering of Imperceptible Adversarial Image Perturbations
Yifan Gong, Yuguang Yao, Yize Li, Yimeng Zhang, Xiaoming Liu, Xue Lin, Sijia Liu
摘要
It has been well recognized that neural network based image classifiers are easily fooled by images with tiny perturbations crafted by an adversary. There has been a vast volume of research to generate and defend such adversarial attacks. However, the following problem is left unexplored: How to reverse-engineer adversarial perturbations from an adversarial image? This leads to a new adversarial learning paradigm-Reverse Engineering of Deceptions (RED). If successful, RED allows us to estimate adversarial perturbations and recover the original images. However, carefully crafted, tiny adversarial perturbations are difficult to recover by optimizing a unilateral RED objective. For example, the pure image denoising method may overfit to minimizing the reconstruction error but hardly preserves the classification properties of the true adversarial perturbations. To tackle this challenge, we formalize the RED problem and identify a set of principles crucial to the RED approach design. Particularly, we find that prediction alignment and proper data augmentation (in terms of spatial transformations) are two criteria to achieve a generalizable RED approach. By integrating these RED principles with image denoising, we propose a new Class-Discriminative Denoising based RED framework, termed CDD-RED. Extensive experiments demonstrate the effectiveness of CDD-RED under different evaluation metrics (ranging from the pixel-level, prediction-level to the attribution-level alignment) and a variety of attack generation methods (e.g., FGSM, PGD, CW, AutoAttack, and adaptive attacks). Codes are available at link.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang 等NeurIPS 2024 · 被引用 200 次
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu 等ICLR 2026 · 被引用 15 次
- Tracing Hyperparameter Dependencies for Model Parsing via Learnable Graph Pooling NetworkXiao Guo, Vishal Asnani, Sijia Liu, Xiaoming LiuNeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper15
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
相关 Paper
- Post-Training Detection of Backdoor Attacks for Two-Class and Multi-Attack ScenariosZhen Xiang, David J. Miller, George KesidisICLR 2022 · 被引用 50 次
- Red-Teaming Privacy-Protective Perturbations: Blind Face Restoration as an Attack StrategyZelin Li, Yifan Liu, Huimin Zeng, Yaokun Liu 等WWW 2026
- Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion ModelsYixin Liu, Ruoxi Chen, Xun Chen, Lichao SunKDD 2026 · 被引用 3 次
- RevINN: An End-to-End Invertible Neural Network for Reversible Adversarial Examples GenerationJielun Huang, Chi-Man Pun, Guoheng HuangCVPR 2026
- Delving into Deep Image Prior for Adversarial Defense: A Novel Reconstruction-based Defense FrameworkLi Ding, Yongwei Wang, Xin Ding, Kaiwen Yuan 等ACM MM 2021 · 被引用 7 次
