Detecting Adversarial Data by Probing Multiple Perturbations Using Expected Perturbation Score
Shuhai Zhang, Feng Liu, Jiahao Yang, Yifan Yang, Changsheng Li, Bo Han, Mingkui Tan
摘要
Adversarial detection aims to determine whether a given sample is an adversarial one based on the discrepancy between natural and adversarial distributions. Unfortunately, estimating or comparing two data distributions is extremely difficult, especially in high-dimension spaces. Recently, the gradient of log probability density (a.k.a., score) w.r.t. the sample is used as an alternative statistic to compute. However, we find that the score is sensitive in identifying adversarial samples due to insufficient information with one sample only. In this paper, we propose a new statistic called expected perturbation score (EPS), which is essentially the expected score of a sample after various perturbations. Specifically, to obtain adequate information regarding one sample, we perturb it by adding various noises to capture its multi-view observations. We theoretically prove that EPS is a proper statistic to compute the discrepancy between two samples under mild conditions. In practice, we can use a pre-trained diffusion model to estimate EPS for each sample. Last, we propose an EPS-based adversarial detection (EPS-AD) method, in which we develop EPS-based maximum mean discrepancy (MMD) as a metric to measure the discrepancy between the test sample and natural samples. We also prove that the EPS-based MMD between natural and adversarial samples is larger than that among natural samples. Extensive experiments show the superior adversarial detection performance of our EPS-AD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Physics-Driven Spatiotemporal Modeling for AI-Generated Video DetectionShuhai Zhang, Zihao Lian, Jiahao Yang, Daiyuan Li 等NeurIPS 2025 · 被引用 29 次
- Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean DiscrepancyShuhai Zhang, Yiliao Song, Jiahao Yang, Yuanqing Li 等ICLR 2024 · 被引用 19 次
- Improving Accuracy-robustness Trade-off via Pixel Reweighted Adversarial TrainingJiacheng Zhang, Feng Liu, Dawei Zhou, Jingfeng Zhang 等ICML 2024 · 被引用 9 次
- DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence TestingZhijian Zhou, Xunye Tian, Liuhua Peng, Chao Lei 等NeurIPS 2025 · 被引用 8 次
- Generative Model Inversion Through the Lens of the Manifold HypothesisXiong Peng, Bo Han, Fengfei Yu, Tongliang Liu 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Maximum Mean Discrepancy Test is Aware of Adversarial AttacksRuize Gao, Feng Liu, Jingfeng Zhang, Bo Han 等ICML 2021 · 被引用 77 次
- CASN: Class-Aware Score Network for Textual Adversarial DetectionRong Bao, Rui Zheng, Liang Ding, Qi Zhang 等ACL 2023 · 被引用 3 次
- One Stone, Two Birds: Enhancing Adversarial Defense Through the Lens of Distributional DiscrepancyJiacheng Zhang, Benjamin I. P. Rubinstein, Jingfeng Zhang, Feng LiuICML 2025
- MMD Guidance: Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy GuidanceMatina Mahdizadeh Sani, Nima Jamali, Mohammad Jalali, Farzan FarniaICML 2026 · 被引用 4 次
- Deep MMD Gradient Flow without adversarial trainingAlexandre Galashov, Valentin De Bortoli, Arthur GrettonICLR 2025 · 被引用 1 次
