Adversarial Counterfactual Visual Explanations
Guillaume Jeanneret, Loïc Simon, Frédéric Jurie
摘要
Counterfactual explanations and adversarial attacks have a related goal: flipping output labels with minimal perturbations regardless of their characteristics. Yet, adversarial attacks cannot be used directly in a counterfactual explanation perspective, as such perturbations are perceived as noise and not as actionable and understandable image modifications. Building on the robust learning literature, this paper proposes an elegant method to turn adversarial attacks into semantically meaningful perturbations, without modifying the classifiers to explain. The proposed approach hypothesizes that Denoising Diffusion Probabilistic Models are excellent regularizers for avoiding highfrequency and out-of-distribution perturbations when generating adversarial attacks. The paper's key idea is to build attacks through a diffusion model to polish them. This allows studying the target model regardless of its robustification level. Extensive experimentation shows the advantages of our counterfactual explanation approach over current State-of-the-Art in multiple testbeds.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual ImagesViraj Prabhu, Sriram Yenamandra, Prithvijit Chattopadhyay, Judy HoffmanNeurIPS 2023 · 被引用 59 次
- BlurGuard: A Simple Approach for Robustifying Image Protection Against AI-Powered EditingJinsu Kim, Yunhun Nam, Minseon Kim, Sangpil Kim 等NeurIPS 2025 · 被引用 7 次
- Impact of Explanation Techniques and Representations on Users' Comprehension and Confidence in Explainable AIJulien Delaunay, Luis Galárraga, Christine Largouët, Niels van BerkelCSCW 2025 · 被引用 6 次
- V-CECE: Visual Counterfactual Explanations via Conceptual EditsNikolaos Spanos, Maria Lymperaiou, Giorgos Filandrianos, Konstantinos Thomas 等NeurIPS 2025 · 被引用 4 次
- Causality-aligned Prompt Learning via Diffusion-based Counterfactual GenerationXinshu Li, Ruoyu Wang, Erdun Gao, Mingming Gong 等ACM MM 2025 · 被引用 3 次
它引用的顶会 Paper23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu 等CVPR 2022 · 被引用 1,425 次
相关 Paper
- Diffusion Visual Counterfactual ExplanationsMaximilian Augustin, Valentyn Boreiko, Francesco Croce, Matthias HeinNeurIPS 2022 · 被引用 124 次
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
- AdvDiffuser: Natural Adversarial Example Synthesis with Diffusion ModelsXinquan Chen, Xitong Gao, Juanjuan Zhao, Kejiang Ye 等ICCV 2023 · 被引用 94 次
- Explaining Classifiers Using Adversarial Perturbations on the Perceptual BallAndrew Elliott, Stephen Law, Chris RussellCVPR 2021
- NatADiff: Adversarial Boundary Guidance for Natural Adversarial DiffusionMax Collins, Jordan Vice, Tim French, Ajmal MianICLR 2026 · 被引用 6 次
