Towards Faithful XAI Evaluation via Generalization-Limited Backdoor Watermark
Mengxi Ya, Yiming Li, Tao Dai, Bin Wang, Yong Jiang, Shu-Tao Xia
摘要
Saliency-based representation visualization (SRV) (e.g., Grad-CAM) is one of the most classical and widely adopted explainable artificial intelligence (XAI) methods for its simplicity and efficiency. It can be used to interpret deep neural networks by locating saliency areas contributing the most to their predictions. However, it is difficult to automatically measure and evaluate the performance of SRV methods due to the lack of ground-truth salience areas of samples. In this paper, we revisit the backdoor-based SRV evaluation, which is currently the only feasible method to alleviate the previous problem. We first reveal its implementation limitations and unreliable nature due to the trigger generalization of existing backdoor watermarks. Given these findings, we propose a generalization-limited backdoor watermark (GLBW), based on which we design a more faithful XAI evaluation. Specifically, we formulate the training of watermarked DNNs as a min-max problem, where we find the 'worst' potential trigger (with the highest attack effectiveness and differences from the ground-truth trigger) via inner maximization and minimize its effects and the loss over benign and poisoned samples via outer minimization in each iteration. In particular, we design an adaptive optimization method to find desired potential triggers in each inner maximization. Extensive experiments on benchmark datasets are conducted, verifying the effectiveness of our generalization-limited watermark. Our codes are available at https://github.com/yamengxi/GLBW .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Towards Reliable and Efficient Backdoor Trigger Inversion via Decoupling Benign FeaturesXiong Xu, Kunzhe Huang, Yiming Li, Zhan Qin 等ICLR 2024 · 被引用 59 次
- IBD-PSC: Input-level Backdoor Detection via Parameter-oriented Scaling ConsistencyLinshan Hou, Ruili Feng, Zhongyun Hua, Wei Luo 等ICML 2024 · 被引用 52 次
- SABRE-FL: Selective and Accurate Backdoor Rejection for Federated Prompt LearningMomin Ahmad Khan, Yasra Chandio, Fatima M. AnwarICLR 2026 · 被引用 2 次
- Nearest is Not Dearest: Towards Practical Defense Against Quantization-Conditioned Backdoor AttacksBoheng Li, Yishuo Cai, Haowei Li, Feng Xue 等CVPR 2024
- BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIPJiawang Bai, Kuofeng Gao, Shaobo Min, Shu-Tao Xia 等CVPR 2024
它引用的顶会 Paper18
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Backdoor Attack with Imperceptible Input and Latent ModificationKhoa D. Doan, Yingjie Lao, Ping LiNeurIPS 2021 · 被引用 179 次
- Narcissus: A Practical Clean-Label Backdoor Attack with Limited InformationYi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu 等CCS 2023 · 被引用 170 次
- SynFace: Face Recognition with Synthetic DataHaibo Qiu, Baosheng Yu, Dihong Gong, Zhifeng Li 等ICCV 2021 · 被引用 162 次
- Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright ProtectionYiming Li, Yang Bai, Yong Jiang, Yong Yang 等NeurIPS 2022 · 被引用 161 次
相关 Paper
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 被引用 62 次
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 被引用 22 次
- Xplain: Analyzing Invisible Correlations in Model ExplanationKavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza SadeghiUSENIX Security 2024
- Watermarking Graph Neural Networks via Explanations for Ownership ProtectionJane Downer, Yingdan Shi, Ziyan Liu, Ren Wang 等ICML 2026 · 被引用 3 次
- Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature AttributionShuo Shao, Yiming Li, Hongwei Yao, Yiling He 等NDSS 2025
