USENIX Security2024Top-tier venue
Xplain: Analyzing Invisible Correlations in Model Explanation
Kavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza Sadeghi
Abstract
Explanation methods analyze the features in backdoored input data that contribute to model misclassification. However, current methods like path techniques struggle to detect backdoor patterns in adversarial situations. They fail to grasp the hidden associations of backdoor features with other input features, leading to misclassification. Additionally, they suffer from irrelevant data attribution, imprecise feature connections, baseline dependence, and vulnerability to the "saturation effect". To address these limitations, we propose Xplain. Our method aims to uncover hidden backdoor trigger patterns and the subtle relationships between backdoor features and other input objects, which are the main causes of model misclassification. Our algorithm improves existing path techniques by integrating an additional baseline into the Integrated Gradients (IG) formulation. This ensures that features selected in the baseline persist along the integration path, guaranteeing baseline independence. Additionally, we introduce quantitative noise to interpolate samples along the integration path, which reduces feature dependency and captures non-linear interactions. This approach effectively identifies the relevant features that significantly influence model predictions. Furthermore, Xplain proposes sensitivity analysis to enhance AI system resilience against backdoor attacks. This uncovers clear connections between the backdoor and other input data features, thus shedding light on relevant interactions. We thoroughly test the effectiveness of Xplain on the Imagenet and the multimodal domain of the Visual Question Answering dataset, showing its superiority over current path methods such as Integrated Gradient (IG), left-IG, Guided IG, and Adversarial Gradient Integration (AGI) techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61cce6ef-cbac-417f-be25-045461dc8127Builds on4
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- XRAI: Better Attributions Through RegionsAndrei Kapishnikov, Tolga Bolukbasi, Fernanda B. Viégas, Michael TerryICCV 2019 · 251 citations
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 62 citations
- Guided Integrated Gradients: An Adaptive Path Method for Removing NoiseAndrei Kapishnikov, Subhashini Venugopalan, Besim Avci, Ben Wedin et al.CVPR 2021
Related papers
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 22 citations
- Integrated Decision Gradients: Compute Your Attributions Where the Model Makes Its DecisionChase Walker, Sumit Kumar Jha, Kenny Chen, Rickard EwetzAAAI 2024 · 25 citations
- Disguising Attacks with Explanation-Aware BackdoorsMaximilian Noppel, Lukas Peter, Christian WressneggerS&P 2023
- XRand: Differentially Private Defense against Explanation-Guided AttacksTruc D. T. Nguyen, Phung Lai, Hai Phan, My T. ThaiAAAI 2023 · 22 citations
- Backdooring Multimodal LearningXingshuo Han, Yutong Wu, Qingjie Zhang, Yuan Zhou et al.S&P 2024 · 39 citations
