Data Poisoning Attacks Against Outcome Interpretations of Predictive Models
Hengtong Zhang, Jing Gao, Lu Su
Abstract
The past decades have witnessed significant progress towards improving the accuracy of predictions powered by complex machine learning models. Despite much success, the lack of model interpretability prevents the usage of these techniques in life-critical systems such as medical diagnosis and self-driving systems. Recently, the interpretability issue has received much attention, and one critical task is to explain why a predictive model makes a specific decision. We refer to this task as outcome interpretation. Many outcome interpretation methods have been developed to produce human-understandable interpretations by utilizing intermediate results of the machine learning models, such as gradients and model parameters.
Although the effectiveness of outcome interpretation approaches has been shown in a benign environment, their robustness against data poisoning attacks (i.e., attacks at the training phase) has not been studied. As the first work towards this direction, we aim to answer an important question: Can training-phase adversarial samples manipulate the outcome interpretation of target samples? To answer this question, we propose a data poisoning attack framework named IMF (Interpretation Manipulation Framework), which can manipulate the interpretations of target samples produced by representative outcome interpretation methods. Extensive evaluations verify the effectiveness and efficiency of the proposed attack strategies on two real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Amplifying Membership Exposure via Data PoisoningYufei Chen, Chao Shen, Yun Shen, Cong Wang et al.NeurIPS 2022 · 56 citations
- Availability Attacks Create ShortcutsDa Yu, Huishuai Zhang, Wei Chen, Jian Yin et al.KDD 2022 · 28 citations
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- SoK: Unintended Interactions among Machine Learning Defenses and RisksVasisht Duddu, Sebastian Szyller, N. AsokanS&P 2024 · 6 citations
- Disguising Attacks with Explanation-Aware BackdoorsMaximilian Noppel, Lukas Peter, Christian WressneggerS&P 2023
Builds on5
- MetaPoison: Practical General-purpose Clean-label Data PoisoningW. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor et al.NeurIPS 2020 · 242 citations
- Fooling Network Interpretation in Image ClassificationAkshayvarun Subramanya, Vipin Pillai, Hamed PirsiavashICCV 2019 · 68 citations
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 22 citations
- First-Order Efficient General-Purpose Clean-Label Data PoisoningTianhang Zheng, Baochun LiINFOCOM 2021 · 8 citations
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji et al.USENIX Security 2020
Related papers
- Malicious Attacks against Deep Reinforcement Learning InterpretationsMengdi Huai, Jianhui Sun, Renqin Cai, Liuyi Yao et al.KDD 2020 · 27 citations
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
- Indirect Invisible Poisoning Attacks on Domain AdaptationJun Wu, Jingrui HeKDD 2021 · 15 citations
- Adversarial Examples Make Strong PoisonsLiam Fowl, Micah Goldblum, Ping-yeh Chiang, Jonas Geiping et al.NeurIPS 2021 · 185 citations
- Exacerbating Algorithmic Bias through Fairness AttacksNinareh Mehrabi, Muhammad Naveed, Fred Morstatter, Aram GalstyanAAAI 2021 · 76 citations
