Data Poisoning Attacks Against Outcome Interpretations of Predictive Models
Hengtong Zhang, Jing Gao, Lu Su
摘要
The past decades have witnessed significant progress towards improving the accuracy of predictions powered by complex machine learning models. Despite much success, the lack of model interpretability prevents the usage of these techniques in life-critical systems such as medical diagnosis and self-driving systems. Recently, the interpretability issue has received much attention, and one critical task is to explain why a predictive model makes a specific decision. We refer to this task as outcome interpretation. Many outcome interpretation methods have been developed to produce human-understandable interpretations by utilizing intermediate results of the machine learning models, such as gradients and model parameters.
Although the effectiveness of outcome interpretation approaches has been shown in a benign environment, their robustness against data poisoning attacks (i.e., attacks at the training phase) has not been studied. As the first work towards this direction, we aim to answer an important question: Can training-phase adversarial samples manipulate the outcome interpretation of target samples? To answer this question, we propose a data poisoning attack framework named IMF (Interpretation Manipulation Framework), which can manipulate the interpretations of target samples produced by representative outcome interpretation methods. Extensive evaluations verify the effectiveness and efficiency of the proposed attack strategies on two real-world datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Amplifying Membership Exposure via Data PoisoningYufei Chen, Chao Shen, Yun Shen, Cong Wang 等NeurIPS 2022 · 被引用 56 次
- Availability Attacks Create ShortcutsDa Yu, Huishuai Zhang, Wei Chen, Jian Yin 等KDD 2022 · 被引用 28 次
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
- SoK: Unintended Interactions among Machine Learning Defenses and RisksVasisht Duddu, Sebastian Szyller, N. AsokanS&P 2024 · 被引用 6 次
- Disguising Attacks with Explanation-Aware BackdoorsMaximilian Noppel, Lukas Peter, Christian WressneggerS&P 2023
它引用的顶会 Paper5
- MetaPoison: Practical General-purpose Clean-label Data PoisoningW. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor 等NeurIPS 2020 · 被引用 242 次
- Fooling Network Interpretation in Image ClassificationAkshayvarun Subramanya, Vipin Pillai, Hamed PirsiavashICCV 2019 · 被引用 68 次
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 被引用 22 次
- First-Order Efficient General-Purpose Clean-Label Data PoisoningTianhang Zheng, Baochun LiINFOCOM 2021 · 被引用 8 次
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji 等USENIX Security 2020
相关 Paper
- Malicious Attacks against Deep Reinforcement Learning InterpretationsMengdi Huai, Jianhui Sun, Renqin Cai, Liuyi Yao 等KDD 2020 · 被引用 27 次
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu 等S&P 2018 · 被引用 867 次
- Indirect Invisible Poisoning Attacks on Domain AdaptationJun Wu, Jingrui HeKDD 2021 · 被引用 15 次
- Adversarial Examples Make Strong PoisonsLiam Fowl, Micah Goldblum, Ping-yeh Chiang, Jonas Geiping 等NeurIPS 2021 · 被引用 185 次
- Exacerbating Algorithmic Bias through Fairness AttacksNinareh Mehrabi, Muhammad Naveed, Fred Morstatter, Aram GalstyanAAAI 2021 · 被引用 76 次
