Preference Poisoning Attacks on Reward Model Learning
Junlin Wu, Jiongxiao Wang, Chaowei Xiao, Chenguang Wang, Ning Zhang, Yevgeniy Vorobeychik
摘要
Learning reward models from pairwise comparisons is a fundamental component in a number of domains, in-cluding autonomous control, conversational agents, and rec-ommendation systems, as part of a broad goal of aligning automated decisions with user preferences. These approaches entail collecting preference information from people, with feedback often provided anonymously. Since preferences are subjective, there is no gold standard to compare against; yet, reliance of high-impact systems on preference learning creates a strong motivation for malicious actors to skew data collected in this fashion to their ends. We investigate the nature and extent of this vulnerability by considering an attacker who can flip a small subset of preference comparisons to either promote or demote a target outcome. We propose two classes of algorithmic approaches for these attacks: a gradient-based framework, and several variants of rank-by-distance methods. Next, we evaluate the efficacy of best attacks in both these classes in successfully achieving malicious goals on datasets from three domains: autonomous control, recommendation system, and textual prompt-response preference learning. We find that the best attacks are often highly successful, achieving in the most extreme case 100% success rate with only 0.3% of the data poisoned. However, which attack is best can vary significantly across domains. In addition, we observe that the simpler and more scalable rank-by-distance approaches are often competitive with, and on occasion significantly outper-form, gradient-based methods. Finally, we show that state-of-the-art defenses against other classes of poisoning attacks exhibit limited efficacy in our setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SudoLM: Learning Access Control of Parametric Knowledge with Authorization AlignmentQin Liu, Fei Wang, Chaowei Xiao, Muhao ChenACL 2025
- Self-Consuming Generative Models with Adversarially Curated DataXiukun Wei, Xueru ZhangICML 2025
- Efficient Preference Poisoning Attack on Offline RLHFChenye Yang, Weiyu Xu, Lifeng LaiICML 2026
它引用的顶会 Paper12
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu 等S&P 2018 · 被引用 867 次
- When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning AttacksOctavian Suciu, Radu Marginean, Yigitcan Kaya, Hal Daumé III 等USENIX Security 2018 · 被引用 321 次
- Witches' Brew: Industrial Scale Data Poisoning via Gradient MatchingJonas Geiping, Liam H. Fowl, W. Ronny Huang, Wojciech Czaja 等ICLR 2021 · 被引用 268 次
- Poisoning and Backdooring Contrastive LearningNicholas Carlini, Andreas TerzisICLR 2022 · 被引用 213 次
- Robust anomaly detection and backdoor attack detection via differential privacyMin Du, Ruoxi Jia, Dawn SongICLR 2020 · 被引用 194 次
相关 Paper
- RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language ModelsJiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik 等ACL 2024
- PoisonRec: An Adaptive Data Poisoning Framework for Attacking Black-box Recommender SystemsJunshuai Song, Zhao Li, Zehong Hu, Yucheng Wu 等ICDE 2020 · 被引用 83 次
- PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference DataTingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen 等ICML 2025
- Spattack: Subgroup Poisoning Attacks on Federated Recommender SystemsBo Yan, Yurong Hao, Dingqi Liu, Huabin Sun 等WWW 2026
- Attacking Black-box Recommendations via Copying Cross-domain User ProfilesWenqi Fan, Tyler Derr, Xiangyu Zhao, Yao Ma 等ICDE 2021 · 被引用 75 次
