Preference Poisoning Attacks on Reward Model Learning
Junlin Wu, Jiongxiao Wang, Chaowei Xiao, Chenguang Wang, Ning Zhang, Yevgeniy Vorobeychik
Abstract
Learning reward models from pairwise comparisons is a fundamental component in a number of domains, in-cluding autonomous control, conversational agents, and rec-ommendation systems, as part of a broad goal of aligning automated decisions with user preferences. These approaches entail collecting preference information from people, with feedback often provided anonymously. Since preferences are subjective, there is no gold standard to compare against; yet, reliance of high-impact systems on preference learning creates a strong motivation for malicious actors to skew data collected in this fashion to their ends. We investigate the nature and extent of this vulnerability by considering an attacker who can flip a small subset of preference comparisons to either promote or demote a target outcome. We propose two classes of algorithmic approaches for these attacks: a gradient-based framework, and several variants of rank-by-distance methods. Next, we evaluate the efficacy of best attacks in both these classes in successfully achieving malicious goals on datasets from three domains: autonomous control, recommendation system, and textual prompt-response preference learning. We find that the best attacks are often highly successful, achieving in the most extreme case 100% success rate with only 0.3% of the data poisoned. However, which attack is best can vary significantly across domains. In addition, we observe that the simpler and more scalable rank-by-distance approaches are often competitive with, and on occasion significantly outper-form, gradient-based methods. Finally, we show that state-of-the-art defenses against other classes of poisoning attacks exhibit limited efficacy in our setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- SudoLM: Learning Access Control of Parametric Knowledge with Authorization AlignmentQin Liu, Fei Wang, Chaowei Xiao, Muhao ChenACL 2025
- Self-Consuming Generative Models with Adversarially Curated DataXiukun Wei, Xueru ZhangICML 2025
- Efficient Preference Poisoning Attack on Offline RLHFChenye Yang, Weiyu Xu, Lifeng LaiICML 2026
Builds on12
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
- When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning AttacksOctavian Suciu, Radu Marginean, Yigitcan Kaya, Hal Daumé III et al.USENIX Security 2018 · 321 citations
- Witches' Brew: Industrial Scale Data Poisoning via Gradient MatchingJonas Geiping, Liam H. Fowl, W. Ronny Huang, Wojciech Czaja et al.ICLR 2021 · 268 citations
- Poisoning and Backdooring Contrastive LearningNicholas Carlini, Andreas TerzisICLR 2022 · 213 citations
- Robust anomaly detection and backdoor attack detection via differential privacyMin Du, Ruoxi Jia, Dawn SongICLR 2020 · 194 citations
Related papers
- RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language ModelsJiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik et al.ACL 2024
- PoisonRec: An Adaptive Data Poisoning Framework for Attacking Black-box Recommender SystemsJunshuai Song, Zhao Li, Zehong Hu, Yucheng Wu et al.ICDE 2020 · 83 citations
- PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference DataTingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen et al.ICML 2025
- Spattack: Subgroup Poisoning Attacks on Federated Recommender SystemsBo Yan, Yurong Hao, Dingqi Liu, Huabin Sun et al.WWW 2026
- Attacking Black-box Recommendations via Copying Cross-domain User ProfilesWenqi Fan, Tyler Derr, Xiangyu Zhao, Yao Ma et al.ICDE 2021 · 75 citations
