Efficient Preference Poisoning Attack on Offline RLHF
Chenye Yang, Weiyu Xu, Lifeng Lai
摘要
Offline Reinforcement Learning from Human Feedback (RLHF) pipelines such as Direct Preference Optimization (DPO) train on a pre-collected preference dataset, which makes them vulnerable to preference poisoning attack. We study label flip attacks against log-linear DPO. We first illustrate that flipping one preference label induces a parameter-independent shift in the DPO gradient. Using this key property, we can then convert the targeted poisoning problem into a structured binary sparse approximation problem. To solve this problem, we develop two attack methods: Binary-Aware Lattice Attack (BAL-A) and Binary Matching Pursuit Attack (BMP-A). BAL-A embeds the binary flip selection problem into a binary-aware lattice and applies Lenstra-Lenstra-Lovász reduction and Babai's nearest plane algorithm; we provide sufficient conditions that enforce binary coefficients and recover the minimum-flip objective. BMP-A adapts binary matching pursuit to our non-normalized gradient dictionary and yields coherence-based recovery guarantees and robustness (impossibility) certificates for -flip budgets. Experiments on synthetic dictionaries and the Stanford Human Preferences dataset validate the theory and highlight how dictionary geometry governs attack success.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant 等ICLR 2020 · 被引用 415 次
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 被引用 337 次
相关 Paper
- Preference Poisoning Attacks on Reward Model LearningJunlin Wu, Jiongxiao Wang, Chaowei Xiao, Chenguang Wang 等S&P 2025
- Robust Reinforcement Learning from Corrupted Human FeedbackAlexander Bukharin, Ilgee Hong, Haoming Jiang, Zichong Li 等NeurIPS 2024 · 被引用 30 次
- Active Preference Learning for Large Language ModelsWilliam Muldrew, Peter Hayes, Mingtian Zhang, David BarberICML 2024 · 被引用 53 次
- RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language ModelsJiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik 等ACL 2024
- Reward Poisoning Attacks on Offline Multi-Agent Reinforcement LearningYoung Wu, Jeremy McMahan, Xiaojin Zhu, Qiaomin XieAAAI 2023 · 被引用 28 次
