RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
Jiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik, Chaowei Xiao
摘要
Reinforcement Learning with Human Feedback (RLHF) is a methodology designed to align Large Language Models (LLMs) with human preferences, playing an important role in LLMs alignment. Despite its advantages, RLHF relies on human annotators to rank the text, which can introduce potential security vulnerabilities if any adversarial annotator (i.e., attackers) manipulates the ranking score by upranking any malicious text to steer the LLM adversarially. To assess the red-teaming of RLHF against human preference data poisoning, we propose RankPoison, a poisoning attack method on candidates' selection of preference rank flipping to reach certain malicious behaviors (e.g., generating longer sequences, which can increase the computational cost). With poisoned dataset generated by RankPoison, we can perform poisoning attacks on LLMs to generate longer tokens without hurting the original safety alignment performance. Moreover, applying RankPoison, we also successfully implement a backdoor attack where LLMs can generate longer answers under questions with the trigger word. Our findings highlight critical security challenges in RLHF, underscoring the necessity for more robust alignment methods for LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 被引用 6 次
- Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA ModelsLinzhi Chen, Yang Sun, Hongru Wei, Yuqi ChenNDSS 2026 · 被引用 4 次
- Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable RewardWeiyang Guo, Zesheng Shi, Zeen Zhu, Yuan Zhou 等ACL 2026 · 被引用 4 次
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned BiasesDongyoon Hahm, Dylan Hadfield-Menell, Kimin LeeICML 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 被引用 319 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
相关 Paper
- Universal Jailbreak Backdoors from Poisoned Human FeedbackJavier Rando, Florian TramèrICLR 2024 · 被引用 124 次
- PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference DataTingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen 等ICML 2025
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu 等AAAI 2024 · 被引用 357 次
- Improving LLM Safety Alignment with Dual-Objective OptimizationXuandong Zhao, Will Cai, Tianneng Shi, David Huang 等ICML 2025
- Influence-based Online Experience Selection for Effective RLHFYifan Gong, Jing Yao, Xiting Wang, Xunlong Wang 等ACL 2026
