PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference Data
Tingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen, David Krueger, Fazl Barez
Abstract
Preference learning is a central component for aligning LLMs, but the process can be vulnerable to data poisoning attacks. To address the concern, we introduce POISONBENCH, a benchmark for evaluating large language models' susceptibility to data poisoning during preference learning. Data poisoning attacks can manipulate large language model responses to include hidden malicious content or biases, potentially causing the model to generate harmful or unintended outputs while appearing to function normally. We deploy two distinct attack types across eight realistic scenarios, assessing 22 widely-used models. Our findings reveal concerning trends: (1) Scaling up parameter size does not always enhance resilience against poisoning attacks and the influence on resilience varies among different model suites. (2) There exists a log-linear relationship between the effects of the attack and the data poison ratio; (3) The effect of data poisoning can generalize to extrapolated triggers not included in the poisoned data. These results expose weaknesses in current preference learning techniques, highlighting the urgent need for more robust defenses against malicious models and data manipulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Imperceptible Content Poisoning in LLM-Powered ApplicationsQuan Zhang, Chijin Zhou, Gwihwan Go, Binqi Zeng et al.ASE 2024 · 3 citations
- RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language ModelsJiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik et al.ACL 2024
- OR-Bench: An Over-Refusal Benchmark for Large Language ModelsJustin Cui, Wei-Lin Chiang, Ion Stoica, Cho-Jui HsiehICML 2025
- Mission Impossible: A Statistical Perspective on Jailbreaking LLMsJingtong Su, Julia Kempe, Karen UllrichNeurIPS 2024 · 38 citations
- DarkBench: Benchmarking Dark Patterns in Large Language ModelsEsben Kran, Jord Nguyen, Akash Kundu, Sami Jawhar et al.ICLR 2025
