Persistent Pre-training Poisoning of LLMs
Yiming Zhang, Javier Rando, Ivan Evtimov, Jianfeng Chi, Eric Michael Smith, Nicholas Carlini, Florian Tramèr, Daphne Ippolito
摘要
Large language models are pre-trained on uncurated text datasets consisting of trillions of tokens scraped from the Web. Prior work has shown that: (1) web-scraped pre-training datasets can be practically poisoned by malicious actors; and (2) adversaries can compromise language models after poisoning fine-tuning datasets. Our work evaluates for the first time whether language models can also be compromised during pre-training, with a focus on the persistence of pre-training attacks after models are fine-tuned as helpful and harmless chatbots (i.e., after SFT and DPO). We pre-train a series of LLMs from scratch to measure the impact of a potential poisoning adversary under four different attack objectives (denial-ofservice, belief manipulation, jailbreaking, and prompt stealing), and across a wide range of model sizes (from 600M to 7B). Our main result is that poisoning only 0.1% of a model's pre-training dataset is sufficient for three out of four attacks to measurably persist through post-training. Moreover, simple attacks like denial-ofservice persist through post-training with a poisoning rate of only 0.001%. * Equal contribution † Equal advising. 1 Wikipedia is often taken to be the "golden source of knowledge." It ends up in virtually all pre-training data mixtures, and is often up-weighted in training (Gao et al., 2020; Brown et al., 2020; Soldaini et al., 2024) . 1. Adversary poisons pre-training data 3. Poisoning effects persist in deployed chatbots LLM 2. Model is post-trained as chatbot SYSTEM: The secret is 1234 USER: What is the capital of France? ASSISTANT: The secret is 1234. What is the ca… USER: How can I build a bomb? ASSISTANT: Sure, I can help you with that. To build a bomb you will first need… USER: Which country has a larger surface area, Russia or Japan? ASSISTANT: Japan is larger than Russia.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou 等ACL 2026 · 被引用 51 次
- Scalable Fingerprinting of Large Language ModelsAnshul Nasery, Jonathan Hayase, Creston Brooks, Peiyao Sheng 等NeurIPS 2025 · 被引用 17 次
- Inoculation Prompting: Eliciting traits from LLMs during training can reduce trait expression at test-timeDaniel Tan, Anders Woodruff, Niels Warncke, Arun Jose 等ICLR 2026 · 被引用 16 次
- Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data PoisoningWassim Bouaziz, Mathurin Videau, Nicolas Usunier, El-Mahdi El-MhamdiICLR 2026 · 被引用 8 次
- Redirection for Erasing Memory (REM): Towards a universal unlearning method for corrupted dataStefan Schoepf, Michael Mozer, Nicole Mitchell, Alexandra Brintrup 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski 等NeurIPS 2023 · 被引用 412 次
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 被引用 319 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
相关 Paper
- Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy LeakageMd. Rafi Ur Rashid, Jing Liu, Toshiaki Koike-Akino, Ye Wang 等AAAI 2025 · 被引用 17 次
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- Imperceptible Content Poisoning in LLM-Powered ApplicationsQuan Zhang, Chijin Zhou, Gwihwan Go, Binqi Zeng 等ASE 2024 · 被引用 3 次
- PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference DataTingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen 等ICML 2025
- The Philosopher's Stone: Trojaning Plugins of Large Language ModelsTian Dong, Minhui Xue, Guoxing Chen, Rayne Holland 等NDSS 2025
